We model contact at the end-effector with a diagonal, time-varying Cartesian
impedance law:
$$
\mathbf{f}(t) = \mathbf{K}(t) \odot \mathbf{e}(t) + \mathbf{D}(t) \odot \dot{\mathbf{e}}(t),
\qquad
\mathbf{e}(t) = \mathbf{x}_{\mathrm{eq}}(t) \ominus \mathbf{x}_f(t),
$$
where $\mathbf{x}_f(t)$ is the measured follower pose, $\mathbf{x}_{\mathrm{eq}}(t)$ is
the operator's unobserved commanded equilibrium, and $\mathbf{f}(t)$ is the
measured contact wrench. Given only the pair $(\mathbf{x}_f, \mathbf{f})$ — all a
pose-only interface ever records — this is unidentifiable: for
any axis and any positive stiffness, a compatible equilibrium exists that reproduces
the observed force exactly. Two unknowns, one equation, every timestep, no amount of
additional data collection fixes it.
Resolving it with an interface, not a model
Four-channel bilateral teleoperation actively couples a leader and a follower arm, so
the leader pose is itself an independent, physically measured proxy for the operator's
intended equilibrium:
$$
x_{\mathrm{eq}}(t) := x_l(t) \quad \text{(leader pose, independently measured)}.
$$
Substituting this collapses the ambiguity — $e(t) = x_l(t) \ominus x_f(t)$ is now
directly observed — and $(K(t), D(t))$ become identifiable by windowed regression
on $(e, \dot e, f)$, subject to an identifiability mask that
excludes (rather than imputes) timesteps with poor regression conditioning,
insufficient excitation, or force below the per-axis sensorless noise floor.
At static poses, the residual between measured joint torque and a zero-payload gravity
model is linear in the unknown tool's mass and mass-weighted center of mass, identified
by least squares over a handful of static poses. A per-axis residual bias model (a
random-Fourier-feature ridge regression over joint position and velocity) removes the
remaining configuration-dependent drift. The per-axis noise floor $\sigma_{f,i}$ used by
the identifiability mask is measured — not assumed — on held-out free-space
sweep sessions.
A SmolVLA backbone (vision-language model frozen, action expert fine-tuned) takes a
scene camera, a wrist camera, a language instruction (“wipe the {left, right}
mark {normally, firmly}”), and a force-history token (a trailing 500 ms of
6-DoF wrench), and outputs an action chunk of absolute target pose and log
stiffness, $[x_{\mathrm{eq}}, \log K]$, trained with flow-matching on pose and a masked
Huber loss on $\log K$. A 1 kHz Cartesian impedance controller tracks this
chunked output safely via temporal ensembling across overlapping chunks, log-space
stiffness rate-limiting, and a passivity-preserving energy tank that freezes stiffness
increases once depleted while always permitting decreases.