Compliance for Free: Learning Identifiable Impedance via Bilateral Teleoperation

Anonymous Author(s)
Anonymous Institution
Submitted to IEEE ICRA 2027 — anonymized for double-blind review

A pose-only teleoperation interface — VR controller, SpaceMouse, UMI gripper — records only where the operator moved the robot. Four-channel bilateral teleoperation already records more: the leader arm's pose is an independent, physical measurement of the operator's intended equilibrium, which is exactly the missing quantity that makes per-timestep, anisotropic compliance identifiable from real demonstrations — using only the manipulator's native joint-torque sensing, no dedicated force/torque hardware, at training or inference time.

0

Dedicated F/T or tactile sensors, training or inference

84.3–96.2%

In-contact timesteps retained by the identifiability mask (translational axes)

d = 0.89

Effect size of manner (“firmly” vs. “normally”) on our policy's realized contact force — the only one of five policies where it's significant

50.0%

Real-robot task success, best of five policies (fixed-stiffness baseline: 33.3%)

Abstract

Vision-language-action (VLA) policies typically output positions, but contact-rich manipulation — wiping, insertion, assembly — requires impedance too: how hard the robot pushes is as much an action as where it goes. We extract compliance without any dedicated force/torque sensor, using only the joint-torque sensing already present on the manipulator. Recent attempts to close this gap treat force as an additional input observation; the handful that produce force outputs must either assume a known task structure or infer compliance targets from simulation with privileged contact state, because pose-only teleoperation (VR controllers, SpaceMouse, UMI grippers) never measures what impedance the human was regulating.

Four-channel bilateral control does. We show that the leader-arm trajectory is an independent measurement of the operator's intended equilibrium pose, which uniquely resolves the equilibrium/stiffness identifiability ambiguity that has forced prior work into simulation. This yields per-timestep, anisotropic, real-world compliance labels at zero additional annotation cost. We use these labels to supervise a compliance head on a pretrained VLA, and show on real hardware that only the resulting compliance-output policy — not a force-as-input baseline, not a fixed-stiffness baseline, not a hybrid force/position baseline — changes what the arm actually does when the instruction says “firmly” instead of “normally.”

Contributions

  • A sensorless extraction procedure that resolves the pose-only identifiability ambiguity via bilateral teleoperation, yielding per-timestep compliance labels using only native joint-torque sensors — applicable, unmodified, to any existing or future bilateral-teleoperation dataset.
  • An offline benchmark showing the labels carry learnable signal: a trained predictor beats a per-axis constant predictor on every translational axis, for a 1.25× aggregate improvement.
  • A compliance-output VLA policy trained on these labels that safely tracks chunked impedance targets and is the only policy of five whose realized contact force responds to the manner instruction.

Method

The identifiability problem

We model contact at the end-effector with a diagonal, time-varying Cartesian impedance law:

$$ \mathbf{f}(t) = \mathbf{K}(t) \odot \mathbf{e}(t) + \mathbf{D}(t) \odot \dot{\mathbf{e}}(t), \qquad \mathbf{e}(t) = \mathbf{x}_{\mathrm{eq}}(t) \ominus \mathbf{x}_f(t), $$

where $\mathbf{x}_f(t)$ is the measured follower pose, $\mathbf{x}_{\mathrm{eq}}(t)$ is the operator's unobserved commanded equilibrium, and $\mathbf{f}(t)$ is the measured contact wrench. Given only the pair $(\mathbf{x}_f, \mathbf{f})$ — all a pose-only interface ever records — this is unidentifiable: for any axis and any positive stiffness, a compatible equilibrium exists that reproduces the observed force exactly. Two unknowns, one equation, every timestep, no amount of additional data collection fixes it.

Resolving it with an interface, not a model

Four-channel bilateral teleoperation actively couples a leader and a follower arm, so the leader pose is itself an independent, physically measured proxy for the operator's intended equilibrium:

$$ x_{\mathrm{eq}}(t) := x_l(t) \quad \text{(leader pose, independently measured)}. $$

Substituting this collapses the ambiguity — $e(t) = x_l(t) \ominus x_f(t)$ is now directly observed — and $(K(t), D(t))$ become identifiable by windowed regression on $(e, \dot e, f)$, subject to an identifiability mask that excludes (rather than imputes) timesteps with poor regression conditioning, insufficient excitation, or force below the per-axis sensorless noise floor.

Pipeline: bilateral teleoperation rig produces leader/follower pose, joint position, velocity and measured torque, which feed sensorless wrench estimation (payload ID, re-zero, RFF-ridge bias model), then windowed regression and an identifiability mask, producing a labeled dataset K(t), D(t), m(t).
From a bilateral rig to a labeled compliance dataset: sensorless wrench estimation, then windowed regression under an identifiability mask — no F/T sensor anywhere in the loop.

Sensorless wrench estimation

At static poses, the residual between measured joint torque and a zero-payload gravity model is linear in the unknown tool's mass and mass-weighted center of mass, identified by least squares over a handful of static poses. A per-axis residual bias model (a random-Fourier-feature ridge regression over joint position and velocity) removes the remaining configuration-dependent drift. The per-axis noise floor $\sigma_{f,i}$ used by the identifiability mask is measured — not assumed — on held-out free-space sweep sessions.

Compliance-output policy & controller

A SmolVLA backbone (vision-language model frozen, action expert fine-tuned) takes a scene camera, a wrist camera, a language instruction (“wipe the {left, right} mark {normally, firmly}”), and a force-history token (a trailing 500 ms of 6-DoF wrench), and outputs an action chunk of absolute target pose and log stiffness, $[x_{\mathrm{eq}}, \log K]$, trained with flow-matching on pose and a masked Huber loss on $\log K$. A 1 kHz Cartesian impedance controller tracks this chunked output safely via temporal ensembling across overlapping chunks, log-space stiffness rate-limiting, and a passivity-preserving energy tank that freezes stiffness increases once depleted while always permitting decreases.

Results

Dataset validity: identifiable, retained, and above the noise floor

Applied to real bilateral wiping demonstrations, the identifiability mask retains 84.3–96.2% of in-contact timesteps on the three translational axes (Table 1) — every axis clears a 25% validity floor by a wide margin. The measured per-axis sensorless noise floor lies comfortably below the labels' working range on every axis, and extracted stiffness traces show clean, phase-separated dynamics through approach, contact, and retract.

Table 1. Per-axis identifiability mask coverage, in-contact timesteps only.
Axis$f_x$$f_y$$f_z$$\tau_x$$\tau_y$$\tau_z$
Coverage (%)96.286.884.382.371.580.3
Extracted stiffness trace K(t) with phase annotation for one firm-wiping demonstration.

Extracted stiffness $K(t)$, one demonstration.

Identifiability mask coverage per axis, broken out by manner, across the full wiping dataset.

Mask coverage, per axis, by manner.

Sensorless wrench noise floor per axis, measured from free-space residuals on held-out sweep sessions.

Measured sensorless noise floor $\sigma_{f,i}$.

Offline compliance-prediction benchmark

Before training a policy on the extracted labels, we check they carry any signal at all: predicting $\log K_i(t)$ from a feature vector that deliberately excludes force. A learned predictor (random-Fourier-feature ridge) beats a per-axis constant baseline on every translational axis on a held-out, cross-session split — a 1.25× aggregate improvement in $\log K$ RMSE ($1.39\times$ on $f_x$, $1.03\times$ on $f_y$, $1.41\times$ on $f_z$) — rejecting the null hypothesis that the labels are pure noise, the necessary precondition for supervising a policy with them.

Offline compliance-prediction benchmark: per-axis RMSE on log K for constant, nearest-neighbour, and learned predictors, held-out test session.

Rotational axes are shown for completeness; their targets are box-saturated so their RMSE is correspondingly small.

The instruction survives, link by link

The central claim is that compliance extracted from bilateral demonstrations is a learnable, instruction-conditioned action. We test this as a three-link chain on the manner contrast (“firmly” vs.\ “normally”) on the in-plane axis $f_x$: does the adverb show up in the extracted label, then in the policy's own output, then in the arm's realized force?

Table 2. Instruction-conditioned compliance, measured end to end (Cohen's $d$, two-sided Mann–Whitney $p$).
LinkQuantityfirmlynormally$d$$p$
L1Extracted label $K_x$146.2127.20.760.002
L2Policy's commanded $K_x$141.2116.40.980.019
L3Realized RMS contact force9.126.400.890.023

All three links carry the instruction: the demonstrations encode the adverb in stiffness ($d=0.76$), the policy reproduces that dependence in its own output ($d=0.98$), and it survives all the way to realized contact force ($d=0.89$). The policy also commits to genuine anisotropy — a median in-plane/normal stiffness ratio of $2.50\times$ in the contact frame, far beyond the ${\approx}1.2\times$ in the pooled label distribution — so this is not regression to the label mean.

But predicting compliance is not by itself sufficient: of five policies, only ours modulates realized force with the instruction (Fig. below). B3 moves its commanded stiffness with manner ($d=2.26$ on $f_y$) but its realized force does not follow ($d=0.17$) — the rest of the policy still has to track what it predicts.

Left: RMS contact force per rollout by manner for all five policies, only the compliance-output policy (B4, ours) shows a significant separation, d=0.89, p=0.023. Right: median contact force over time in rollout for our policy, normally vs firmly, showing sustained separation through the wiping phase.

The manner word changes what the arm does — and only for the compliance-output policy. The separation is sustained through the wiping phase, not driven by isolated peaks, and both manners converge on approach/retract where the instruction shouldn't matter.

Real-robot task performance

All policies are evaluated on a whiteboard-wiping task, $n=24$ rollouts per policy (3 seeds, a $2\times2$ referent$\times$manner instruction grid). B0: position output, fixed stiffness. B1: position output, oracle per-axis constant stiffness fit to our label distribution. B2: ForceVLA-style force-as-input, position output. B3: ForceVLA2/Force-Policy-style hybrid force+position output. B4 (ours): compliance output, force input.

Table 3. In-distribution task performance, five policies ($n=24$/policy, 3 seeds).
PolicySuccess (M1)Ink removal % (M2)
B0 — pos., fixed $K$33.3%32.6 ± 34.9
B1 — pos., oracle const. $K$0.0%1.3 ± 3.0
B2 — force-input, pos. out41.7%37.3 ± 31.5
B3 — force-input, hybrid out20.8%18.4 ± 28.8
B4 (ours) — compliance out50.0%40.8 ± 35.3

Contact-force behavior is more informative than success alone: our policy attains the highest ink removal at a peak force 8.6 N lower and an RMS force 3.1 N lower than B0, tripping one protective stop in 24 rollouts against B0's six.

Table 4. Contact-force behavior, same rollouts as Table 3.
PolicyPeak force (N)RMS force (N)Protective stops
B034.9 ± 14.010.9 ± 2.26/24
B121.4 ± 7.26.5 ± 2.30/24
B232.3 ± 9.012.9 ± 6.52/24
B319.6 ± 8.27.1 ± 3.10/24
B4 (ours)26.3 ± 9.17.8 ± 3.31/24
Qualitative rollouts of the compliance-output policy, scene and wrist camera, across the 2x2 referent by manner instruction grid, five phases per rollout: approach, pre-contact, contact, wiping, retract.

Qualitative rollouts of the compliance-output policy across the referent×manner instruction grid — scene camera (top) and wrist camera (bottom), five phases per rollout.

Controller validation

Before trusting any of the above, we verify the low-level controller actually responds to a commanded stiffness: a 20 s sinusoidal probe on all six axes shows realized stiffness tightly tracking the commanded signal, with the energy-tank trace confirming the passivity budget behaves as designed.

Controller validation on real hardware: commanded vs. realized stiffness under a 20 second sinusoidal probe on all six axes, with the accompanying energy tank trace.

Is the force channel actually used?

We freeze the force-history token at its first observed value for the remainder of the episode and re-run. All three force-consuming policies degrade when force is frozen (B4 by $-8.3$ points), consistent with the live force signal being causally used rather than ignored — though with $n=24$ per policy none of these drops individually reaches significance ($p \geq 0.73$).

Table 5. Causal force ablation: force history frozen at its first-observed value vs. the live condition.
PolicyLive (M1)FrozenΔ$p$
B241.7% (10/24)35.7% (5/14)−6.0 pts1.00
B320.8% (5/24)16.7% (2/12)−4.2 pts1.00
B450.0% (12/24)41.7% (5/12)−8.3 pts0.73

Why B1's oracle stiffness fails

Grid-searching, per axis, the single constant stiffness that best fits our label distribution gives unremarkable translational constants (141–158 N/m, well inside the controller's realizable range) but rotational constants that land on the extraction's lower bound — the rotational impedance in these demonstrations is softer than the controller's minimum commandable rotational stiffness. Deployed, B1 cannot hold orientation against the board, which is why it scores 0% and not evidence against constant impedance as such: the honest constant-impedance control in this study is B0 (hand-set, well-posed, 33.3%). The claim is narrower and, we think, sharper — a well-tuned constant impedance still cannot modulate force on instruction (Table 2).

Table 6. Constant-stiffness fit (B1), per-axis best constant (contact frame, grid-searched on validation split).
AxisBest constant $K_i$Realizable range
$f_x$141.1 N/m[50, 1500] N/m
$f_y$149.5 N/m[50, 1500] N/m
$f_z$158.4 N/m[50, 1500] N/m
$\tau_x, \tau_y, \tau_z$all at the 5 N·m/rad lower bound

Discussion & Limitations

The comparison against B2 (ForceVLA-style: force as input, position as output) isolates the distinction this paper cares about: observing force is not the same as acting compliantly. Facing the wiping task's rigid physical constraints, position-only policies either under-press (poor ink removal) or inject excessive energy at chunk boundaries (protective stops). Actively modulating the impedance target $K$ lets our policy absorb these geometric mismatches natively rather than fighting them.

  • All results are on a single task (whiteboard wiping); the identifiability argument and extraction method are not claimed to be task-specific, but this has not been verified on a second task.
  • Sensorless wrench estimation is configuration-dependent, so both calibration and teleoperation were constrained to a fixed nullspace posture and working volume — a real kinematic trade-off against dedicated F/T hardware, traded here for zero end-effector bulk, wiring, or cost.
  • The offline benchmark uses a small, session-level, cross-operator split; the per-axis improvement margin is thin on some axes and should be re-verified as more sessions are collected.
  • The M7 causal force ablation shows the expected direction of degradation on every force-consuming policy, but at $n=24$ per policy none of the individual drops reaches significance.
  • The wrench noise floor is a free-space measurement, not an independently verified calibration against a known applied force.

Pose-only interfaces remain the most accessible way to collect kinematic demonstration data, but they are mathematically insufficient for impedance learning — no amount of scale fixes an identifiability problem. Because the extraction procedure consumes only signals every four-channel bilateral rig already logs, it applies unchanged to any existing or future bilateral-teleoperation dataset: an identifiable compliance signal that can be recovered without new data collection.

Conclusion

Standard pose-only teleoperation interfaces record only realized poses, which leaves the operator's intended equilibrium and stiffness confounded — a mathematical bottleneck, not a data-scale one. Four-channel bilateral teleoperation collapses this ambiguity for free: the leader arm already measures what the operator intended. Supervising a compliance head on these extracted labels lets a VLA policy dynamically regulate contact impedance and outperform rigid, force-input baselines on real hardware. Future work: extending this pipeline to more dynamic, impulsive contact tasks, and cross-embodiment distillation of bilateral-derived compliance priors onto lower-cost hardware.

BibTeX

@inproceedings{anonymous2027complianceforfree,
  title     = {Compliance for Free: Learning Identifiable Impedance via
               Bilateral Teleoperation},
  author    = {Anonymous Author(s)},
  booktitle = {Under review, IEEE International Conference on Robotics and
               Automation (ICRA)},
  year      = {2027},
  note      = {Anonymized for double-blind review}
}