← Servo7 R&D Index
Tests 2026-08-16 · Finalized 2026-08-17 FR3 fw v3.9.6 (from v3.9.5) · FR20 via r1lite atlas PR #204 ↗

The Fairino ServoJ Runaway: Reproduced, Explained, Fixed

A setpoint jump of ~1–1.5° inside one 10 ms tick latches a controller fault — and the controller then executes an unbounded sweep anyway. This report maps the trigger on the real FR3 and FR20, and describes the acceleration-bounded stream shaper and lag-compensated divergence guard that now make it impossible to trigger from the Atlas side.

Test conditions. The diagnosis runs used a standalone Rust tool talking directly to the controller — no Atlas software and no software safety in the loop, by design. The operator held the physical e-stop at all times. The report ends with the new Atlas safety system built and validated from these findings.
Runaway trigger
~1–1.5°
setpoint jump in one 10 ms tick (a Δv of 100–150 °/s)
Worst unshaped sweep
88.6°
from a 2.5° step, peaking at 196 °/s, uncommandable for ~0.9 s
Shaped 10° jumps
0 faults
10 of 10 runs clean on the real FR3, arrival in ≈328 ms
Teleop after fix
100 °/s
sustained, zero trips — 2.5× the old false-trip ceiling
01 · Conclusions

Eleven findings, each measured on real hardware

1

The Fairino ServoJ runaway is reproduced, measured, and explained. The trigger is a velocity discontinuity in the setpoint stream: a jump of ~1.0–1.5° inside one 10 ms command period, implying a velocity change of 100–150 °/s in a single tick. The controller detects it, latches a fault — and then still executes an unbounded sweep.

2

The transport does not matter. The runaway is identical over UDP (cmdType=1) and XML-RPC (cmdType=0): a 2.5° step produced an 88.57° sweep on UDP and 88.56° on XML-RPC. The defect lives in the ServoJ execution engine, not in a communication path.

3

The threshold band is unstable. At a 1.0° step the arm tracked cleanly once and ran away 13.9° on an equal step minutes later. At 1.5° and above the fault fires every time. Safety margins must therefore sit below 1.0° per tick.

4

High speed is safe when the motion is smooth. With a trapezoid profile (acceleration ≤ 600 °/s²) the arm ran clean at 40–120 °/s. At 120 °/s the setpoint advances 1.2° per tick (inside the step trigger band) and the arm trails the setpoint by 9.3° (six times the trigger size) — nothing bad happened, twice. The trigger is acceleration — not speed, and not the tracking gap.

5

Very large steps are safe in a different way. A 30° step is refused outright: every command is rejected and the arm stays frozen. The dangerous zone is mid-size steps, measured from ~1.5° to 4.8° (the 5–30° band is untested raw).

6

Once the runaway starts, no command can stop it. The controller rejects all commands with errcode 14 for ~0.9 s while the arm moves. Only the physical e-stop or the end of the controller's own motion stops the arm.

7

The real FR3's tracking lag is 76–79 ms, constant from 5 to 120 °/s. The SimMachine shows 37–40 ms. This lag makes Atlas's old command guard trip falsely at 38–48 °/s, because the guard measures lag × speed, not true divergence.

8

The fix is validated on the real arm. An acceleration-bounded stream shaper on the follower side (state = last sent pose and velocity; per tick it clamps velocity at 130 °/s and velocity change at 500 °/s² × cmdT, with a braking bound so it stops on target without overshoot) makes the runaway impossible to trigger. Raw jumps of 10° (ten repeats) and 15° — 4–10× the runaway threshold — passed through the shaper onto the real FR3: zero rejects, zero controller faults, smooth arrivals in 328 and 379 ms. Unshaped, a 2.5° jump gave an 88.6° sweep at 196 °/s. The concept is shipped in the Atlas node driver.

9

The runaway is joint-independent, but not joint-uniform. The proximal joints (indices 0–2: base, shoulder, elbow) trigger earliest; the wrist joints have more room, closer to 5.5° per step. When the gap is large on several joints at once, several joints sweep at once.

10

The real FR20's tracking lag is measured: L = 109–111 ms, constant across joints (J1 and J6), speeds (5–20 °/s), and directions — the same pure-transport-delay model as the FR3 (78 ms) and the SimMachine (39 ms). The old ~130 ms estimate from session logs was close but high.

11

The Atlas safety system was rebuilt from these numbers and is live on the Rust node path: a velocity/acceleration-bounded stream shaper at the wire plus a lag-compensated divergence guard replace the old command guards, which measured lag as if it were divergence. Validated on the real FR3 at a sustained ~100 °/s measured, with zero trips and zero controller faults. See PR #204.

02 · Background

The system under test

The FAIRINO controller has three independent interfaces:

Host
Atlas node driver
or servoj-lag (this report)
FAIRINO controller
FR3 fw v3.9.6
FR20 (via r1lite)
XML-RPC (TCP 20003) — commands and getters: RobotEnable, Mode, ServoMoveStart, ResetAllError, plus a synchronous ServoJ call (the SDK's cmdType=0).  ·  CNDE (TCP 20005) — a binary state frame pushed every 10 ms: joint positions, robot mode, robot state, and the fault codes main_code/sub_code.  ·  ServoJ stream (UDP 20007) — one joint setpoint per command period (cmdT, 10 ms in all tests); the controller replies per datagram, and a reply with errcode:N (N ≠ 0) is a rejection.

Atlas streams ServoJ setpoints at 100 Hz. A safety guard (command_guard / MaxJointDeltasGuard) compares every outgoing setpoint with the last measured pose; if the difference exceeds a per-joint limit (3.0° in our configs), it raises an e-stop. Production had seen an unexplained "random motion runaway" that this guard caught, with an unknown trigger. A separate theory said the guard's delta is dominated by controller lag, so it limits speed instead of detecting divergence. Both questions are now answered: the trigger is a velocity discontinuity (finding 1), and the old guard indeed measured lag, not divergence (finding 7).

03 · Instrumentation

The test tool: servoj-lag

servoj-lag is a standalone binary in the Atlas repo at rust/fairino-sdk/src/bin/servoj_lag.rs. It talks straight to the controller and contains no safety of any kind — deliberately, because the goal was to observe the runaway. Both the sender and the recorder use one host clock, so there is no clock-sync problem: the measured lag is send-instant to change-observed-on-CNDE, the exact vantage point the Atlas guard has.

What it does under the hood, in order
  1. Init — the "posture walk", a mirror of the production driver: ServoMoveEnd (close any old servo window), Mode(manual), RobotEnable(1), Mode(automatic), RobotEnable(1), ServoMoveStart, all over XML-RPC, each retried on the controller's transient NOT_READY code (14). A real controller rejects all ServoJ without this walk; the SimMachine does not care.
  2. CNDE recorder thread: records every 10 ms state frame with a host timestamp — the six joints plus robot_mode, robot_state, main_code, sub_code.
  3. Acceptance-gated warm-up: the sender streams "hold the current pose" until the controller accepts the stream (an accept reply, or 300 ms with zero rejections), then holds one more second. Reason: RobotEnable(1) returns 0 immediately, but the robot needs ~1.4 s before it executes ServoJ (it answers errcode 101, "robot is not enabled", in that window). Without this gate a test starts against a dead stream, the setpoints run ahead, and the enable window itself creates a runaway trigger — which is exactly how we met the runaway the first time.
  4. The experiment — step mode or ramp mode, below.
  5. Teardown: ServoMoveEnd, then 1.5 s of extra recording (to catch post-stop drift), then disable, unless the keep flag is set.
  6. Output: a console report plus a CSV per run, named servoj_lag_<mode>_<UTC timestamp>.csv, with three row kinds: send (each setpoint with its send time), state (each CNDE frame with the status words), delta (|setpoint − last measured| at each send tick — exactly the quantity the Atlas guard computes). Header lines record the full run config.

Step mode — the gap injector

servoj-lag <IP> [joint] [step_deg] [reps] [cmd_ms] [keep] [rpc] # example: servoj-lag 192.168.58.3 0 1.5 1 10 keep

Arguments: joint index (0 = J1, the base), step size in degrees, number of steps, command period in ms, keep = leave the robot enabled at exit, rpc = send each setpoint as an XML-RPC call (cmdType=0) instead of UDP. After the warm-up, the streamed setpoint of the chosen joint changes by the step size in one tick, and the new value is re-sent every 10 ms for 600 ms; then it steps back. With reps > 1 the level cycle is +X, 0, −X, 0. There is no ramp between levels — the change is instantaneous by design, injecting a gap of exactly X degrees at a known instant. Reported per step: first_move_ms (first CNDE frame that left the pre-step pose), arrive_ms (first frame within 10% of the target), peak_delta (max |setpoint − measured|), and guard_lag_ms (how long the guard-view delta stayed elevated).

Ramp mode — the smooth-motion prober

servoj-lag <IP> ramp [joint] [sweep_deg] [speeds] [cmd_ms] [keep] [rpc] # example: servoj-lag 192.168.58.3 ramp 0 60 120 10 keep

Sweeps the joint from its current pose out by sweep_deg degrees and back, once per listed speed (comma list, °/s). Each pass is a trapezoid: velocity rises linearly over 200 ms, cruises constant, falls over 200 ms. The velocity change per tick is only v/20, so there is never a discontinuity. Reported per pass: mean and peak guard delta over the cruise, the implied lag L = delta / v, a midpoint-crossing lag cross-check, and the overshoot after the stop. The summary fits delta ≈ v × L; a constant L across speeds means the guard delta is pure lag.

04 · Chronological path

Eight experiments, from baseline to fix

1SimMachine baseline

Against the v3.9.6 Docker SimMachine: step-response lag ~33 ms; sustained-motion L = 39–40 ms, constant from 5 to 40 °/s, identical on the FR3 and FR20 sim models, identical on both transports (37.5 ms via XML-RPC). Delta = v × L exactly. No step size ever caused a runaway in sim — the sim executes any jump almost instantly after a fixed delay. Conclusion: in sim the guard delta is 100% lag, 0% divergence. Also disproved: an old note claimed "75 ms sim lag"; that number came from an earlier probe at cmdT = 20 ms and does not apply at cmdT = 10 ms.

2First runaway encounter

A ramp run started while the robot was still inside its enable window (~1.4 s of errcode 101 after RobotEnable returns 0). The setpoints advanced ~4.8° before the controller began to execute. When it did, the arm swept ~158° at 180–209 °/s, straight past the +30° target, and only the operator's e-stop stopped it (errcode 31, "e-stop button", requiring a control-box power cycle). Reproduction: a run that starts with the robot disabled produces the runaway (2 out of 2); a run that starts with the robot already enabled is clean at the same speeds, with the same L. The controlling variable is the gap at the moment execution starts. This finding produced the acceptance-gated warm-up and the keep flag in the tool.

3Controlled gap injection over UDP

One step per run, joint 0, robot enabled and tracking before each step:

StepResultPeak speedTravel
0.5°clean, smoothed (arrive 141 ms)n/a0.5°
1.0°clean once, then RUNAWAY on an equal return step75 °/s13.9°
1.5°RUNAWAY142 °/s33°
2.5°RUNAWAY196 °/s88.6°
~4.8°RUNAWAY (from experiment 2)180–209 °/s158°+, e-stopped

The constant pattern in the CSVs: the controller latches fault main_code=1, sub_code=32 at the step instant, before any motion. Then it rejects every command with errcode 14 for ~0.9 s while the arm sweeps. The peak speed equals the step's implied velocity (step / cmdT), clamped at the joint maximum. The travel is roughly 14–35× the step. The motion stops on its own.

4Same test over XML-RPC (cmdType=0)

The tool's rpc flag sends every setpoint as a synchronous XML-RPC ServoJ call; each call returns its own accept/reject code, so the tally is exact. At 1.5° the fault fired but the motion stayed bounded on target (the threshold band is a coin flip). At 2.5°: runaway, travel 88.56° — identical to the UDP run's 88.57°. Lag via RPC equals lag via UDP. cmdType does not matter; the runaway is transport-agnostic.

5The 30° step

A missing ramp keyword ran step mode with a 30° step. The controller refused. The teach pendant showed the fault in plain text: "the command speed in the joint space of axis 1 exceeds the speed limit and can be reset." The controller's speed check works at giant steps and fails dangerously at mid-size steps.

6Smooth ramps to high speed

Trapezoid ramps, robot enabled at start, one speed per run, sweep 30° (60° for the fast runs so the cruise stays long enough to measure):

SpeedPer-tick deltaTracking gap (v×L)Result
40 °/s0.4°3.1°clean, L = 77–79 ms
60 °/s0.6°4.8°clean
80 °/s0.8°6.2°clean (see note)
100 °/s1.0°7.7°clean
120 °/s1.2°9.3°clean, twice, overshoot ≤ 0.02°

Note: one 80 °/s run was stopped by the controller's own Cartesian TCP speed safety, set conservatively for the cell. That is a separate, configurable protection layer; after adjustment the run was clean.

Read the 120 °/s row against the step table: the per-tick delta (1.2°) is inside the step trigger band, and the tracking gap (9.3°) is six times the runaway threshold. Nothing happened. The trigger is the velocity discontinuity — the acceleration impulse — not the per-tick delta and not the gap. L stayed 76–79 ms across the full range, so the lag is a pure transport delay.

7The fix: acceleration-bounded stream shaping

The tool gained a shaped flag. With it, every frame passes through a stream shaper before the wire. The shaper works per joint, per tick:

  1. Its state is the last sent setpoint and its velocity. It never reads the measured pose, so the 78 ms lag cannot blind it.
  2. It computes the velocity toward the raw target and clamps it three times: at v_max (130 °/s), at the discrete braking bound (so it can always stop exactly on the target, no overshoot), and at a velocity change of a_max × cmdT with a_max = 500 °/s² — the proven-clean acceleration.
  3. When the input already respects the bounds, the shaper changes nothing. When the input jumps, the wire sees a clean trapezoid instead of a discontinuity.

Unit tests assert the bounds on every tick of a 2000-tick random shock stream (jumps to ±30°, random freezes). Live results on the real FR3, joint 0, raw jumps injected through the shaper:

Raw jumpRunsRejectsController faultsArrival
10°100nonemean 328 ms (322–335)
15°10none379 ms

The same class of input, unshaped, at only 2.5°, produced an 88.6° sweep at 196 °/s. The shaped runs also probe the previously untested 5–30° band safely, because the wire never sees the raw jump. Two numbers from these runs matter for the next design step:

8The FR20, measured with the same protocol

After the new system shipped, the same lag measurement ran on the real FR20 (from the r1lite mini PC). Safety rules for a production arm: ramp mode only (a ramp cannot create the velocity discontinuity that triggers the runaway; its per-tick velocity change is v/20), the shaped flag as a second bound on the tool's own output, small sweeps (10–15°), low speeds (5–20 °/s — enough, because L is speed-independent), the wrist joint J6 first (smallest motion envelope), then J1 to confirm.

Result: L = 109–111 ms on every pass (J6 and J1, 5/10/20 °/s, both directions), zero rejects, zero overshoot. The midpoint cross-check gave the same 103–117 ms. The FR20 controller also accepted UDP ServoJ directly, so its firmware needs no upgrade.

05 · Interactive

The shaper, live

This is the exact algorithm now sitting at the wire, running in your browser at the real tick rate (cmdT = 10 ms). Inject a raw setpoint jump and watch what the controller receives: the raw stream steps instantly — the discontinuity that latches the fault — while the shaped stream turns the same input into a bounded trapezoid. Drag the sliders or pick a preset from the measured campaign.

This jump sent raw

This jump through the shaper

Table view — every tick of this run
tickt (ms)raw (°)shaped (°)vel (°/s)Δv/tick (°/s)

Simulated ideal wire trajectory from the shipped equations (§08). The physical arm follows the shaped stream ~78 ms behind on the FR3 (~111 ms on the FR20), so measured arrivals — 328 ms mean for the live 10° runs — sit slightly above the wire arrival shown here. Hover or focus the chart and use ←/→ to read exact values; every value is also in the table view.

06 · Method

How the lag L is actually measured

The number every guard decision rests on comes from the tool's ramp mode:

  1. One host clock stamps everything. The sender thread stamps each outgoing setpoint; the recorder thread stamps each incoming CNDE state frame. There is no clock-sync problem, and the measurement includes exactly what the guard experiences: controller execution lag plus state-reporting delay.
  2. The tool streams a trapezoid sweep at a constant cruise speed v. At every 10 ms send tick it computes delta = |setpoint just sent − latest measured joint|.
  3. During the cruise, the arm is always executing the setpoint from L seconds ago. In those L seconds the setpoints advanced by exactly v × L. So a healthy arm shows a constant delta during cruise, and that delta is the lag in disguise: L = mean cruise delta / v. One pass gives one estimate; the summary fits delta ≈ v × L through the origin across all passes.
  4. Two validity checks. First, an independent cross-check: the time shift between the commanded stream crossing the sweep midpoint and the measured stream crossing it must equal the same L. Second, the model check: L must come out constant across speeds, joints, and directions. If L grew with speed, the "lag" would contain real dynamics and the simple compensation would be wrong. On all three platforms it is constant: sim 39 ms, FR3 78 ms, FR20 111 ms.
07 · Productized

The new Atlas safety system

This is what changed in the Atlas safety envelope, and why — the productized form of everything above, shipped in PR #204.

What was removed, and why

Three mechanisms are deleted from the Rust node path (code, configs, and the safety policy):

All three shared one flaw: they compared the raw commanded target against the measured pose, and that difference is dominated by physics, not by danger. At speed it contains the tracking lag (v × L) and, once the shaper exists, also the shaper's braking trail (v²/2a). Measured consequence: the old guard falsely e-stopped clean teleop at 27–48 °/s (depending on arm), and cap_delta was worse than useless — after a network stall it emitted a ~3° instantaneous step, which is inside the measured runaway trigger band. The protections themselves carried the trigger.

The old wire names (command_guard, command_speed) still exist as graceful no-ops: an old frontend toggle or runtime edit gets a "retired" warning instead of a crash.

What replaced them: two components with strictly separated roles

Producers
teleop · playout · hold
go-home · replay
Stream shaper
never trips, always shapes
|v| ≤ v_max · |Δv| ≤ a_max·T
braking bound, no overshoot
Arm
tracks L behind
CNDE measured pose
Divergence guard
never shapes, only trips
|s − m| > margin + |v|·lag_s → e-stop
Independent nets (kept)
joint-limits · workarea
joint_speed · eef_speed
Open-loop bounding at the wire; closed-loop tripping against the measurement; measured-state observers underneath. A Fairino cannot boot half-protected: the config refuses to construct a shaper without its divergence guard.

Component 1 — the stream shaper (never trips, always shapes). It sits at the single point every producer passes before ServoJ: teleop, playout, hold frames, go-home, replay. Its state is the last sent setpoint and its velocity; it never reads the measured pose, so lag cannot blind it. Per tick it bounds the wire three ways: velocity ≤ v_max, velocity change ≤ a_max × cmdT, and a braking bound so it stops on the target without overshoot (full math in §08). Any input, however broken, becomes a physically legal trapezoid. It primes on the measured pose whenever a stream engages — which closes the enable-window trigger — and its state is dropped on stream exit, so a re-engage never ramps out of a stale pre-stop target. A raw input jump beyond the maximum braking trail (~14°) logs a warning; it never trips.

Component 2 — the divergence guard (never shapes, only trips). The closed-loop check the shaper cannot provide, because the shaper cannot see the arm. It trips when

$$|s(t) - m(t)| \;>\; \text{margin} \;+\; |v_{\text{shaper}}| \cdot \text{lag}_s$$

with s the shaped setpoint just sent, m the measured pose, and v the shaper's own commanded velocity — already in its state, so the lag term costs one multiplication; no ring buffer, no time shift. The allowance breathes with the motion: at rest it is the bare margin (~3°, as tight as the old guard); at full speed it grows by exactly the healthy lag term. Clean tracking can never trip it; an arm that moves while we command stillness trips within the margin; a fast controller runaway (140–200 °/s) crosses the allowance within one or two ticks (~20 ms detection).

Two floors protect the tuning against reintroducing false trips:

Kept unchanged, as the independent nets: the joint-limits gate and the workarea gate on the command side (they validate the raw target and trip), and the measured-state observers joint_speed (trip above the wire ceiling: nothing we can command exceeds v_max, so a faster measured arm means the controller acts on its own) and eef_speed. These watch the arm, not the commands, so they hold even if everything above them is misconfigured.

Fail-closed pairing: a config with a shaper but no divergence guard refuses to construct, and the safety policy (config/safety_policy.yaml) requires joint_speed, stream_shaper, and divergence_guard keys for every strict robot type. A Fairino cannot boot half-protected.

Envelope reporting: no nameless e-stops

The mechanics of both components run on the driver's write path, but the SafetyEnvelope remains the single reporting surface. It carries their configuration echo and the divergence trip latch: any_tripped includes the divergence guard, the telemetry safety dict gained the wire keys stream_shaper (bounds) and divergence (margin, lag, tripped flag, last reason), and the operator reset clears the latch. Because the bridge, the frontend, and the test harness all identify tripped guards by scanning that dict, a divergence e-stop is always named: the "e-stop with an empty guard list" pattern is gone.

Per-robot tuning, all from measured numbers

PlatformMeasured LWire v_maxWire a_maxlag_smarginjoint_speed net
FR3 real78 ms110 °/s500 °/s²0.093.0120
FR20 real111 ms100 °/s450 °/s²0.133.5110
SimMachine39 msas per familyas per family0.053.0110–120

The leader (SpaceMouse) caps were raised to match the follower ceilings; they are comfort shaping only — the follower does not trust them.

Validation state

  1. Real FR3, SpaceMouse teleop through the full Atlas stack: sustained ~100 °/s measured joint speed, smooth, zero trips, zero controller faults — two and a half times the old 40 °/s ceiling.
  2. Real FR20: lag measured and tuned (experiment 8); the full teleop session is the next step.
  3. CI: the old command-guard harness checks were replaced by harness_shaper_check_rust.py with the inverted contract — a 30° raw punch into the node must be absorbed (speed-bounded wire, clean arrival, zero trips). The divergence trip side is covered by unit tests plus a driver-level test where the simulated arm "moves on its own" and the trip is asserted by name.

Honest limits

08 · Theory

The math and physics of the fix

This section states what the shaper is, in equations, and how it relates to classical control concepts.

The equations

Per joint. The state is the last sent position \(s_k\) and its velocity \(v_k\). The constants are the velocity limit \(v_{max}\), the acceleration limit \(a_{max}\), and the tick time \(T\) (cmdT, 10 ms). The input is the raw target \(x_k\), which can jump in any way. Each tick:

$$d_k = x_k - s_k \qquad \text{(position error)}$$

$$v_{brk} = \sqrt{2\, a_{max}\, |d_k|} \qquad \text{(braking curve)}$$

$$v_{des} = \mathrm{sign}(d_k)\cdot \min\!\left(v_{max},\; v_{brk}\right)$$

$$v_{k+1} = \mathrm{clamp}\!\left(v_{des},\; v_k - a_{max}T,\; v_k + a_{max}T\right)$$

$$s_{k+1} = s_k + v_{k+1}\, T$$

The guarantees hold for any input sequence: \(|v_k| \le v_{max}\), \(|v_{k+1} - v_k| \le a_{max} T\) (bounded acceleration), no overshoot of a static target, and the identity function when the input already respects the bounds.

Where the square root comes from

Uniform-acceleration kinematics: a mass that brakes at constant \(a\) from velocity \(v\) needs the distance \(d = v^2 / (2a)\) to stop. Turned around, the highest velocity from which you can still stop within the remaining distance \(d\) is

$$v = \sqrt{2\, a_{max}\, d}.$$

That curve is the braking limit. As long as the current velocity is below it, full acceleration toward the target is allowed. On the curve, the shaper must brake at full \(a_{max}\) to land exactly on the target with zero velocity.

In discrete time the bound is slightly stricter, because the shaper moves one full tick at the chosen velocity before it can brake. Require \(vT + v^2/(2a_{max}) \le d\) and solve the quadratic for \(v\):

$$v_{brk} = -a_{max}T + \sqrt{a_{max}^2T^2 + 2\,a_{max}\,|d_k|}.$$

The implementation uses this discrete form. For \(d \gg a_{max}T^2\) it converges to the simple square root.

The physical picture

The last two equations form a double integrator: we choose the second derivative of the sent trajectory, bounded by \(a_{max}\), and integrate twice:

$$\ddot{s} = u, \qquad |u| \le a_{max}, \qquad |\dot{s}| \le v_{max}.$$

The physical model is a unit mass pushed by a bounded force, \(|F| \le m\,a_{max}\). In the phase plane \((d, v)\) the braking curve \(v = \mathrm{sign}(d)\sqrt{2 a_{max} |d|}\) is the switching curve of minimum-time (bang-bang) control: push at full force toward the target until the state touches the curve, then ride the curve into the origin. With the velocity saturation on top, the resulting profile is the familiar trapezoid. The shaper is not "just clamping" — it is the time-optimal member of the second-order filter family. It reaches the target as fast as the bounds allow, and never faster.

Blue: the shaped trajectory from the demo above, starting at d = the raw jump, v = 0. Gray curve: the switching curve v = √(2·a_max·d). The trajectory accelerates at full a_max, clips at v_max when the jump is long enough, then rides the braking curve into the origin — the bang-bang minimum-time solution. Change the sliders in §05 and this plot follows.

The one cost of the no-overshoot guarantee: when the target itself moves at velocity \(v\), the shaper trails it by the braking distance,

$$\Delta_{trail} \approx \frac{v^2}{2\,a_{max}} + vT,$$

because it always keeps enough margin to stop if the target stops. At \(a_{max}=500\) °/s² and \(v=120\) °/s that is about 15° (~130 ms). Raising \(a_{max}\) shrinks this quadratically; the safe ceiling between 600 (proven) and ~10,000 °/s² (the trigger) is still open.

Relation to a mass-damper admittance

A mass-spring-damper admittance is the linear second-order filter:

$$M\ddot{s} + D\dot{s} = K\,(x - s).$$

Its acceleration is proportional to the error: \(\ddot{s} = \left(K(x - s) - D\dot{s}\right)/M\). Feed it a 30° jump and the initial acceleration is \(30K/M\) — large, unless the spring is soft. A soft spring bounds acceleration only in practice, never by guarantee, and it makes the filter sluggish for all inputs, all the time. The shaper is the saturated, time-optimal member of the same family:

PropertyMass-damper (linear)Shaper (bang-bang)
Accelerationproportional to error, unboundedhard bound a_max, by construction
Compliant inputalways filtered, always some lagidentity, zero added lag
Large jumplarge force spike, or sluggish tuningtime-optimal trapezoid
Guarantee typetuning outcomemathematical invariant

Both have a place, and they are complementary, not competing. The mass-damper belongs on the leader: it shapes the feel of the human input (Atlas already runs an admittance filter on the SpaceMouse). But nothing on the leader can guarantee what reaches the wire: network stalls, the latest-wins mailbox, go-home, replay, and bugs all enter downstream of it. The shaper sits at the follower's last point before the wire and bounds whatever arrives, from any source. Because it is transparent for compliant input, a well-tuned leader admittance means the shaper normally does nothing at all. It is the constraint, not the controller.

09 · Appendix

Reproduction quick reference

Build (in rust/ of the Atlas repo):

cargo build --release -p fairino-sdk --bin servoj-lag

Lag measurement — safe on any arm; this is how the tuning table was produced. Wrist joint with a small sweep first, then J1 to confirm. Read L from the fit line. Stop the Atlas node first: the CNDE stream allows one client.

servoj-lag <IP> ramp 5 10 5,10 10 keep shaped servoj-lag <IP> ramp 0 15 5,10 10 keep shaped

Runaway trigger — test arms only, never a production arm. One 1.5° step; e-stop in hand, clear arc on the joint, power-cycle after any code-31 stop.

servoj-lag <IP> 0 1.5 1 10 keep

Fix demonstration — the same class of jump through the shaper; expect zero rejects and a ~330 ms smooth arrival.

servoj-lag <IP> 0 10 1 10 keep shaped

Every run writes a timestamped CSV with the full config in its header.