A setpoint jump of ~1–1.5° inside one 10 ms tick latches a controller fault — and the controller then executes an unbounded sweep anyway. This report maps the trigger on the real FR3 and FR20, and describes the acceleration-bounded stream shaper and lag-compensated divergence guard that now make it impossible to trigger from the Atlas side.
The Fairino ServoJ runaway is reproduced, measured, and explained. The trigger is a velocity discontinuity in the setpoint stream: a jump of ~1.0–1.5° inside one 10 ms command period, implying a velocity change of 100–150 °/s in a single tick. The controller detects it, latches a fault — and then still executes an unbounded sweep.
The transport does not matter. The runaway is identical over UDP (cmdType=1) and XML-RPC (cmdType=0): a 2.5° step produced an 88.57° sweep on UDP and 88.56° on XML-RPC. The defect lives in the ServoJ execution engine, not in a communication path.
The threshold band is unstable. At a 1.0° step the arm tracked cleanly once and ran away 13.9° on an equal step minutes later. At 1.5° and above the fault fires every time. Safety margins must therefore sit below 1.0° per tick.
High speed is safe when the motion is smooth. With a trapezoid profile (acceleration ≤ 600 °/s²) the arm ran clean at 40–120 °/s. At 120 °/s the setpoint advances 1.2° per tick (inside the step trigger band) and the arm trails the setpoint by 9.3° (six times the trigger size) — nothing bad happened, twice. The trigger is acceleration — not speed, and not the tracking gap.
Very large steps are safe in a different way. A 30° step is refused outright: every command is rejected and the arm stays frozen. The dangerous zone is mid-size steps, measured from ~1.5° to 4.8° (the 5–30° band is untested raw).
Once the runaway starts, no command can stop it. The controller rejects all commands with errcode 14 for ~0.9 s while the arm moves. Only the physical e-stop or the end of the controller's own motion stops the arm.
The real FR3's tracking lag is 76–79 ms, constant from 5 to 120 °/s. The SimMachine shows 37–40 ms. This lag makes Atlas's old command guard trip falsely at 38–48 °/s, because the guard measures lag × speed, not true divergence.
The fix is validated on the real arm. An acceleration-bounded stream shaper on the follower side (state = last sent pose and velocity; per tick it clamps velocity at 130 °/s and velocity change at 500 °/s² × cmdT, with a braking bound so it stops on target without overshoot) makes the runaway impossible to trigger. Raw jumps of 10° (ten repeats) and 15° — 4–10× the runaway threshold — passed through the shaper onto the real FR3: zero rejects, zero controller faults, smooth arrivals in 328 and 379 ms. Unshaped, a 2.5° jump gave an 88.6° sweep at 196 °/s. The concept is shipped in the Atlas node driver.
The runaway is joint-independent, but not joint-uniform. The proximal joints (indices 0–2: base, shoulder, elbow) trigger earliest; the wrist joints have more room, closer to 5.5° per step. When the gap is large on several joints at once, several joints sweep at once.
The real FR20's tracking lag is measured: L = 109–111 ms, constant across joints (J1 and J6), speeds (5–20 °/s), and directions — the same pure-transport-delay model as the FR3 (78 ms) and the SimMachine (39 ms). The old ~130 ms estimate from session logs was close but high.
The Atlas safety system was rebuilt from these numbers and is live on the Rust node path: a velocity/acceleration-bounded stream shaper at the wire plus a lag-compensated divergence guard replace the old command guards, which measured lag as if it were divergence. Validated on the real FR3 at a sustained ~100 °/s measured, with zero trips and zero controller faults. See PR #204.
The FAIRINO controller has three independent interfaces:
RobotEnable, Mode, ServoMoveStart, ResetAllError, plus a synchronous ServoJ call (the SDK's cmdType=0).
· CNDE (TCP 20005) — a binary state frame pushed every 10 ms: joint positions, robot mode, robot state, and the fault codes main_code/sub_code.
· ServoJ stream (UDP 20007) — one joint setpoint per command period (cmdT, 10 ms in all tests); the controller replies per datagram, and a reply with errcode:N (N ≠ 0) is a rejection.
Atlas streams ServoJ setpoints at 100 Hz. A safety guard (command_guard /
MaxJointDeltasGuard) compares every outgoing setpoint with the last measured pose;
if the difference exceeds a per-joint limit (3.0° in our configs), it raises an e-stop.
Production had seen an unexplained "random motion runaway" that this guard caught, with an
unknown trigger. A separate theory said the guard's delta is dominated by controller lag, so it
limits speed instead of detecting divergence. Both questions are now answered: the trigger is
a velocity discontinuity (finding 1), and the old guard indeed measured lag, not divergence
(finding 7).
servoj-lag is a standalone binary in the Atlas repo at
rust/fairino-sdk/src/bin/servoj_lag.rs. It talks straight to the controller and
contains no safety of any kind — deliberately, because the goal was to observe
the runaway. Both the sender and the recorder use one host clock, so there is no clock-sync
problem: the measured lag is send-instant to change-observed-on-CNDE, the exact vantage point
the Atlas guard has.
ServoMoveEnd (close any old servo window), Mode(manual), RobotEnable(1), Mode(automatic), RobotEnable(1), ServoMoveStart, all over XML-RPC, each retried on the controller's transient NOT_READY code (14). A real controller rejects all ServoJ without this walk; the SimMachine does not care.robot_mode, robot_state, main_code, sub_code.RobotEnable(1) returns 0 immediately, but the robot needs ~1.4 s before it executes ServoJ (it answers errcode 101, "robot is not enabled", in that window). Without this gate a test starts against a dead stream, the setpoints run ahead, and the enable window itself creates a runaway trigger — which is exactly how we met the runaway the first time.ServoMoveEnd, then 1.5 s of extra recording (to catch post-stop drift), then disable, unless the keep flag is set.servoj_lag_<mode>_<UTC timestamp>.csv, with three row kinds: send (each setpoint with its send time), state (each CNDE frame with the status words), delta (|setpoint − last measured| at each send tick — exactly the quantity the Atlas guard computes). Header lines record the full run config.
Arguments: joint index (0 = J1, the base), step size in degrees, number of steps, command
period in ms, keep = leave the robot enabled at exit, rpc = send each
setpoint as an XML-RPC call (cmdType=0) instead of UDP. After the warm-up, the streamed
setpoint of the chosen joint changes by the step size in one tick, and the new
value is re-sent every 10 ms for 600 ms; then it steps back. With reps > 1 the level cycle
is +X, 0, −X, 0. There is no ramp between levels — the change is instantaneous by design,
injecting a gap of exactly X degrees at a known instant. Reported per step:
first_move_ms (first CNDE frame that left the pre-step pose),
arrive_ms (first frame within 10% of the target), peak_delta
(max |setpoint − measured|), and guard_lag_ms (how long the guard-view delta
stayed elevated).
Sweeps the joint from its current pose out by sweep_deg degrees and back, once per
listed speed (comma list, °/s). Each pass is a trapezoid: velocity rises linearly over 200 ms,
cruises constant, falls over 200 ms. The velocity change per tick is only v/20, so there is
never a discontinuity. Reported per pass: mean and peak guard delta over the cruise, the
implied lag L = delta / v, a midpoint-crossing lag cross-check, and the overshoot after the
stop. The summary fits delta ≈ v × L; a constant L across speeds means the guard delta is pure
lag.
Against the v3.9.6 Docker SimMachine: step-response lag ~33 ms; sustained-motion L = 39–40 ms, constant from 5 to 40 °/s, identical on the FR3 and FR20 sim models, identical on both transports (37.5 ms via XML-RPC). Delta = v × L exactly. No step size ever caused a runaway in sim — the sim executes any jump almost instantly after a fixed delay. Conclusion: in sim the guard delta is 100% lag, 0% divergence. Also disproved: an old note claimed "75 ms sim lag"; that number came from an earlier probe at cmdT = 20 ms and does not apply at cmdT = 10 ms.
A ramp run started while the robot was still inside its enable window (~1.4 s of errcode 101
after RobotEnable returns 0). The setpoints advanced ~4.8° before the controller
began to execute. When it did, the arm swept ~158° at 180–209 °/s, straight past the +30°
target, and only the operator's e-stop stopped it (errcode 31, "e-stop button", requiring a
control-box power cycle). Reproduction: a run that starts with the robot disabled produces the
runaway (2 out of 2); a run that starts with the robot already enabled is clean at the same
speeds, with the same L. The controlling variable is the gap at the moment execution starts.
This finding produced the acceptance-gated warm-up and the keep flag in the tool.
One step per run, joint 0, robot enabled and tracking before each step:
| Step | Result | Peak speed | Travel |
|---|---|---|---|
| 0.5° | clean, smoothed (arrive 141 ms) | n/a | 0.5° |
| 1.0° | clean once, then RUNAWAY on an equal return step | 75 °/s | 13.9° |
| 1.5° | RUNAWAY | 142 °/s | 33° |
| 2.5° | RUNAWAY | 196 °/s | 88.6° |
| ~4.8° | RUNAWAY (from experiment 2) | 180–209 °/s | 158°+, e-stopped |
The constant pattern in the CSVs: the controller latches fault main_code=1,
sub_code=32 at the step instant, before any motion. Then it rejects
every command with errcode 14 for ~0.9 s while the arm sweeps. The peak speed equals the
step's implied velocity (step / cmdT), clamped at the joint maximum. The travel is roughly
14–35× the step. The motion stops on its own.
The tool's rpc flag sends every setpoint as a synchronous XML-RPC
ServoJ call; each call returns its own accept/reject code, so the tally is exact.
At 1.5° the fault fired but the motion stayed bounded on target (the threshold band is a coin
flip). At 2.5°: runaway, travel 88.56° — identical to the UDP run's 88.57°. Lag via RPC equals
lag via UDP. cmdType does not matter; the runaway is transport-agnostic.
A missing ramp keyword ran step mode with a 30° step. The controller refused. The
teach pendant showed the fault in plain text: "the command speed in the joint space of
axis 1 exceeds the speed limit and can be reset." The controller's speed check works at
giant steps and fails dangerously at mid-size steps.
Trapezoid ramps, robot enabled at start, one speed per run, sweep 30° (60° for the fast runs so the cruise stays long enough to measure):
| Speed | Per-tick delta | Tracking gap (v×L) | Result |
|---|---|---|---|
| 40 °/s | 0.4° | 3.1° | clean, L = 77–79 ms |
| 60 °/s | 0.6° | 4.8° | clean |
| 80 °/s | 0.8° | 6.2° | clean (see note) |
| 100 °/s | 1.0° | 7.7° | clean |
| 120 °/s | 1.2° | 9.3° | clean, twice, overshoot ≤ 0.02° |
Note: one 80 °/s run was stopped by the controller's own Cartesian TCP speed safety, set conservatively for the cell. That is a separate, configurable protection layer; after adjustment the run was clean.
The tool gained a shaped flag. With it, every frame passes through a stream shaper
before the wire. The shaper works per joint, per tick:
Unit tests assert the bounds on every tick of a 2000-tick random shock stream (jumps to ±30°, random freezes). Live results on the real FR3, joint 0, raw jumps injected through the shaper:
| Raw jump | Runs | Rejects | Controller faults | Arrival |
|---|---|---|---|---|
| 10° | 10 | 0 | none | mean 328 ms (322–335) |
| 15° | 1 | 0 | none | 379 ms |
The same class of input, unshaped, at only 2.5°, produced an 88.6° sweep at 196 °/s. The shaped runs also probe the previously untested 5–30° band safely, because the wire never sees the raw jump. Two numbers from these runs matter for the next design step:
After the new system shipped, the same lag measurement ran on the real FR20 (from the r1lite
mini PC). Safety rules for a production arm: ramp mode only (a ramp cannot
create the velocity discontinuity that triggers the runaway; its per-tick velocity change is
v/20), the shaped flag as a second bound on the tool's own output, small sweeps
(10–15°), low speeds (5–20 °/s — enough, because L is speed-independent), the wrist joint J6
first (smallest motion envelope), then J1 to confirm.
Result: L = 109–111 ms on every pass (J6 and J1, 5/10/20 °/s, both directions), zero rejects, zero overshoot. The midpoint cross-check gave the same 103–117 ms. The FR20 controller also accepted UDP ServoJ directly, so its firmware needs no upgrade.
This is the exact algorithm now sitting at the wire, running in your browser at the real tick rate (cmdT = 10 ms). Inject a raw setpoint jump and watch what the controller receives: the raw stream steps instantly — the discontinuity that latches the fault — while the shaped stream turns the same input into a bounded trapezoid. Drag the sliders or pick a preset from the measured campaign.
| tick | t (ms) | raw (°) | shaped (°) | vel (°/s) | Δv/tick (°/s) |
|---|
Simulated ideal wire trajectory from the shipped equations (§08). The physical arm follows the shaped stream ~78 ms behind on the FR3 (~111 ms on the FR20), so measured arrivals — 328 ms mean for the live 10° runs — sit slightly above the wire arrival shown here. Hover or focus the chart and use ←/→ to read exact values; every value is also in the table view.
The number every guard decision rests on comes from the tool's ramp mode:
delta = |setpoint just sent − latest measured joint|.This is what changed in the Atlas safety envelope, and why — the productized form of everything above, shipped in PR #204.
Three mechanisms are deleted from the Rust node path (code, configs, and the safety policy):
command_guard (MaxJointDeltasGuard) tripped when |raw setpoint − measured| exceeded ~3°.command_speed tripped on the same quantity divided by cmdT — the same signal in different units.cap_delta clamped the raw setpoint to measured ± ~3°.cap_delta was worse than useless — after a network stall it emitted a ~3°
instantaneous step, which is inside the measured runaway trigger band. The
protections themselves carried the trigger.
The old wire names (command_guard, command_speed) still exist as
graceful no-ops: an old frontend toggle or runtime edit gets a "retired" warning instead of a
crash.
Component 1 — the stream shaper (never trips, always shapes). It sits at the single point every producer passes before ServoJ: teleop, playout, hold frames, go-home, replay. Its state is the last sent setpoint and its velocity; it never reads the measured pose, so lag cannot blind it. Per tick it bounds the wire three ways: velocity ≤ v_max, velocity change ≤ a_max × cmdT, and a braking bound so it stops on the target without overshoot (full math in §08). Any input, however broken, becomes a physically legal trapezoid. It primes on the measured pose whenever a stream engages — which closes the enable-window trigger — and its state is dropped on stream exit, so a re-engage never ramps out of a stale pre-stop target. A raw input jump beyond the maximum braking trail (~14°) logs a warning; it never trips.
Component 2 — the divergence guard (never shapes, only trips). The closed-loop check the shaper cannot provide, because the shaper cannot see the arm. It trips when
$$|s(t) - m(t)| \;>\; \text{margin} \;+\; |v_{\text{shaper}}| \cdot \text{lag}_s$$
with s the shaped setpoint just sent, m the measured pose, and v the shaper's own commanded velocity — already in its state, so the lag term costs one multiplication; no ring buffer, no time shift. The allowance breathes with the motion: at rest it is the bare margin (~3°, as tight as the old guard); at full speed it grows by exactly the healthy lag term. Clean tracking can never trip it; an arm that moves while we command stillness trips within the margin; a fast controller runaway (140–200 °/s) crosses the allowance within one or two ticks (~20 ms detection).
Two floors protect the tuning against reintroducing false trips:
lag_s must be at least the measured L, rounded up. Below it, the allowance grows slower than the healthy gap and clean cruise trips the guard.margin_deg must be at least a_max × L² / 2 plus noise. That term is the braking transient: during a hard stop the arm briefly lags by more than v × L, and at the stop instant the allowance has collapsed to the bare margin.
Kept unchanged, as the independent nets: the joint-limits gate and the
workarea gate on the command side (they validate the raw target and trip), and the
measured-state observers joint_speed (trip above the wire ceiling: nothing we can
command exceeds v_max, so a faster measured arm means the controller acts on its own) and
eef_speed. These watch the arm, not the commands, so they hold
even if everything above them is misconfigured.
Fail-closed pairing: a config with a shaper but no divergence guard refuses to
construct, and the safety policy (config/safety_policy.yaml) requires
joint_speed, stream_shaper, and divergence_guard keys
for every strict robot type. A Fairino cannot boot half-protected.
The mechanics of both components run on the driver's write path, but the SafetyEnvelope
remains the single reporting surface. It carries their configuration echo and the divergence
trip latch: any_tripped includes the divergence guard, the telemetry safety dict
gained the wire keys stream_shaper (bounds) and divergence (margin,
lag, tripped flag, last reason), and the operator reset clears the latch. Because the bridge,
the frontend, and the test harness all identify tripped guards by scanning that dict, a
divergence e-stop is always named: the "e-stop with an empty guard list"
pattern is gone.
| Platform | Measured L | Wire v_max | Wire a_max | lag_s | margin | joint_speed net |
|---|---|---|---|---|---|---|
| FR3 real | 78 ms | 110 °/s | 500 °/s² | 0.09 | 3.0 | 120 |
| FR20 real | 111 ms | 100 °/s | 450 °/s² | 0.13 | 3.5 | 110 |
| SimMachine | 39 ms | as per family | as per family | 0.05 | 3.0 | 110–120 |
The leader (SpaceMouse) caps were raised to match the follower ceilings; they are comfort shaping only — the follower does not trust them.
harness_shaper_check_rust.py with the inverted contract — a 30° raw punch into the node must be absorbed (speed-bounded wire, clean arrival, zero trips). The divergence trip side is covered by unit tests plus a driver-level test where the simulated arm "moves on its own" and the trip is asserted by name.This section states what the shaper is, in equations, and how it relates to classical control concepts.
Per joint. The state is the last sent position \(s_k\) and its velocity \(v_k\). The constants are the velocity limit \(v_{max}\), the acceleration limit \(a_{max}\), and the tick time \(T\) (cmdT, 10 ms). The input is the raw target \(x_k\), which can jump in any way. Each tick:
$$d_k = x_k - s_k \qquad \text{(position error)}$$
$$v_{brk} = \sqrt{2\, a_{max}\, |d_k|} \qquad \text{(braking curve)}$$
$$v_{des} = \mathrm{sign}(d_k)\cdot \min\!\left(v_{max},\; v_{brk}\right)$$
$$v_{k+1} = \mathrm{clamp}\!\left(v_{des},\; v_k - a_{max}T,\; v_k + a_{max}T\right)$$
$$s_{k+1} = s_k + v_{k+1}\, T$$
The guarantees hold for any input sequence: \(|v_k| \le v_{max}\), \(|v_{k+1} - v_k| \le a_{max} T\) (bounded acceleration), no overshoot of a static target, and the identity function when the input already respects the bounds.
Uniform-acceleration kinematics: a mass that brakes at constant \(a\) from velocity \(v\) needs the distance \(d = v^2 / (2a)\) to stop. Turned around, the highest velocity from which you can still stop within the remaining distance \(d\) is
$$v = \sqrt{2\, a_{max}\, d}.$$
That curve is the braking limit. As long as the current velocity is below it, full acceleration toward the target is allowed. On the curve, the shaper must brake at full \(a_{max}\) to land exactly on the target with zero velocity.
In discrete time the bound is slightly stricter, because the shaper moves one full tick at the chosen velocity before it can brake. Require \(vT + v^2/(2a_{max}) \le d\) and solve the quadratic for \(v\):
$$v_{brk} = -a_{max}T + \sqrt{a_{max}^2T^2 + 2\,a_{max}\,|d_k|}.$$
The implementation uses this discrete form. For \(d \gg a_{max}T^2\) it converges to the simple square root.
The last two equations form a double integrator: we choose the second derivative of the sent trajectory, bounded by \(a_{max}\), and integrate twice:
$$\ddot{s} = u, \qquad |u| \le a_{max}, \qquad |\dot{s}| \le v_{max}.$$
The physical model is a unit mass pushed by a bounded force, \(|F| \le m\,a_{max}\). In the phase plane \((d, v)\) the braking curve \(v = \mathrm{sign}(d)\sqrt{2 a_{max} |d|}\) is the switching curve of minimum-time (bang-bang) control: push at full force toward the target until the state touches the curve, then ride the curve into the origin. With the velocity saturation on top, the resulting profile is the familiar trapezoid. The shaper is not "just clamping" — it is the time-optimal member of the second-order filter family. It reaches the target as fast as the bounds allow, and never faster.
Blue: the shaped trajectory from the demo above, starting at d = the raw jump, v = 0. Gray curve: the switching curve v = √(2·a_max·d). The trajectory accelerates at full a_max, clips at v_max when the jump is long enough, then rides the braking curve into the origin — the bang-bang minimum-time solution. Change the sliders in §05 and this plot follows.
The one cost of the no-overshoot guarantee: when the target itself moves at velocity \(v\), the shaper trails it by the braking distance,
$$\Delta_{trail} \approx \frac{v^2}{2\,a_{max}} + vT,$$
because it always keeps enough margin to stop if the target stops. At \(a_{max}=500\) °/s² and \(v=120\) °/s that is about 15° (~130 ms). Raising \(a_{max}\) shrinks this quadratically; the safe ceiling between 600 (proven) and ~10,000 °/s² (the trigger) is still open.
A mass-spring-damper admittance is the linear second-order filter:
$$M\ddot{s} + D\dot{s} = K\,(x - s).$$
Its acceleration is proportional to the error: \(\ddot{s} = \left(K(x - s) - D\dot{s}\right)/M\). Feed it a 30° jump and the initial acceleration is \(30K/M\) — large, unless the spring is soft. A soft spring bounds acceleration only in practice, never by guarantee, and it makes the filter sluggish for all inputs, all the time. The shaper is the saturated, time-optimal member of the same family:
| Property | Mass-damper (linear) | Shaper (bang-bang) |
|---|---|---|
| Acceleration | proportional to error, unbounded | hard bound a_max, by construction |
| Compliant input | always filtered, always some lag | identity, zero added lag |
| Large jump | large force spike, or sluggish tuning | time-optimal trapezoid |
| Guarantee type | tuning outcome | mathematical invariant |
Both have a place, and they are complementary, not competing. The mass-damper belongs on the leader: it shapes the feel of the human input (Atlas already runs an admittance filter on the SpaceMouse). But nothing on the leader can guarantee what reaches the wire: network stalls, the latest-wins mailbox, go-home, replay, and bugs all enter downstream of it. The shaper sits at the follower's last point before the wire and bounds whatever arrives, from any source. Because it is transparent for compliant input, a well-tuned leader admittance means the shaper normally does nothing at all. It is the constraint, not the controller.
Build (in rust/ of the Atlas repo):
Lag measurement — safe on any arm; this is how the tuning table was produced.
Wrist joint with a small sweep first, then J1 to confirm. Read L from the fit
line. Stop the Atlas node first: the CNDE stream allows one client.
Runaway trigger — test arms only, never a production arm. One 1.5° step; e-stop in hand, clear arc on the joint, power-cycle after any code-31 stop.
Fix demonstration — the same class of jump through the shaper; expect zero rejects and a ~330 ms smooth arrival.
Every run writes a timestamped CSV with the full config in its header.