Ethgar · robot lab notebook

Robots that plan live, measured in the open

Drones, a real SO-101 arm, a cable robot that spans a room, Unitree G1 humanoids and a winery digital twin, built solo in 2025–2026 on a laptop and rented GPUs. Far from perfect. The failures are written down next to the wins, with numbers.

The thread through all of it: re-plan every tick with MPPI over a world model, write the task as cost terms anyone can read, switch phases on measured events (never a clock), and draw the robot's plan while it moves. Learned policies (ACT, SmolVLA, PPO) are trained and measured alongside, as the comparison.
11
projects
210+
commits
34
public models & datasets
4/4
real pen grasps, live MPPI
~$1
per G1 policy on HF Jobs

Stringman: a cable robot learns its own body, works in any room

Oct 8, 2026sim, real rooms + 3D roomsworks in sim

Stringman (Neufangled) hangs a gripper from four lines between the corners of a room and tidies the floor. The field teaches robots like it with hours of teleoperation and GPU-trained policies on pixels. Here the SkyJEPA recipe instead: a tiny world model of the machine, trained on simulated motion in random rooms with no task and no demonstration, then live MPPI where each mission is cost terms. One camera at the top of the room says where things are; the bin is a named place.

0
demonstrations or real pixels in training
28 k · 11 min
parameters · training on a laptop CPU (data: 1.5 CPU-h)
33–34 / 36
cubes binned in two real rooms (imported from naavox) it never saw
20–24 / 24
in furnished 3D rooms, 0 furniture contacts
  • Any room: the model never sees the room's position frame, only the gantry's height and each line's direction and length from the 4 calibrated mount points. Trained on a new random room every trajectory. A playroom-only version had binned 5/36 in the bedroom.
  • Ties the physics simulator: 20 mm and 0.8–0.9° after 2 s of open-loop prediction in unseen rooms, as MuJoCo in the same room (which knows the room but not the latency, the miscalibration or the payload).
  • Against the teleop route (naavox, as published): 5,984 teleop episodes (35 h) and 56 models from ACT (52 M params) to pi0.5 (4.1 B) trained on 4090s, H100s and H200s; on the same laptop one decision costs 87 ms (ACT) to 3.8 s (pi0.5, est.), ours 35 ms on the CPU.
  • Real rooms from their recordings: anchors, named places and the floor view imported from naavox/playroom-sep29-2 and bedroom-sep6. 3D rooms: the motors sit in the corners, so a room builds the robot; furniture is only in the mission's costs.
Known limitsSimulation only: the world is a MuJoCo twin whose masses, line stretch and spool acceleration are still guesses. Rigid 4 cm cubes, not socks; the image differencing misses a pale cube on beige carpet. The model's open-loop test fails its own 20 mm bar by 0.2–3 mm. Power not measured; the naavox policies' success rates are not published, so this compares cost, not a head-to-head score.
Furnished 3D playroom, 4× speed: 12/12 cubes around the sofa, table and ball pit. Left: the purple plan. Middle: the one overhead camera. Right: the gripper camera.
Mission map in the furnished playroom
The whole mission from above: one colour per pick, all routed around the furniture.
Real bedroom floor (naavox), never seen in training: 11/12.
Furnished 3D bedroom: 11/12, bed, wardrobe and desk in the way.
Prediction error against horizon, learned vs MuJoCo vs backbone
In the real bedroom it never saw: the 28k-parameter model (purple) on top of MuJoCo (orange).

Two arms hand a pen to each other

Oct 8, 2026sim5/10 round trips

Two SO-101s side by side like a person's two arms, one camera between and behind them like eyes. Arm A picks up the red pen and offers it, arm B takes it, then gives it back. Each arm runs its own live MPPI over its own copy of the learned arm model; a shared "nervous system" carries each hand's servo readings to the other.

5 / 10
full round trips A → B → A in sim (best batch 7/10)
~16 s
pen from the table into B's hand; back in A's at ~22 s
1.3 s
both hands on the pen during the exchange
40–80 ms
per tick for both plans together, on a laptop
  • The giver opens only when the taker's servos say it holds: jaws stalled on the pen, load felt, for 0.4 s. A release is an event on the other hand's sensors, never a clock.
  • The taker aims where the giver's body says the pen is: the giver's joints × where the pen sits in its jaws. The one camera sees the pen as a line and corrects that belief (direction and depth).
  • The pinch feels the pen's thickness: pinched on its curve instead of across its 11 mm diameter, the jaws close further. "Too thin" means re-grasp lower before it slips.
  • Each wrist camera faces away from the other hand: it sticks out 7–10 cm along the pen, and the grip spots are 58 mm apart.
  • The reverse hand-over is the same logic with the roles swapped.
Known failuresA taker that misses can knock the pen out of the giver's gentle grip. The giver opens too fast to recover if the taker slips (a two-step release is next). A hand holding the pen alone 55 mm from its centre lets it pivot out. Sim only: there is one real arm so far, and the simulated servo load saturates, so the "feel" is mostly the jaw angle.
A → B at 16.8 s, B → A at 24.2 s. Left: the one camera. Right: close-up. Purple = A's plan, cyan = B's.
13.0 s13.5 s14.0 s14.5 s15.0 s15.5 s16.0 s16.5 sopen0on the penJAW ANGLE (rad)A holds, offeringB open, approachingA's pads on the penB's pads on the penboth hands on the pen: 1.6 sB closesB: holding, feltA starts openingA clear
One exchange from the sensors: jaw angles, and which hand's pads touch the pen. B reports "holding" at 14.0 s; A starts opening 0.1 s later.
Failure: A's grip slips 0.3 s into B's release; B squeezes again too late.
Failure: B's closing jaw knocks the pen out of A's hand.
pads past the equator-8-4+0+4+8+12+16-0.06-0.08-0.10-0.12-0.14-0.16"feels thin" below −0.085squeezed out at the closeheld through a wrist shakepen centre past the fingertip site (mm) · + = pads higher on the 11 mm barreljaw stall (rad)
Where across the pen the jaws pinch, and what they feel: the stall angle jumps on the curve.

Hedwig: a drone carrying the SO-101

Oct 2026simworks in sim, not real time yet

A quadrotor with the real SO-101 arm (its CAD parts, minus the base: the drone's yaw replaces the pan servo) under its belly. One whole-body MPPI flies four rotors and five arm joints together, so the arm's moving weight is planned for, not fought.

Compute question: the planner was born on a Jetson-equipped drone that saw no images. How much does vision add? A survey mission switches the camera on only while hovering: take off, hover at A and look for the red mailbox, fly blind to a stand-off point, hover and look again, land.

~97%
of each tick is the world model (MuJoCo rollouts, 256×15×20 steps)
0.5 ms
mailbox detector per frame
31 / 185
ticks with vision on (hover-only policy)
16 → 5.5 cm
mailbox error: first look (2.5 m) → second look (1 m)
What it taughtThe bill is the world model, not the camera: ~180 ms per tick against a 50 ms budget on an M2. Gating vision saves <1% with a colour detector; it starts to matter with a neural detector on the GPU. A live compute meter (ms per stage, cores, and watts from tegrastats on a Jetson) now runs in every flight. Jetson Orin numbers: next.
Survey mission. Purple: the hull's plan; orange: the claw's. Inset: the drone's camera, VISION ON only while hovering.
Stacked compute per tick, dominated by rollouts
Live compute log: ms per stage per tick vs the 50 ms budget; purple bands = vision on.

armpilot: live MPPI on a real SO-101

Oct 6–7, 2026real armworks

The arm as a seeker, not a replayed episode: every tick it re-reads the wrist camera, re-estimates the pen, re-plans over a MuJoCo model and executes only the first action. If it closes on nothing, it notices from the gripper angle, re-opens and tries elsewhere.

4/4
real grasps + 5 cm lift (10.1, 10.6, 17.0 with a slip-retry, 12.5 s)
2/2
pen dropped into a cup (42 s, 56 s), 10–11 mm off axis
24 vs 37 s
first real A/B: square vs linear "carry upright" cost
7/10
sim development seeds grasped and lifted
  • Events, never clocks: seen → approach → close → lift; "closed on nothing", "jaw blocked", "slipped" are read from the servo's angle and current.
  • Sense of grip: a gentle hold, slip felt by the gripper servo, squeeze harder only when needed. Same loop grasps a screwdriver.
  • The table measured by touch: the arm probes the desk; the model read it 14.5 mm too low.
  • Talk to it: a sentence → the arm's own skills, in a simulated pen / highlighter / cup world.
Known failuresGrasps that land on the pen's edge and slip at the lift; pens at the edge of the reach take ~20 s; the twin's desk is 15 mm off on one side. Planning runs at 4 Hz on the laptop (oracle model); the learned arm model is the next milestone.
Real wrist camera, first real grasp. Purple: the plan projected on the image.
Gentle grip at home: hold, feel, lift.
Screwdriver: same vortex, its own detector; second close succeeds.
Sim: closes on nothing, notices, retries.
Sim: the pen jumps mid-approach, it re-targets.
Sim: pen into the cup.

barpilot: a two-arm G1 bartender

Oct 7, 2026simpaused

A Unitree G1 upper body on a rail behind a bar: what worked on the SO-101 (live MPPI, named costs, events) taken to two 7-DoF arms, a waist and a neck camera.

0.4 mm
median glass position error from head-camera shapes (72 detections, labels 100%)
3/4
coupes placed upright on the counter, 1.4–4.8 mm

The glass is found by shape (depth → voxel clusters → radius profile matched to a catalogue), never from ground truth. Each fix is logged with its measured effect, including the rejected ones.

Why pausedGrasping was a weld triggered on measured touch: a useful abstraction, but it does not transfer to real hands. Work went back to the real arm, where touch is real.
B1: find the coupe, take it, turn, place it on the counter.
Head-camera view of the bar with labelled glasses and bottles
What the head camera recognises, by shape.

sim2real: a desk twin and imitation learning

Oct 5, 2026sim + reallearned a lot, kept the twin

A measured copy of the real desk (1.20 × 0.70 m), the official SO-101 model with its wrist camera, and the BIC 4 Colours pen. Demonstrations generated in the twin became LeRobot datasets; ACT and SmolVLA were trained on HF and evaluated on 50 sim episodes each.

14%
ACT, plain desk (50 demos)
80%
SmolVLA, plain desk (50 demos)
72%
SmolVLA trained on 300 "office" episodes, tested on the office desk
0/20
the plain-desk SmolVLA on the office desk

On the real arm SmolVLA aligned on the pen but often slipped or landed offset, and nothing in it retries. That is why armpilot exists: explicit closed-loop seeking with costs you can read and edit. The twin lives on as armpilot's world.

SmolVLA (office-300): success.
Plain-desk SmolVLA on the office desk: fails.
ACT (sim-50): a typical failure.
Real wrist camera next to the simulated one
Real wrist camera vs the twin's.

toolbot: a G1 humanoid learning to carry

Sep 30 – Oct 2, 2026sim, GPU on HF Jobsteachers work, students fail

A 29-DoF Unitree G1 in mjlab. Every policy trains on Hugging Face Jobs (L4, ~$1 each) and is evaluated on the Mac with energy measured on every rollout: mechanical work and motor heat.

  • T0 walk / stand: 2.3 cm drift standing 10 s, walks 0.44 m/s, no falls (66 min on an L4, ~$0.90).
  • T1 hold a staff: stands, walks and turns with it, but squeezes: 0.9–1.3 kW of heat from the wrists. Nothing in the reward taught it to relax.
  • T3 carry a real box (from PRISM's motion clips): picks up, carries, sets down; box 2.0 cm from its reference. Again 1.1 kW at the wrists.
  • T5 one generalist for box, ball, barrel, bin: within ~1 cm of the specialists, and carries the barrel its specialist dropped.
  • T6 our video → motion: falls at 10%. The lifted reference slides its feet 53 cm/s and puts the box inside the body 84% of frames: plausible, not physical.
  • T7 student (no reference clip): falls at ~25%, at the pick. It cannot tell when to crouch.
  • Doors: motion capture → door rebuilt from the hand's arc → G1 IK with the scene: 16 of 20 door variants accepted.
  • Maps: a real-apartment scan (ReplicaCAD) in MuJoCo; the walking policy goes 8 m around furniture in 18 s; "pick the box on A, drop it on B" planned into skills.
Compute splitEvery model pass on HF, the laptop only orchestrates: a SAM 2 pass that froze the Mac for 30+ min took 58 s on an L4, ~$0.03 per video. 11 jobs in a day, ~$7.
T3: reference | ours | PRISM's policy.
T5: specialists vs one generalist.
Box from A to B in a scanned apartment.
Walking policy + A* in the apartment.
The door bank.
T1: turning with the staff (and squeezing it).

Factory Gym: a winery as a folder of files

Sep 29, 2026Blender, glTF, USDv0: see the plant

A scale-accurate harvest, press and tank hall, modelled on a working plant (anonymised here). The site is YAML; the 3D scene, glTF and USD exports and the renders are generated from it, so editing the YAML and rebuilding is the whole authoring loop.

Field → tractor and 3 t trailer → 1 m² reception pad and auger → membrane press → 4 receiving tanks + 24 filtered tanks, with gendered hose couplings and pumps. Cameras include an arm's-eye view of the vine row and a humanoid's eye height in the working aisle: the environment future policies have to work in.

3 t → 22.4 hl
mass balance checked on every build (80% juice, 600 kg pomace)
20 s
animated batch: every fill level driven by that balance
Why it mattersYield is arithmetic, not a guess: the validator refuses a scene where a tank or the waste trailer would overflow on top of what it already holds.
Overview of the winery site
Overview: parcel → hall.
Tank aisle at human eye height
The working aisle at humanoid eye height.
Vine row with grapes
An arm's-eye view of the vine row.
Press and auger
Auger → drop pipe → press.

SkyJEPA: a 5.6k-parameter world model that flies

Jul 2026simreference model

A clean-room reimplementation of SkyJEPA (Rao et al., 2026): two causal-TCN encoders and a GRU predictor trained as a JEPA with SIGReg, a physics prober that corrects a differentiable rigid-body integrator, and batched MPPI that flies through the learned model. State only (GPS/IMU), trained purely on domain-randomised sim.

5.6k
parameters in the encoder + predictor stack
20k
trajectories, 500-parameter domain randomisation
11.7 vs 30.3 m
peak error after a 0.3 kg mid-hover payload release: v2 recovers, v1 diverges
+51%
v2's figure-8 error: why it was not promoted
What it taughtDistribution beats scale only half-way: the mission mix helps on disturbances (payload, wind, battery sag) and hurts sustained attitude excitation. The knob is the mixture ratio (~0.4 mission), not more data.
Figure-8 flown by MPPI through the learned model.
Payload release, v1.
Payload release, v2 (recovers).

SkyFall: tasks are only cost functions

Jul – Aug 2026simthe method everything else reuses

A lab built around one rule: the world model is trained once and frozen; a task exists only as a cost function evaluated on its rollouts inside the planner. New behaviours take minutes (edit a cost, rerun, watch the 3D video), never a retrain.

Return to base (~0.15 m final distance on the nominal domain), terrain gates in maps the drone cannot see, ball drops into canyons, craters and between pillars, drop-and-return missions, and a real-time MuJoCo version. The MPPI, task interface and "spaghetti" plan drawing here are what armpilot, barpilot and Hedwig were ported from.

Ball drop over terrain, MuJoCo.
Mission demo.

More experiments

Skyfull: alpine rescue trajectories

Jun 2026

Real terrain from Chamonix, Zermatt, Aosta and the Dolomites encoded as an implicit neural function; 10,000 rescue episodes per tile; a tiny MLP maps a victim's position to a 40-waypoint flight. No pixels, only coordinates.

First real SO-101 datasets

Jun – Aug 2026

Teleoperated LeRobot datasets on the real arm (a lighter, a blue pen, a red rubber) and an ACT policy trained on two cameras: where the arm work started, before world models.

Tools and side quests

LeCut: a local timeline editor for LeRobot datasets: cut pauses and jitter out of teleoperated demos before training. Orchestrated CartPole: five AI agents (environment, algorithm, training, evaluation, docs) built a PPO CartPole in mjlab together, with a post-mortem of what broke in the hand-offs.