Stringman: a cable robot learns its own body, works in any room
Oct 8, 2026sim, real rooms + 3D roomsworks in simStringman (Neufangled) hangs a gripper from four lines between the corners of a room and tidies the floor. The field teaches robots like it with hours of teleoperation and GPU-trained policies on pixels. Here the SkyJEPA recipe instead: a tiny world model of the machine, trained on simulated motion in random rooms with no task and no demonstration, then live MPPI where each mission is cost terms. One camera at the top of the room says where things are; the bin is a named place.
- Any room: the model never sees the room's position frame, only the gantry's height and each line's direction and length from the 4 calibrated mount points. Trained on a new random room every trajectory. A playroom-only version had binned 5/36 in the bedroom.
- Ties the physics simulator: 20 mm and 0.8–0.9° after 2 s of open-loop prediction in unseen rooms, as MuJoCo in the same room (which knows the room but not the latency, the miscalibration or the payload).
- Against the teleop route (naavox, as published): 5,984 teleop episodes (35 h) and 56 models from ACT (52 M params) to pi0.5 (4.1 B) trained on 4090s, H100s and H200s; on the same laptop one decision costs 87 ms (ACT) to 3.8 s (pi0.5, est.), ours 35 ms on the CPU.
- Real rooms from their recordings: anchors, named places and the floor view imported from naavox/playroom-sep29-2 and bedroom-sep6. 3D rooms: the motors sit in the corners, so a room builds the robot; furniture is only in the mission's costs.


Two arms hand a pen to each other
Oct 8, 2026sim5/10 round tripsTwo SO-101s side by side like a person's two arms, one camera between and behind them like eyes. Arm A picks up the red pen and offers it, arm B takes it, then gives it back. Each arm runs its own live MPPI over its own copy of the learned arm model; a shared "nervous system" carries each hand's servo readings to the other.
- The giver opens only when the taker's servos say it holds: jaws stalled on the pen, load felt, for 0.4 s. A release is an event on the other hand's sensors, never a clock.
- The taker aims where the giver's body says the pen is: the giver's joints × where the pen sits in its jaws. The one camera sees the pen as a line and corrects that belief (direction and depth).
- The pinch feels the pen's thickness: pinched on its curve instead of across its 11 mm diameter, the jaws close further. "Too thin" means re-grasp lower before it slips.
- Each wrist camera faces away from the other hand: it sticks out 7–10 cm along the pen, and the grip spots are 58 mm apart.
- The reverse hand-over is the same logic with the roles swapped.
Hedwig: a drone carrying the SO-101
Oct 2026simworks in sim, not real time yetA quadrotor with the real SO-101 arm (its CAD parts, minus the base: the drone's yaw replaces the pan servo) under its belly. One whole-body MPPI flies four rotors and five arm joints together, so the arm's moving weight is planned for, not fought.
Compute question: the planner was born on a Jetson-equipped drone that saw no images. How much does vision add? A survey mission switches the camera on only while hovering: take off, hover at A and look for the red mailbox, fly blind to a stand-off point, hover and look again, land.
tegrastats on a Jetson) now runs in every flight. Jetson Orin numbers: next.
armpilot: live MPPI on a real SO-101
Oct 6–7, 2026real armworksThe arm as a seeker, not a replayed episode: every tick it re-reads the wrist camera, re-estimates the pen, re-plans over a MuJoCo model and executes only the first action. If it closes on nothing, it notices from the gripper angle, re-opens and tries elsewhere.
- Events, never clocks: seen → approach → close → lift; "closed on nothing", "jaw blocked", "slipped" are read from the servo's angle and current.
- Sense of grip: a gentle hold, slip felt by the gripper servo, squeeze harder only when needed. Same loop grasps a screwdriver.
- The table measured by touch: the arm probes the desk; the model read it 14.5 mm too low.
- Talk to it: a sentence → the arm's own skills, in a simulated pen / highlighter / cup world.
barpilot: a two-arm G1 bartender
Oct 7, 2026simpausedA Unitree G1 upper body on a rail behind a bar: what worked on the SO-101 (live MPPI, named costs, events) taken to two 7-DoF arms, a waist and a neck camera.
The glass is found by shape (depth → voxel clusters → radius profile matched to a catalogue), never from ground truth. Each fix is logged with its measured effect, including the rejected ones.

sim2real: a desk twin and imitation learning
Oct 5, 2026sim + reallearned a lot, kept the twinA measured copy of the real desk (1.20 × 0.70 m), the official SO-101 model with its wrist camera, and the BIC 4 Colours pen. Demonstrations generated in the twin became LeRobot datasets; ACT and SmolVLA were trained on HF and evaluated on 50 sim episodes each.
On the real arm SmolVLA aligned on the pen but often slipped or landed offset, and nothing in it retries. That is why armpilot exists: explicit closed-loop seeking with costs you can read and edit. The twin lives on as armpilot's world.

toolbot: a G1 humanoid learning to carry
Sep 30 – Oct 2, 2026sim, GPU on HF Jobsteachers work, students failA 29-DoF Unitree G1 in mjlab. Every policy trains on Hugging Face Jobs (L4, ~$1 each) and is evaluated on the Mac with energy measured on every rollout: mechanical work and motor heat.
- T0 walk / stand: 2.3 cm drift standing 10 s, walks 0.44 m/s, no falls (66 min on an L4, ~$0.90).
- T1 hold a staff: stands, walks and turns with it, but squeezes: 0.9–1.3 kW of heat from the wrists. Nothing in the reward taught it to relax.
- T3 carry a real box (from PRISM's motion clips): picks up, carries, sets down; box 2.0 cm from its reference. Again 1.1 kW at the wrists.
- T5 one generalist for box, ball, barrel, bin: within ~1 cm of the specialists, and carries the barrel its specialist dropped.
- T6 our video → motion: falls at 10%. The lifted reference slides its feet 53 cm/s and puts the box inside the body 84% of frames: plausible, not physical.
- T7 student (no reference clip): falls at ~25%, at the pick. It cannot tell when to crouch.
- Doors: motion capture → door rebuilt from the hand's arc → G1 IK with the scene: 16 of 20 door variants accepted.
- Maps: a real-apartment scan (ReplicaCAD) in MuJoCo; the walking policy goes 8 m around furniture in 18 s; "pick the box on A, drop it on B" planned into skills.
Factory Gym: a winery as a folder of files
Sep 29, 2026Blender, glTF, USDv0: see the plantA scale-accurate harvest, press and tank hall, modelled on a working plant (anonymised here). The site is YAML; the 3D scene, glTF and USD exports and the renders are generated from it, so editing the YAML and rebuilding is the whole authoring loop.
Field → tractor and 3 t trailer → 1 m² reception pad and auger → membrane press → 4 receiving tanks + 24 filtered tanks, with gendered hose couplings and pumps. Cameras include an arm's-eye view of the vine row and a humanoid's eye height in the working aisle: the environment future policies have to work in.




SkyJEPA: a 5.6k-parameter world model that flies
Jul 2026simreference modelA clean-room reimplementation of SkyJEPA (Rao et al., 2026): two causal-TCN encoders and a GRU predictor trained as a JEPA with SIGReg, a physics prober that corrects a differentiable rigid-body integrator, and batched MPPI that flies through the learned model. State only (GPS/IMU), trained purely on domain-randomised sim.
SkyFall: tasks are only cost functions
Jul – Aug 2026simthe method everything else reusesA lab built around one rule: the world model is trained once and frozen; a task exists only as a cost function evaluated on its rollouts inside the planner. New behaviours take minutes (edit a cost, rerun, watch the 3D video), never a retrain.
Return to base (~0.15 m final distance on the nominal domain), terrain gates in maps the drone cannot see, ball drops into canyons, craters and between pillars, drop-and-return missions, and a real-time MuJoCo version. The MPPI, task interface and "spaghetti" plan drawing here are what armpilot, barpilot and Hedwig were ported from.
More experiments
Skyfull: alpine rescue trajectories
Jun 2026Real terrain from Chamonix, Zermatt, Aosta and the Dolomites encoded as an implicit neural function; 10,000 rescue episodes per tile; a tiny MLP maps a victim's position to a 40-waypoint flight. No pixels, only coordinates.
First real SO-101 datasets
Jun – Aug 2026Teleoperated LeRobot datasets on the real arm (a lighter, a blue pen, a red rubber) and an ACT policy trained on two cameras: where the arm work started, before world models.
Tools and side quests
LeCut: a local timeline editor for LeRobot datasets: cut pauses and jitter out of teleoperated demos before training. Orchestrated CartPole: five AI agents (environment, algorithm, training, evaluation, docs) built a PPO CartPole in mjlab together, with a post-mortem of what broke in the hand-offs.