← Index

Humanoid Robustness Ladder

Humanoid learning and evaluation across Isaac Lab and MuJoCo, with sensor ablations and shared-world team tasks.

FIG. 03 — MuJoCo inspection: a frozen gait with oracle routing and sensor-reactive braking, not the new Isaac policy.

Problem

Training reward and a good-looking rollout do not establish reliable robot behavior. I needed repeatable tests that distinguish locomotion, sensor use, task completion, simulator transfer, and failure—and refuse success when a camera or control is broken.

Approach

Train PPO in Isaac Lab, export ONNX, and evaluate compatible frozen policies in MuJoCo. Run seeded curricula and sensor ablations through Slurm, with per-episode contact, fall, camera-health, and mission checks. Separate new Isaac training from older frozen-gait MuJoCo experiments: the inspection mission uses ordered stations and sensor-reactive braking; two- and three-robot teams share a physical airlock. Controls test sensor outage, a wrong branch, skipped coordination, and missing teammates. Oracle routing and localization remain explicit.

What I built

Built the cross-simulator scoring pipeline, task supervisors, simulated sensor adapters, seeded failure controls, Slurm experiment runners, and evidence ledger. Kept upstream robot assets and policy interfaces, while separating learned control from scripted supervision and retaining failed experiments.

Result

New Isaac biped Full checkpoints: 379/384 first episodes pass across four sensor arms, three training seeds, 32 envs each; fixed route/layout, noise off. Separately, older frozen-gait MuJoCo runs pass inspection 3/3 and each team size 5/5. These are small, bounded simulation tests, not transfer of the new checkpoints or hardware validation.

Complete a mission, then challenge its assumptions

Two ordered dwell stations, two travel turns, a dead end, and an exit test a 22-DoF frozen gait. The torso stays east-facing; the northbound leg is a sidestep. Simulated IMU, lidar, and idealized paired depth feed braking—not stereo matching or SSD. The baseline passes 3/3; outage and wrong-branch controls each pass 0/3. Map, pose, and waypoints are privileged.

FIG. 03A — Wrong-route control: dead-end rejection at 7 s. Deliberate negative control, not a spontaneous policy failure.

Shared physics, explicit team coordination

All robots must occupy assigned inspection pads before a physical airlock opens, then cross one at a time. Two- and three-robot teams each pass 5/5 seeds. Both the no-wait and withheld-teammate controls pass 0/5 for each size. This demonstrates coordinated navigation, not cooperative grasping or carrying.

FIG. 03B — Three-robot success rollout; the five-seed experiment and controls supply the score.

Keep the evaluation engines separate

The new 12-DoF Isaac biped result is 379/384 first episodes, not MuJoCo transfer. An older flat-ground transfer study scored 90 episodes per condition: unrandomized policies fell 23% of the time versus 0% for default randomization. Neither result establishes unseen-maze generalization or physical locomotion.

FIG. 03C — Historical terrain-transfer diagnostic on a composed lab floor; separate from the new Isaac runs.

Failure analysis is part of the deliverable

A historical cooperative-carry policy exploited collapse rather than learning a stable lift: the humanoids dropped about 41 cm before cube contact. Cloth folding remains research, with no improvement in the latest strict short-pants comparison. Retractions, renderer failures, and incomplete cells remain in the evidence ledger.

Where the next experiment starts

The inspection route is known, sensor noise is not calibrated to hardware, and paired depth is idealized ray casting. An SSD detection interface is not a trained detector. The next evidence needed is unseen layouts, calibrated sensor noise, and export and evaluation of the new Isaac checkpoints in MuJoCo—not a claim that these are already solved.

The team benchmark specification defines contact and completion gates. The folding investigation separates strict fresh-camera evaluation from historical stale-image results.

Stack

Links