Abstract
Give a coding agent a command line to a robot and its camera images, and it becomes a physical coding agent: it works on the robot the way it works on a code repository.
We ask what such an agent can do if it may also extend this interface with tools it writes, keeps and fixes from its own failures, and where this stops helping. RoboShell is a harness built to measure this: one program on the command line, three camera views with depth and a manual of one page for a simulated robot with two arms, with no planning or perception of its own. After a failed episode, a second agent running the same model reads the true object poses, finds the cause and writes or fixes a tool; the tool may use only what the robot observes.
With Codex and GPT-6 on all 42 RoboDojo tasks, RoboShell reaches a success rate of 41.5% and a progress score of 48.1%, against 32.3% and 39.3% for VPP2 and 31.4% and 36.3% for PhysicalRSI, the two strongest leaderboard entries, and 22.5% and 29.0% for GPT-6 alone.
Introduction
A coding agent can now be pointed at a repository and left alone: it reads the code, writes a program, runs it, reads the error and tries again until a test passes. Nothing in that loop is specific to software. Give the same agent a command line to a robot and its camera images: now the program moves an arm, the test is a benchmark’s success checker, and the error is a gripper that closed on air. We call this setting physical coding and study it on a full manipulation benchmark.
Robot learning has mostly taken two other routes. Learned visuomotor policies [4, 5, 24] replace the agent with a network trained on demonstrations; language-model robot systems [2, 20, 26] keep the model but surround it with perception modules, skill libraries and planners written by hand.
RoboShell inverts this division of labour: the agent writes the primitives, they persist across episodes and are what is optimized, and the per-episode program is short.
First, evolved tools lift the same model by 19 points and past the strongest learned systems.
Second, tools move the boundary where geometry is enough and little elsewhere.
Third, failures have signatures, and some are missing information rather than missing capability.
The 19-point comparison uses the current-tools aggregate. The evaluation protocols are distinguished below.
RoboShell
RoboShell has two parts (Fig. 3): a harness that turns commands into motion (Sec. 3.1) and an evolution loop that turns failures into tools (Sec. 3.3). Four principles shaped both. The harness translates and never decides. Ground truth is isolated by architecture: the container the acting agent runs in holds no simulator, task or scoring code. Success is judged only by the benchmark’s own checker, which we do not modify. The interface description states interfaces only; strategy lives in tools.
The harness
The acting agent is a stock coding agent CLI running in a container with no direct network access, no API key and no host directories. It is given one executable, robo, a directory obs/ that is refreshed on request with three 640×480 RGB-D views (head and two wrist cameras, depth in metres, OpenCV intrinsics and extrinsics), the two tool centre point (TCP) poses, and an interface description of about sixty lines.
Tools
A tool is a directory with tool.py and interface.md. The Python file declares a name, its sub-commands and their arguments, and a run(api, command, args) function over an EpisodeAPI that offers joint angles, TCP poses, the three RGB-D images with calibration, and the same motion primitives the base commands use. Tools cannot import the simulator, cannot read object poses, and are loaded only if listed in the task’s enabled_tools.txt.
The evolution loop
One command starts an unattended run for one task on one GPU. The loop uses the first ten layouts of evaluation seed 0 (for Generalization tasks, five standard and five random, which share tools) and proceeds layout by layout. The acting agent attempts the layout with the current tool pool. On success, the optimizer agent reads the trajectory and appends a short entry to playbook.md on which tools worked and why.
On failure, it reads the true object trajectory, the command log, the per-command camera frames and the current tool code, writes a diagnosis to skill.md, and adds or repairs a tool. A delivery check (check-task) then verifies that files are complete, tools load, and the interface text contains no task or object words that would leak strategy into the prompt; if the check fails the edit is reverted.
After the tenth layout the optimizer consolidates the three documents, the final tool set is frozen, and every layout is retested once with no optimizer; the retest count is the only number we report as a result.
Experiments
Setup
Both agents are the unmodified Codex command-line agent with GPT-6 at medium reasoning effort; each task gets one evolution on one RTX 4090 with a 12-hour development cap.
Round 1 is the fully automatic loop on all 42 tasks, scored by a frozen retest of one episode on each of the ten development layouts (the first ten of seed 0; five standard and five random for Generalization tasks).
Official runs frozen tools through the benchmark’s own evaluation client on all three seeds with the native episode counts.
Round 2 repeats the loop on the tasks where PhysicalRSI was clearly ahead, adding a human-written note on what the checker requires (Sec. 5.3).
Main results
Fully automatic, round 1 reaches a macro-average of 37.6% over the 54 variants, against 26.1% for PhysicalRSI and 24.3% for the same model zero-shot (38.0%, 31.4% and 22.5% on the leaderboard’s metric).
With the current tools the average is 40.9% against 26.1% for PhysicalRSI and 30.6% for VPP2 (41.5% against 31.4% and 32.3% on the leaderboard’s metric).
Success rate (%)
0–100 scale
- RoboShell
- 41.5
- VPP2
- 32.3
- PhysicalRSI
- 31.4
- GPT-6 Astra, zero-shot
- 22.5
Progress score
0–100 scale
- RoboShell
- 48.1
- VPP2
- 39.3
- PhysicalRSI
- 36.3
- GPT-6 Astra, zero-shot
- 29.0
Overall values from Figure 2. RoboShell combines official-client results for 27 variants with development-layout retests for the remainder; current tools include second-round interventions. Baselines use their published protocols.
The zero-shot baseline uses a different harness, so we also compare within ours to check that the gain comes from the tools and not from the harness. The first development episode of each run had no task tools, only the shared pick tool.
On the 37 tasks whose log records it, the agent passed 4 of these layouts; the frozen retest of the same layouts, with the same harness and model, passed 16, turning 12 failures into successes and none the other way (exact McNemar test, p < 0.001).
Under the official evaluation client
This section covers the round-1 tools of sixteen tasks, the fifteen with a round-1 retest of at least 7/10 and fold clothes; six second-round tasks follow in Sec. 5.3, for 3,286 episodes over 27 variants in all (Tab. 3 and Fig. 5, supplementary).
These are runs on our machines, not leaderboard entries.
With round-1 tools RoboShell reaches an unweighted macro-average of 73.2% over the 20 variants (2,391 episodes), against 44.7% for PhysicalRSI and 48.3% for the zero-shot model on the same variants, and is higher than PhysicalRSI on 13, tied on one and lower on six.
The official runs contain the development layouts, the first ten of seed 0, and on them the official client passes 76.2% (macro-average over variants), against 73.0% on all other layouts; most of the gap is run-to-run variation of the same tools on the same scenes.
Tools that compute everything from the observed geometry transfer almost unchanged (cover blocks 150/150, stack blocks 147/150 over both variants), while play tic tac toe falls from 9/10 to 33% and insert tubes from 8/10 to 52%.
Clean versus cluttered layouts
In round 1 RoboShell passes 25 of 60 standard and 20 of 60 cluttered layouts, keeping four fifths of its success; PhysicalRSI keeps about one fifth (25.7% to 5.6%).
Much of this robustness belongs to the model, which keeps 87% of its clean-layout success zero-shot (32.7% to 28.3%). With the current tools RoboShell keeps 79% at a higher level (46.1% to 36.3%), whereas the leaderboard’s 21 VLA and other learned-policy entries keep 1–68% (median 22%) and its five WAM entries 5–25%.
What the agents built
Across the 42 tasks the round-1 optimizer wrote 96 perception commands and 141 execution commands; the shared pool still contains only the pick tool.
The library is highly redundant: a surface measurement from a seed pixel was reinvented in nineteen tasks and a transfer macro in twenty, with interfaces that differ in argument names more than in meaning.
Model picks the pixel, code measures the world; this split keeps the tools deployable and auditable, and, as Sec. 5 shows, it is also where their limits come from.
Where the boundary lies
A capability map
The physical interaction decides. Rigid manipulation reaches 57.7% and pushing 75.0%; tight insertion reaches 21.7%, deformable and fluid material 20.0%, and articulated objects and tool use 0%.
A tool fixes one interpretation and one open-loop motion; an agent acting step by step can still notice and correct. Self-evolution trades that adaptivity for repeatability, which pays where the motion can be planned from geometry and costs where it must be found during the motion.
How episodes fail
Of the 420 retest episodes, 262 failed: 56% exhausted the budget, 32% ended with the agent declaring completion and the checker disagreeing, 11% ended with the agent giving up while a quarter or more of the budget remained, and two were terminated by the checker for a rule violation.
Missing information or missing capability?
For the eight tasks on which PhysicalRSI was clearly ahead after round 1 we reran the loop with one addition: a person reads the public task definition and writes a short note on what the success checker requires and where the unattended run fell short, given to the optimizer at every round.
Where the note supplies information the instruction lacks, three tasks go from 0–40% to 95–100%.
These are development-layout results; the paper reports official-client follow-up results separately.
Where the gap is capability, fold clothes does not move: the optimizer rebuilt the folding tool from the checker’s keypoint windows and fold order, but retests passed 2/10 and 4/10 against 3/10 before; a note cannot supply a closed-loop skill.
build tower, rebuilt by hand by a team member, went from 0/10 to 5/10 and reaches 56.0% under the official client (PhysicalRSI: 28.7%).
Time and cost
A retest episode takes a median 520 s of wall-clock time for 22.6 s of robot time, almost all model latency, which RoboDojo does not charge.
Limitations
Round-1 numbers come from ten seed-0 layouts per task and overstate official results (81.5% against 73.2%). The optimizer reads true poses, so the loop cannot yet run on a real robot without motion capture.
Conclusion
A stock coding agent behind a harness that only translates commands into motion, improved by a second copy of itself writing tools from its failures, rises from 22.5% to 38.0% on 42 tasks, past the strongest learned entries.
Where the right motion follows from measured geometry, tools hold across layouts and clutter; where success must be felt or watched, a frozen macro is no substitute for closed-loop control.
Next, the same loop will run in RoboCasa [33] and RoboTwin [9], then on real robots, and its traces will train a coding model of our own.
Abridged from the paper. Section, figure, and reference numbers within excerpts refer to the full PDF.