All posts

Research

RoboShell: A Thin Harness for Measuring What Coding Agents Can Do with a Robot

Higher scores than RoboDojo’s leading model

RoboShell · Paper-reported results · 42 tasks · 54 variants

Success rate
41.5%
+9.2 pp vs. VPP2’s 32.3%
Progress score
48.1%
+8.8 pp vs. VPP2’s 39.3%

RoboShell’s reported results exceed VPP2, the highest-scoring model on the published RoboDojo leaderboard, on both metrics.

Leaderboard checked 11 Oct 2026 ↗ · Values rounded to one decimal.

On this page

Abstract

Give a coding agent a command line to a robot and its camera images, and it becomes a physical coding agent: it works on the robot the way it works on a code repository.

We ask what such an agent can do if it may also extend this interface with tools it writes, keeps and fixes from its own failures, and where this stops helping. RoboShell is a harness built to measure this: one program on the command line, three camera views with depth and a manual of one page for a simulated robot with two arms, with no planning or perception of its own. After a failed episode, a second agent running the same model reads the true object poses, finds the cause and writes or fixes a tool; the tool may use only what the robot observes.

With Codex and GPT-6 on all 42 RoboDojo tasks, RoboShell reaches a success rate of 41.5% and a progress score of 48.1%, against 32.3% and 39.3% for VPP2 and 31.4% and 36.3% for PhysicalRSI, the two strongest leaderboard entries, and 22.5% and 29.0% for GPT-6 alone.

The RoboShell acting agent, thin harness, robot rollout, and failure-driven tool optimizer; examples of rigid, planar, insertion, cluttered, and cloth tasks.
Figure 1.RoboShell turns robot failures into reusable tools.A coding agent operates a robot with two arms through a thin harness. After a failure, an optimizer agent diagnoses the episode with the true object poses and writes or fixes tools that use only what the robot observes.

Introduction

A coding agent can now be pointed at a repository and left alone: it reads the code, writes a program, runs it, reads the error and tries again until a test passes. Nothing in that loop is specific to software. Give the same agent a command line to a robot and its camera images: now the program moves an arm, the test is a benchmark’s success checker, and the error is a gripper that closed on air. We call this setting physical coding and study it on a full manipulation benchmark.

Robot learning has mostly taken two other routes. Learned visuomotor policies [4, 5, 24] replace the agent with a network trained on demonstrations; language-model robot systems [2, 20, 26] keep the model but surround it with perception modules, skill libraries and planners written by hand.

RoboShell inverts this division of labour: the agent writes the primitives, they persist across episodes and are what is optimized, and the per-episode program is short.

First, evolved tools lift the same model by 19 points and past the strongest learned systems.

Second, tools move the boundary where geometry is enough and little elsewhere.

Third, failures have signatures, and some are missing information rather than missing capability.

The 19-point comparison uses the current-tools aggregate. The evaluation protocols are distinguished below.

RoboShell

RoboShell has two parts (Fig. 3): a harness that turns commands into motion (Sec. 3.1) and an evolution loop that turns failures into tools (Sec. 3.3). Four principles shaped both. The harness translates and never decides. Ground truth is isolated by architecture: the container the acting agent runs in holds no simulator, task or scoring code. Success is judged only by the benchmark’s own checker, which we do not modify. The interface description states interfaces only; strategy lives in tools.

Architecture showing an observation-only acting agent above a tool-development loop with privileged failure logs.
Figure 3.RoboShell.Top, one episode: ① a coding agent in a sandbox issues robo commands; the server validates each one, turns it into joint motion and returns feedback and RGB-D observations.

The harness

The acting agent is a stock coding agent CLI running in a container with no direct network access, no API key and no host directories. It is given one executable, robo, a directory obs/ that is refreshed on request with three 640×480 RGB-D views (head and two wrist cameras, depth in metres, OpenCV intrinsics and extrinsics), the two tool centre point (TCP) poses, and an interface description of about sixty lines.

Tools

A tool is a directory with tool.py and interface.md. The Python file declares a name, its sub-commands and their arguments, and a run(api, command, args) function over an EpisodeAPI that offers joint angles, TCP poses, the three RGB-D images with calibration, and the same motion primitives the base commands use. Tools cannot import the simulator, cannot read object poses, and are loaded only if listed in the task’s enabled_tools.txt.

The evolution loop

One command starts an unattended run for one task on one GPU. The loop uses the first ten layouts of evaluation seed 0 (for Generalization tasks, five standard and five random, which share tools) and proceeds layout by layout. The acting agent attempts the layout with the current tool pool. On success, the optimizer agent reads the trajectory and appends a short entry to playbook.md on which tools worked and why.

On failure, it reads the true object trajectory, the command log, the per-command camera frames and the current tool code, writes a diagnosis to skill.md, and adds or repairs a tool. A delivery check (check-task) then verifies that files are complete, tools load, and the interface text contains no task or object words that would leak strategy into the prompt; if the check fails the edit is reverted.

After the tenth layout the optimizer consolidates the three documents, the final tool set is frozen, and every layout is retested once with no optimizer; the retest count is the only number we report as a result.

Experiments

Setup

Both agents are the unmodified Codex command-line agent with GPT-6 at medium reasoning effort; each task gets one evolution on one RTX 4090 with a 12-hour development cap.

Round 1 is the fully automatic loop on all 42 tasks, scored by a frozen retest of one episode on each of the ten development layouts (the first ten of seed 0; five standard and five random for Generalization tasks).

Official runs frozen tools through the benchmark’s own evaluation client on all three seeds with the native episode counts.

Round 2 repeats the loop on the tasks where PhysicalRSI was clearly ahead, adding a human-written note on what the checker requires (Sec. 5.3).

Main results

Fully automatic, round 1 reaches a macro-average of 37.6% over the 54 variants, against 26.1% for PhysicalRSI and 24.3% for the same model zero-shot (38.0%, 31.4% and 22.5% on the leaderboard’s metric).

With the current tools the average is 40.9% against 26.1% for PhysicalRSI and 30.6% for VPP2 (41.5% against 31.4% and 32.3% on the leaderboard’s metric).

RoboDojo comparison on the leaderboard’s metrics

Success rate (%)

0–100 scale

RoboShell
41.5
VPP2
32.3
PhysicalRSI
31.4
GPT-6 Astra, zero-shot
22.5

Progress score

0–100 scale

RoboShell
48.1
VPP2
39.3
PhysicalRSI
36.3
GPT-6 Astra, zero-shot
29.0

Overall values from Figure 2. RoboShell combines official-client results for 27 variants with development-layout retests for the remainder; current tools include second-round interventions. Baselines use their published protocols.

Grouped bar charts comparing success rate and progress score by category. RoboShell overall success is 41.5% and progress is 48.1%.
Figure 2.RoboDojo comparison on the leaderboard’s metrics: success rate (a) and progress score (b).Overall is the leaderboard’s score, the mean of its five task types (Generalization: mean of its standard and cluttered variants). RoboShell uses frozen tools, scored under the official client for 27 of its 54 variants and on the development layouts for the rest (Tab. 2).

The zero-shot baseline uses a different harness, so we also compare within ours to check that the gain comes from the tools and not from the harness. The first development episode of each run had no task tools, only the shared pick tool.

On the 37 tasks whose log records it, the agent passed 4 of these layouts; the frozen retest of the same layouts, with the same harness and model, passed 16, turning 12 failures into successes and none the other way (exact McNemar test, p < 0.001).

Initial attempts4 / 37
Frozen tools, same layouts16 / 37

Under the official evaluation client

This section covers the round-1 tools of sixteen tasks, the fifteen with a round-1 retest of at least 7/10 and fold clothes; six second-round tasks follow in Sec. 5.3, for 3,286 episodes over 27 variants in all (Tab. 3 and Fig. 5, supplementary).

These are runs on our machines, not leaderboard entries.

With round-1 tools RoboShell reaches an unweighted macro-average of 73.2% over the 20 variants (2,391 episodes), against 44.7% for PhysicalRSI and 48.3% for the zero-shot model on the same variants, and is higher than PhysicalRSI on 13, tied on one and lower on six.

Development layouts76.2%
Unseen layouts73.0%

The official runs contain the development layouts, the first ten of seed 0, and on them the official client passes 76.2% (macro-average over variants), against 73.0% on all other layouts; most of the gap is run-to-run variation of the same tools on the same scenes.

Tools that compute everything from the observed geometry transfer almost unchanged (cover blocks 150/150, stack blocks 147/150 over both variants), while play tic tac toe falls from 9/10 to 33% and insert tubes from 8/10 to 52%.

Clean versus cluttered layouts

In round 1 RoboShell passes 25 of 60 standard and 20 of 60 cluttered layouts, keeping four fifths of its success; PhysicalRSI keeps about one fifth (25.7% to 5.6%).

Much of this robustness belongs to the model, which keeps 87% of its clean-layout success zero-shot (32.7% to 28.3%). With the current tools RoboShell keeps 79% at a higher level (46.1% to 36.3%), whereas the leaderboard’s 21 VLA and other learned-policy entries keep 1–68% (median 22%) and its five WAM entries 5–25%.

What the agents built

Across the 42 tasks the round-1 optimizer wrote 96 perception commands and 141 execution commands; the shared pool still contains only the pick tool.

The library is highly redundant: a surface measurement from a seed pixel was reinvented in nineteen tasks and a transfer macro in twenty, with interfaces that differ in argument names more than in meaning.

Model picks the pixel, code measures the world; this split keeps the tools deployable and auditable, and, as Sec. 5 shows, it is also where their limits come from.

Where the boundary lies

A capability map

The physical interaction decides. Rigid manipulation reaches 57.7% and pushing 75.0%; tight insertion reaches 21.7%, deformable and fluid material 20.0%, and articulated objects and tool use 0%.

A tool fixes one interpretation and one open-loop motion; an agent acting step by step can still notice and correct. Self-evolution trades that adaptivity for repeatability, which pays where the motion can be planned from geometry and costs where it must be found during the motion.

How episodes fail

Of the 420 retest episodes, 262 failed: 56% exhausted the budget, 32% ended with the agent declaring completion and the checker disagreeing, 11% ended with the agent giving up while a quarter or more of the budget remained, and two were terminated by the checker for a rule violation.

Missing information or missing capability?

For the eight tasks on which PhysicalRSI was clearly ahead after round 1 we reran the loop with one addition: a person reads the public task definition and writes a short note on what the success checker requires and where the unattended run fell short, given to the optimizer at every round.

Where the note supplies information the instruction lacks, three tasks go from 0–40% to 95–100%.

These are development-layout results; the paper reports official-client follow-up results separately.

Where the gap is capability, fold clothes does not move: the optimizer rebuilt the folding tool from the checker’s keypoint windows and fold order, but retests passed 2/10 and 4/10 against 3/10 before; a note cannot supply a closed-loop skill.

build tower, rebuilt by hand by a team member, went from 0/10 to 5/10 and reaches 56.0% under the official client (PhysicalRSI: 28.7%).

Time and cost

A retest episode takes a median 520 s of wall-clock time for 22.6 s of robot time, almost all model latency, which RoboDojo does not charge.

Limitations

Round-1 numbers come from ten seed-0 layouts per task and overstate official results (81.5% against 73.2%). The optimizer reads true poses, so the loop cannot yet run on a real robot without motion capture.

Conclusion

A stock coding agent behind a harness that only translates commands into motion, improved by a second copy of itself writing tools from its failures, rises from 22.5% to 38.0% on 42 tasks, past the strongest learned entries.

Where the right motion follows from measured geometry, tools hold across layouts and clutter; where success must be felt or watched, a frozen macro is no substitute for closed-loop control.

Next, the same loop will run in RoboCasa [33] and RoboTwin [9], then on real robots, and its traces will train a coding model of our own.

Abridged from the paper. Section, figure, and reference numbers within excerpts refer to the full PDF.

Research figure