All posts

Research

Make code as policy great again: Frontier agents write, call, and evolve robot tools

A pale-blue isometric workbench with two robot arms folding a blue cloth, a code terminal, and a robot tool controller.
On this page

Building on Physical Coding, our paper Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools introduces the Universal Robot–Agent Interface (URAI), through which frontier models and people share the same robot tools. A programming agent writes and improves tool code; an execution agent selects tools from observations and uses execution feedback to decide what to do next. Validated tool improvements persist without updating foundation-model weights, providing a concrete path for physical-world agents to accumulate capability through execution, feedback, and code revision. Across the tested simulation tasks, URAI improves success rates over direct end-effector control for all four execution models, while three also run faster and use fewer output tokens. We further examine the design through seven real-robot tasks and a tool-evolution case study.

URAI system introductionThe shared interface for selecting and executing robot tools.

01 · Step-by-step decisions are slow and costly

Frontier models can already control robots from visual observations. The next question is how much action a single model decision should cover—and who should decide what happens after that action finishes. Three approaches draw this boundary differently.

Direct end-effector control

One model call produces one end-effector command. A model such as GPT-6 Astra observes camera images, chooses a target position, orientation, and gripper state, then inspects the result before choosing another action. This keeps the model involved throughout execution.

The cost is repeated waiting and reasoning. A pick-and-place operation contains several phases: approach, grasp, transfer, place, and retreat. When these phases are split into many separately chosen commands, every observation–reasoning–action cycle adds latency and token use. In an independent Robocurve block-placement evaluation cited in our paper, GPT-6 Astra succeeded in 19 of 20 trials, averaging 2.5 minutes and 2.1k output tokens per episode.

Code as Policies

The model first writes an executable policy. The program calls perception and action tools, and uses conditions and loops to organize execution. What to check, how to interpret the result, and what to run next are determined by the generated code. This reduces repeated model calls, but its ability to respond is bounded by the checks and recovery logic encoded in that program.

Can a robot instead benefit from the iterative programming process familiar from Vibe Coding—using intent and feedback to build tools, while letting the model choose a complete motion and reconsider the situation afterward?

Code as Policy with URAI

URAI packages motions as executable tools. The model selects a tool and its parameters from the current observation; one call lets local code carry out a complete, multi-stage motion. After it finishes, the model receives fresh observations and execution feedback before selecting its next call. Code executes the motion; the model retains decision-making authority between motions.

What one model call controls. Direct end-effector control, Code as Policies, and Code as Policy with URAI assign different responsibilities to one model call.
Figure 1. What one model call controls Direct end-effector control, Code as Policies, and Code as Policy with URAI assign different responsibilities to one model call.

02 · Models decide, code acts

URAI separates writing tools from using them. These two roles operate on different timescales.

The programming agent starts from task intent and writes reusable or task-specific robot tools. Execution feedback and human natural-language guidance inform revisions. Candidate changes are validated before the resulting tool version is frozen for subsequent evaluation.

The execution agent selects tools and fills in their arguments from observations. After each call, it uses the new evidence to decide whether to continue, change its approach, or stop. Foundation-model weights remain fixed throughout this process.

The agent stays in the loop through URAI. A programming agent builds validated tools; the execution agent selects calls from observations and feedback. People and agents share the GUI/API backend.
Figure 2. The agent stays in the loop through URAI A programming agent builds validated tools; the execution agent selects calls from observations and feedback. People and agents share the GUI/API backend.

Consider pick-and-place again. The programming agent implements approach, grasp, transfer, placement, and retreat as a validated tool. The execution agent identifies the operation in the camera image, specifies the relevant pixels and arm, and calls that tool. The URAI backend grounds the pixels in 3D, plans and checks the motion, and executes the timed trajectory locally. The model reads the next observation only after the tool has returned.

This differs from handing an entire episode to a generated program. The motion is implemented in code, but the next decision remains with the execution model. Our same-tool comparison below tests the value of this boundary.

The same tools for people and agents

The GUI and API share one backend and the same execution semantics. A person marks an operation on a camera image; an agent submits the corresponding pixel coordinates and parameters through the API. Both invoke the same planning and control code. This also creates a useful diagnostic comparison: if a person succeeds with a tool that an agent struggles to use, the problem may lie in tool selection or argument grounding rather than the motion implementation.

Improvements that persist in tools

Validated revisions persist across episodes as tool code and execution strategies. This is a form of system improvement without retraining the foundation model. During evaluation, the accepted tool version is held fixed, separating development from the measurement of its performance.

Tool development and closed-loop execution. Tools evolve between episodes, while the execution agent makes decisions within each episode. Foundation-model weights stay fixed.
Figure 3. Tool development and closed-loop execution Tools evolve between episodes, while the execution agent makes decisions within each episode. Foundation-model weights stay fixed.

03 · Simulation experiments

We test URAI in RoboDojo on five tasks: pouring liquid into a cup, pouring balls into a vase, folding clothes, swapping blocks, and stacking blocks. The execution models are DeepSeek-V4-Flash, GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5.

Each model is evaluated on five seeds per task using both the native direct-control interface and URAI: 25 episodes per model and condition, or 100 episodes for each interface overall. The two conditions use matched task seeds and simulation-step budgets. Success is determined by RoboDojo’s official terminal evaluator.

Success rates on five RoboDojo tasks. Results for native control, URAI, and programs written in advance over the same tools. Five evaluation seeds per task and model.
Figure 4. Success rates on five RoboDojo tasks Results for native control, URAI, and programs written in advance over the same tools. Five evaluation seeds per task and model.

All four models improve with URAI. Across the tested episodes, overall success rises from 18% to 53%.

Source: the paper’s simulation results. This is a five-task study, not an evaluation on the full RoboDojo leaderboard suite.

Do tools alone explain the improvement?

To isolate the role of runtime decisions, we give GPT-6 Astra and Claude Opus 5.5 the same URAI tools but ask them to inspect the scene and write a complete program in advance. That program can branch on tool return values, but it cannot call the model again during execution.

The prewritten programs achieve 24% success. Keeping the model in the loop after each tool call achieves 56% across the same two models. Tools provide a useful action interface; retaining model judgment between tool executions contributes additional value.

RoboDojo execution costs. Mean episode time and execution-agent output tokens, including failed episodes and reasoning tokens. Tool-development costs are excluded.
Figure 5. RoboDojo execution costs Mean episode time and execution-agent output tokens, including failed episodes and reasoning tokens. Tool-development costs are excluded.

Execution time and token use

We also measure mean wall-clock time and the execution agent’s output tokens, including reasoning tokens and failed episodes. GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 are approximately 1.3–1.5× faster with URAI and use approximately 1.5–1.7× fewer output tokens. DeepSeek-V4-Flash shows smaller reductions, indicating that observation and reasoning costs still matter even when fewer action calls are needed.

These are execution costs; they do not include the programming agent’s tool-development cost.

The comparison also has task-specific limits. For SwapBlocks, URAI bundles pad placement and button interaction into tool semantics, so the improvement cannot be attributed purely to coarser action granularity. Excluding that task, success is 42.5% with URAI versus 22.5% with native control.

04 · Real-robot experiments

We deploy URAI on an AgileX PiPER-X dual-arm robot with four RGB-D cameras. GPT-6 Astra acts as the execution agent across seven tabletop tasks.

1. Unscrew a bottle cap

3/3 successful trials. Mean execution time is 5.1 minutes, approximately 3.5× faster than the cited published run with the same model. The first trial’s initial diagnosis-and-repair phase is excluded from the reported cost.

Unscrew a bottle cap3/3 successful trials in the reported evaluation.

2. Play tic-tac-toe

3/3 successful trials. Mean execution time is 5.0 minutes, approximately 2.7× faster than the published reference. Wall-clock time includes the human opponent’s moves.

Play tic-tac-toe3/3 successful trials in the reported evaluation.

3. Place a block into a bowl

3/3 successful trials. The paper reports a 2.4× speedup over the published reference and 4.7× fewer output tokens, averaging 443 output tokens per episode.

Place a block into a bowl3/3 successful trials in the reported evaluation.

4. Serve a hot dog

10/10 successful trials.

Serve a hot dog10/10 successful trials in the reported evaluation.

5. Pour blocks

10/10 successful trials.

Pour blocks10/10 successful trials in the reported evaluation.

6. Fold clothes

8/10 successful trials, with 90% mean task progress.

Fold clothes8/10 successful trials; 90% mean task progress.

7. Toss blocks into a bowl

7/10 successful trials, with 91.7% mean task progress.

Toss blocks into a bowl7/10 successful trials; 91.7% mean task progress.

The first three timing comparisons use published runs of the same model on other robot platforms. They are not paired, same-hardware comparisons: hardware, control settings, and operating conditions differ. The trial counts and these differences limit what the speed ratios establish.

05 · Tool evolution

Because tools are code, a failure observed during execution can become the basis of the next revision. In RoboDojo’s simulated clothes-folding task, we do not implement a dedicated folding routine. Instead, under human natural-language guidance, the programming agent iterates seven times on an existing two-point pick-and-place tool.

The initial tool treats clothing as a rigid object and rejects all cloth-operation requests. Revisions remove that obstacle and expose a different limitation. With the final tool, a person choosing operation points through the GUI succeeds in 4 of 5 trials. When the model selects those points autonomously, it succeeds in 3 of 20 and 2 of 20 trials on two evaluation seed sets.

Failures often begin at the same step: the first sleeve grasp targets the sleeve tip, and the gripper closes without catching the fabric. The tool can carry out a fold, but selecting an effective grasp point remains difficult. Sharing the same tools between people and agents helps separate motion-tool capability from visual argument-grounding errors.

Tool evolutionIterative revision of the cloth manipulation tool.

An execution produces feedback; feedback informs a code revision; a validated revision becomes available to later executions. The foundation model’s weights stay fixed. What persists is an improved tool implementation or execution strategy.

What comes next

URAI is one concrete realization of Physical Coding: use an agent’s programming ability to build robot tools, preserve the model’s judgment during execution, and let physical feedback inform the next development cycle.

The current results are a starting point. We will extend the task set and robot platforms, include tool-development costs in evaluation, and test how broadly these reusable improvements transfer. Better visual grounding, independent validation, and accounting for the full cost of iteration will be essential to understanding when tool evolution produces reliable, lasting gains.

For the broader research direction, read Self-Evolving Coding Agents. The full methods and results are available in the URAI paper.

Research figure