All posts

Position Paper

Self-Evolving Coding Agents

From Digital Programs to Physical-World Intelligence

A physical coding laboratory An isometric laboratory combines a physical code terminal, a compact robot placing a block into a shallow bowl, a pendulum, a spring experiment, and two coupled oscillators. Animation-ready groups are static by default.
Contents

Abstract

Physical Coding represents task state and execution as code, making the conditions behind physical actions explicit and revisable. Code as World records objects, relations, observations, constraints, and progress; Code as Policy organizes planning, tool calls, verification, and recovery. HexaAnything implements this interface through a Harness that connects language models to perception, planning, and control tools, including VLA/WAM policies. Fresh observations and independent evaluation determine whether execution should continue, recover, or stop. The resulting traces preserve how decisions were made and checked, so experience can return as reusable programs, memory, or training data.

On RoboCasa365, HexaAnything raises Composite-Unseen success from 34.3% to 38.3% and overall success from 56.6% to 61.1% over the native XR-1 VLA. HexaModel v0.1, a 27B model trained on Harness-returned and general-domain data, improves over its base model on every split and reaches 39.5% Composite-Unseen success. Separate experiments revise tools against a fixed evaluator, carry out quantitative science in the simulated PhyBench laboratory, and transfer the interface to a dual-arm AgileX robot, where five of seven tabletop tasks succeed in all three trials. Together, these results provide initial evidence for Harness, data, model, and tool improvement. The broader program connects these updates to better representations, environments, and embodiments through independently evaluated, versioned changes.

Introduction

VLA and WAM systems connect observations and instructions to robot actions [7, 8, 24, 40, 42, 106]. This gives robots a compact interface from perception to motion, but a long-horizon task also requires a record of which objects have moved, which constraints still hold, which subgoals remain, and when recovery is necessary. A controller can finish moving while an ingredient remains outside the oven; closing the door then makes the remaining placement impossible. An action chunk and stop signal do not expose the precondition that should have prevented that decision.

Benchmark success can also hide dependence on the initial scene. In the paper's LIBERO experiment, a policy trained with instructions masked still succeeds in 92.3% of trials, compared with 96.2% when instructions are present. Related robustness studies report substantial drops under changes to viewpoints, robot state, or object layouts [20, 104]. These observations motivate an interface that makes task requirements and intermediate progress available for inspection throughout execution.

Digital coding agents offer a precedent: models write programs, call tools, inspect results, and revise procedures against external tests [12, 85, 95]. Code can preserve state, express branches and loops, expose the step responsible for a failure, and retain a tested correction as a versioned artifact. Physical execution makes that loop more demanding because actions may be delayed, partially observed, noisy, or irreversible. State must therefore be grounded in images, depth, proprioception, and tool outcomes, with the source of each observation retained. Completion must be checked independently of the model that proposed the action.

HexaAnything connects these components through the world–policy interface in Fig. 1. A successful execution can become a reusable program, memory entry, or training record. A failure can identify a missing predicate, an unsuitable tool, or an incomplete recovery branch and motivate a specific edit. This connects the execution loop to a longer model–Harness–environment loop: interaction produces evidence, evidence updates the system, and the updated system determines what happens in the next round.

The evaluation tests this idea at several levels. Manipulation experiments hold the underlying action model fixed to measure workflow, verification, and recovery. A separate data-to-model experiment places a trained checkpoint back in the same Harness. Tool-revision experiments retain a fixed task and evaluator, while PhyBench and the real robot test how the interface extends to measurements and physical execution. These are controlled, partial steps toward self-evolution; coordinated improvement of all the system's components remains a longer-term goal.

Physical Coding: an executable world–policy interface
Figure 1. Physical Coding: an executable world–policy interface Code as World and Code as Policy connect task-relevant state, planning, action tools, verification, and persistent records. Original Figure 1 from the paper. In the paper ↗

Background: From Action Models to Coding Agents

Benchmark success need not establish instruction following. A QwenGR00T policy using a Qwen3-VL-4B backbone [5] and GR00T-style action head [7], trained in StarVLA [74] on LIBERO [49], succeeds in 92.3% of trials with instructions masked, against 96.2% with instructions. These results average three suites; the masked policy trails by at most 6.0 percentage points on any suite. LangForce reports rates within 1.1 points of ours [46]. When the initial scene determines the task, a scene-to-trajectory mapping can achieve high success.

Instruction ablation on LIBERO
Figure 2(a). Instruction ablation on LIBERO Instruction-conditioned and instruction-masked QwenGR00T policies, evaluated on 500 trials per suite. Redrawn from Figure 2(a), with all reported values preserved. Differences are in percentage points; the separate rollout panel is not shown. In the paper ↗Download PNG ↓

Code as Policies composes perception and control primitives in executable programs [47]; related agents maintain scene graphs, state records, or spatial constraints [25, 33, 34, 66, 99]. Physical Coding extends this approach with explicit verification and recovery, while retaining VLA/WAM policies as action tools. Video prediction alone does not supply checkable task state: visual realism and physical understanding can diverge [37, 57].

VCode provides a complementary route, representing images through executable SVG [48]. HexaAnything builds its world program from images without reading simulator state. Observation provenance and independent verification connect that representation to evidence. Persistent skills or better scaffolding must also be distinguished from model learning: changing a Harness does not itself change model weights [76, 81, 101].

Physical Coding

Code as World and Code as Policy

“Coding” includes behaviors, controllers, evaluators, and environment transformations expressed as executable specifications. Two coupled representations make this useful for physical tasks.

Code as World. An executable account of task-relevant objects, relations, state variables, observations, constraints, and progress predicates.

Code as Policy. An executable account of task decomposition, tool selection, action calls, progress checks, recovery, and replanning.

A policy can call a controller but needs world predicates to determine whether an ingredient remains outside the oven. A world description needs a procedure to act or recover. Together, they make progress inspectable, failures localizable, and proposed changes testable.

Two tasks make the connection concrete. In LoadKebabSandwich, the world program records whether the bread and kebab are inside the oven; the policy allows the door to close only when both conditions hold. In a Hooke's-law experiment, the world program records each ruler reading together with its applied load and source observation; the policy chooses the loads, operates the arm, reads the ruler, and fits the spring constant. Both tasks link a physical action to an explicit state condition and the evidence needed to check it.

Coupling world state, policy execution, and evidence
Figure 3. Coupling world state, policy execution, and evidence Observations and tool outputs update Code as World; Code as Policy plans, acts, verifies, and recovers. Fresh evidence determines whether to continue, re-observe, recover, or stop. Original Figure 3, including manipulation and scientific-experiment examples. In the paper ↗

Formally, the model writes a world program W and policy program P; the Harness executes P using W, the environment, and the task specification. The resulting trajectory supplies evidence to an independent verifier. The model provides interpretation and synthesis; the Harness provides tools, permissions, execution state, and rollback.

World entries are updated from scene observations, instrument measurements, and robot or tool state. Each retains its source and observation time; a measurement also retains the experimental condition that produced it. After a tool call, the agent observes again, updates the relevant entries, and checks the predicates the action was meant to change. If evidence cannot determine whether a condition holds, the next step can be another observation rather than an unsupported decision.

Verifier verdicts distinguish pass, fail, insufficient evidence, blocked, and safety stop. Traces enter training only after independent verification and contamination checks. Tool execution alone does not establish physical completion.

Because these programs govern the agent’s own execution, failures can motivate edits to its observation abstractions, predicates, tool compositions, or recovery procedures. State, procedure, and their revision share an executable form. A failed episode can become a localized program change; a verified trace can become memory or training data. The same interface supports scientific experiments, where the agent must operate instruments and check whether measurements support a conclusion.

Physical Self-Evolution

Self-evolution follows a generate–evaluate–select loop [2, 52, 68, 78, 100]. A round fixes tasks, environment, seeds, budgets, and evaluation. Execution traces expose failures; diagnosis proposes bounded edits to a versioned bundle of world and policy programs, Harness, data, memory, and evaluator configuration. Candidates undergo static checks, regression tests, and held-out comparison against a frozen parent.

One successful rollout is insufficient. Admission requires independent evidence of improvement, preserved provenance and compatibility, and no regression on protected suites. Evaluator changes need independent testing to prevent score inflation. Rejected edits leave counterexamples and retain the parent as a rollback point; accepted traces can support the next program or model update.

Four coupled targets organize these updates:

  • Harness: tools, observation abstractions, predicates, workflows, verification, and recovery.
  • Data and environments: demonstrations, memory, task distributions, curricula, simulators, and counterexamples.
  • Model: planning and tool use, adapters, post-training, and weights; architecture and routing are longer-term targets.
  • Embodiment and compute: sensors, actuators, morphology, control interfaces, and training or inference hardware.
Four coupled targets of physical self-evolution
Figure 4. Four coupled targets of physical self-evolution A seven-step evaluation cycle governs changes to the Harness, data and environments, model and algorithms, and embodiment and compute. Original Figure 4; broader representation and hardware evolution remain research directions. In the paper ↗

Simulation supplies controlled perturbations and inexpensive resets. Physical environments contribute sensor noise, calibration error, contact variation, latency, and failures simulation may omit. Both supply experiences and counterexamples that can revise the executable interface.

Representation evolution cuts across all four. The agent may eventually revise object types, constraints, progress predicates, tool schemas, or executable language constructs when its existing interface cannot explain a failure. These changes require evaluation with the tasks, model, Harness, and embodiment that will use them.

Attribution remains essential: stronger recovery, easier tasks, or a permissive evaluator can raise success without improving the model. Reduced Harness assistance and held-out compositions test whether trained capability has been internalized. The present evidence covers Harness and tool revision and a first data-to-model update. Sustained co-evolution of representations, model architectures, hardware, and physical deployment remains a research program, subject to the same independent evaluation and safety gates.

HexaAnything: A Physical Coding Agent

HexaAnything couples world and policy programs with a Harness for tools, workflow, verification, recovery, and memory. The action model supplies a tool; the Harness owns the execution loop. It updates state from observations, checks outcomes independently of completion claims, and decides whether to continue, recover, or stop. During execution, the verifier checks world predicates against available sensor and tool evidence; in simulation, the benchmark's official success check assigns the trial label. Disagreements between the two become counterexamples for later improvement.

The task and acceptance reference remain fixed; observations retain their evidence source. A step limit records a stopping condition, not automatically a task failure. Verification distinguishes insufficient evidence and safety stops from success and failure.

The proposed roadmap begins by establishing tools, observation, verification, recovery, and logging with existing models. A second stage uses failures to revise tools and workflows in a sandbox, then trains on validated traces so the next model can organize the next round of improvement. A third stage brings the loop into constrained physical workflows, where engineers, sensors, and execution logs reveal calibration, contact, and latency failures that simulation may omit. Independent evaluators, provenance, regression checks, human oversight, and rollback govern candidate changes throughout this process.

HexaAnything: model and Harness
Figure 5. HexaAnything: model and Harness The model writes world and policy programs; the Harness connects observation, tool use, verification, recovery, and execution evidence. Verified execution products can support data, memory, and system revisions. Original Figure 5 from the paper. In the paper ↗

Experiments: Harness, Model, and Tool

Harness: Organizing Action

On RoboCasa365, XR-1 [19] runs natively or as a tool. Conditions share maximum environment-step budgets and, unless noted, tasks and seeds.

38.3% Composite-Unseen HexaAnything with GPT-5.6-Sol improves on native XR-1’s 34.3% by 4.0 points; same-model Codex reaches 34.1%, 4.2 points lower.

61.5% Composite-Seen Harness decomposition, verification, and recovery improve on native XR-1’s 54.8% by 6.7 points.

61.1% Overall The Harness improves on native XR-1’s 56.6% and same-model Codex's 59.5% across the three splits.

With Qwen3.8-27B as planner, Composite-Unseen success reaches 37.3%, 3.0 points above native XR-1. The Harness advantage is strongest on composite tasks: Codex leads HexaAnything by 0.8 points on Atomic-Seen, while HexaAnything leads by 1.7 points on Composite-Seen and 4.2 points on Composite-Unseen.

Table 1. Harness comparison on RoboCasa365. XR-1 runs natively or is called as a tool by a coding agent; both coding agents are driven by GPT-5.6-Sol. Entries are success rates. Each task is evaluated on 50 seeds across 18 Atomic-Seen, 16 Composite-Seen, and 16 Composite-Unseen tasks; Overall pools all trials.

SplitXR-1 (native)HexaAnythingCodex
Atomic-Seen78.0%80.9%81.7%
Composite-Seen54.8%61.5%59.8%
Composite-Unseen34.3%38.3%34.1%
Overall56.6%61.1%59.5%

On three long-horizon tasks, gains are 31.0, 15.0, and 17.0 points. The unchanged VLA executes calls; the coding agent maintains and verifies subgoals.

Table 2. RoboCasa365 Composite-Unseen case study. Entries are success rates over 100 seeds. HexaAnything calls the same XR-1 as a tool inside an editable workflow.

TaskXR-1 (native)HexaAnything
LoadKebabSandwich17.0%48.0%
PortionHotDogs22.0%37.0%
WaffleReheat53.0%70.0%

In LoadKebabSandwich, native runs sometimes close the oven after placing one ingredient, leaving the other unreachable. The agent verifies both ingredients are inside before closing. In PortionHotDogs, native runs may leave sausage in the bowl, repeatedly grasp while a plate remains empty, or stall after placing the breads. The agent rechecks incomplete plates and recovers from drops, misplaced items, and stalled calls instead of accepting the VLA’s termination signal (Appendix B).

LoadKebabSandwich: native XR-1 and HexaAnything
Figure 6. LoadKebabSandwich: native XR-1 and HexaAnything A simulation rollout from the same initial state (seed 3). Native XR-1 closes the oven with kebab outside; HexaAnything checks the intermediate state before closure. Original frames and labels from the updated Figure 6. In the paper ↗

Doubling the native PortionHotDogs budget yielded 21/100 successes, versus 22/100 normally and 37/100 with the coding agent. Monitor–interrupt–replan also raises CloseFridge from 50.0% to 79.0% and TurnOffStove from 0 of 8 baseline trials to 18.0%; these use different protocols and are not pooled. In one representative trace, execution summaries reduced active context from roughly 190K to an estimated 30–35K tokens, retaining the full trace on disk.

Model: Learning from Traces

HexaModel v0.1 fine-tunes Qwen3.8-27B on Harness-returned and general-domain data. The Harness supplies 9.3K VQA examples and 1.0K agent traces. The reported gains measure this combined training recipe; a data-source ablation would be needed to isolate the Harness data's contribution.

Training data and protocol

Traces include 703 successful base-model runs, 204 GPT-5.6-Sol recovery runs, and 103 recovery runs from an earlier fine-tuned checkpoint. Harness-returned data supply 38.8% of 178.6M training tokens. Thus, part of the training set comes from a previous model-update round; the remaining tokens cover embodied VQA, scene captions, math, code, and instruction following.

Table 3. Training data of HexaModel v0.1, trained for one epoch. The agent traces and the 9.3K RoboCasa365 VQA examples are returned by the Harness.

StreamContentRecordsTokens
Agent tracesSubgoals, tool calls, verifier outcomes, recovery decisions1,01042.4M
VQA with CoTRoboCasa365 state and task-completion judgments; general embodied VQA12,07541.7M
Scene captionsRoboCasa365 scene descriptions6,7642.0M
General dataMath, code, and instruction following24,49192.5M
Total44,340178.6M

In the same Harness, gains over the base model are 0.2, 1.3, and 2.2 points across the three splits, and 1.2 overall. The largest improvement is on Composite-Unseen, which places greater demands on decomposition and recovery. HexaModel exceeds GPT-5.6-Sol by 1.2 points on that split and 0.6 overall, with gains of 0.1–1.2 points across all splits.

Table 4. Model evolution on RoboCasa365. Each row is the planner inside HexaAnything, except XR-1 (native), which runs the VLA without the Harness. Entries are success rates. Each task is evaluated on 50 seeds (18 Atomic-Seen, 16 Composite-Seen, and 16 Composite-Unseen tasks), and Overall pools all 2,500 trials.

PlannerAtomic-SeenComposite-SeenComposite-UnseenOverall
XR-1 (native, no Harness)78.0%54.8%34.3%56.6%
Qwen3.8-27B (base)80.8%61.0%37.3%60.5%
GPT-5.6-Sol80.9%61.5%38.3%61.1%
HexaModel v0.181.0%62.3%39.5%61.7%
RoboCasa365 model and Harness comparison
Figure 7. RoboCasa365 model and Harness comparison Native XR-1 runs without the Harness; the other models plan within HexaAnything. Redrawn from the updated Figure 7 and Table 4 with a zero-based 0–100% axis. Values and the HexaModel v0.1 label follow the latest paper. In the paper ↗Download PNG ↓

Tools: Revising the Code

On RoboDojo [13], the agent diagnoses failed traces and edits tool code and its calling workflow while tasks, evaluator, and weights remain fixed. Improvement is non-monotonic: Pour vase initially falls from 40.0% to 0.0% before reaching 100.0% after the second revision.

Table 5. Tool self-evolution on three RoboDojo tasks. Each cell is the task success rate after the corresponding revision round; the task definition and the evaluator are the same in every round.

Revision roundFold clothPour vasePress by number
Round 00.0%40.0%0.0%
Round 160.0%0.0%0.0%
Round 280.0%100.0%100.0%

For Fold cloth, the agent revises the reusable urai_pick_place tool. The task requires both sleeves and the hem to satisfy RoboDojo’s evaluator. Human and agent share the tool interface; the agent improves a general two-point pick-and-place operation rather than writing a folding routine. Episodes contain three calls, each preceded by a fresh observation: left sleeve, right sleeve, and body fold.

The initial version rejects a shirt treated as a 60 cm-wide rigid object. Seven revisions, ending at development round 18, add thin-layer pinching, synchronized transport, wrist separation, and improved release, drop height, and clearance. Versions preserve the interface and carry unit tests; no branches target particular seeds, garments, or tasks.

On development seeds 0–4, eight versions succeed in 0, 0, 0, 0, 3, 2, 2, and 4 of 5 episodes. Revisions exchange failures; seed 3 always fails at the workspace edge. The operator had consulted diagnostic garment keypoints, which share their source with the official evaluator, when selecting pixels for three development seeds. The final tool was therefore frozen and re-evaluated using only public RGB-D images, camera calibration, and tool returns, with keypoint output removed and no evaluator source or internal state consulted during evaluation.

Under that constraint, the final tool succeeds on 3 of 5 development seeds and 4 of 5 held-out seeds, each run once in a fresh container without retries. The operator had read the evaluator's source in earlier development sessions and still selected the pixels for each call, so these results measure the tool and interface rather than autonomous perception. Claude Fable 5.1 performed the one-day development history, and Claude Opus 5 completed the re-evaluation.

Cloth-folding tool evolution
Figure 8. Cloth-folding tool evolution Eight tool versions across seven code revisions. Redrawn from Figure 8, preserving development-set success counts and the regression after v4. Diagnostic garment keypoints were used on three of five development seeds; pixel targets were operator-selected. Separate re-evaluation without those keypoints achieved 3/5 on development seeds and 4/5 on held-out seeds. In the paper ↗Download PNG ↓

Beyond Manipulation: Scientific Experiment with Physical Coding

PhyBench supplies objectives and apparatus; the agent designs procedures, operates the robot, records measurements, and fits models. The Isaac Sim laboratory [61] hides seeded reference values until submission and binds evidence to each episode.

From an initial observation, the agent identifies instruments and their labels, chooses experimental conditions, and plans the analysis. It records measurements with their conditions, checks residuals, and decides whether to measure again. PhyBench supplies apparatus and evidence-based evaluation independently of the participating system; scoring depends on physical outcomes, not tool choice.

  • Hooke’s law: load three to five weights, read a ruler, and fit force–extension data for stiffness. Weights must be physically supported and released, required loading conditions covered, and readings stable; the benchmark does not supply ruler readings. Success means ≤15% relative error.
  • Simple pendulum: use calibrated P1–P3 lengths and optical-gate periods to estimate gravity. Complete post-warm-up cycles, no robot contact, and amplitude below 5° are required; success means ≤3% error.
  • Coupled oscillators: identify two normal-mode frequencies from synchronized 120 Hz, 0.1 mm displacement records. Two independent initial displacements and free recordings of at least 40 s with measurable motion on both channels are required; each frequency must have ≤10% error.
PhyBench simulated scientific experiments
Figure 9. PhyBench simulated scientific experiments Original simulation frames from Figure 9: Hooke’s law, pendulum, and coupled-oscillator experiments. These illustrate a feasible process; the tasks do not prescribe this sequence of steps. In the paper ↗

Runs receive participant-side guidance on recording measurements, reading scales, and fitting in code, then proceed without intervention under 7,200 s and 18,000-frame limits. The model chooses its procedure and final estimate. A valid run satisfies the evidence protocol and has finite error; validity does not imply success. Mean relative error averages quantities, then valid runs. Failed or interrupted trials remain in N without numerical errors.

Table 6. Scientific experiment execution with HexaAnything in simulation. Entries give mean relative error (%; lower is better) over valid runs and the valid-run count n/N. A valid run is not necessarily below the task's success threshold.

Harness + ModelHooke’s law: k error ↓Simple pendulum: g error ↓Coupled oscillators: ω₁,₂ error ↓
HexaAnything + GPT-6-Astra2.3% (10/10)1.2% (10/10)0.8% (10/10)
HexaAnything + Opus 5.51.3% (10/10)1.0% (10/10)1.7% (10/10)
HexaAnything + GPT-5.6-Sol4.8% (9/10)1.9% (8/10)1.2% (10/10)
HexaAnything + Qwen3.8-Max12.4% (5/10)2.4% (3/10)3.7% (3/10)

Astra and Opus obtain valid results in all 30 trials each; Sol in 27 of 30. All three achieve mean errors below 5% on every task. Opus has the lowest error on Hooke’s law and pendulum, and Astra on coupled oscillators. Qwen3.8-Max produces valid estimates in 5 of 10 Hooke's-law trials and 3 of 10 trials on each remaining task, with mean errors of 12.4%, 2.4%, and 3.7%. Both measurement accuracy and completion depend on the underlying model. Appendix C provides additional example plans and runs.

Following one experiment from plan to estimate. In a Hooke's-law run with Opus 5.5, the agent planned five load levels from 0 to 0.20 kg, moved the arm away from the ruler before each reading, and required two images taken 2 s apart to agree. Each reading was stored with its load and source observation. A weighted linear fit with a free intercept estimated the spring constant as 24.82 ± 0.40 N/m, compared with the evaluator's reference of 25.16 N/m revealed after submission. The 1.36% error is close to the model's 1.3% task mean; the largest residual, 0.024 cm, is below the 0.10 cm reading uncertainty. The final quantitative claim can thus be traced back to five camera frames and the code that analyzed them.

From robot execution to a measured spring constant
Figure 10. From robot execution to a measured spring constant One Opus 5.5 Hooke’s-law run: five load levels, source-linked ruler readings, and a weighted linear fit yield k = 24.82 ± 0.40 N/m against a 25.16 N/m reference (1.4% rounded error). Original Figure 10; this is one run within Table 6, not its aggregate result. In the paper ↗

Physical Deployment, Limitations, and Open Problems

The AgileX PiPER-X has two 6-DoF arms and four RGB-D cameras. Through URAI, humans draw strokes on camera images; agents submit the same pixels and parameters through an API. Both use shared planning and control code.

In Vibe as Policy, a programming agent validates and freezes tools between episodes; a frozen execution agent composes them. Tools run at the robot’s control rate. GPT-6-Astra executed seven tasks:

AgileX dual-arm tabletop tasks
Figure 11. AgileX dual-arm tabletop tasks Original photographs from Figure 11, showing five of the seven tasks: unscrewing, hot-dog serving, tossing, folding, and pouring. These are initial setups, not successful terminal states. In the paper ↗

Table 7. Seven tabletop tasks on the AgileX dual-arm platform, each run three times with GPT-6-Astra as the execution agent calling URAI tools. Progress is the fraction of the task completed (blocks in the bowl out of six for toss blocks; the operator’s completion score for fold clothes). Tokens are the execution agent’s output tokens per episode, reasoning included; time is wall-clock from task release to a confirmed outcome and includes model latency, robot motion, and, in tic-tac-toe, the human’s moves. Published references use the same model on other hardware (GPT-Policy on its own robot; Robocurve on YAM arms with a 25% speed cap and at most 20 model calls) and are not same-robot measurements. The first unscrew-bottle-cap trial excludes an initial diagnosis-and-repair phase.

TaskSuccessProgressOutput tokensTime (min)Published reference (same model)
Unscrew bottle cap3/3100%1.52k5.1GPT-Policy R3: 3/3, 17.9 min
Tic-tac-toe3/3100%3.26k5.0GPT-Policy R8: 3/3, 13.6 min
Block into bowl3/3100%4431.0Robocurve: 19/20, 2.1k tokens, 2.5 min
Hot-dog serving3/3100%2.40k7.9—
Pour blocks3/3100%2.29k3.1—
Toss blocks into bowl1/377.8%4.69k8.1—
Fold clothes1/383.3%1.41k4.1—

Five tasks succeed in all three trials. Hardware and operating limits confound timing comparisons; three trials provide limited evidence. A remaining block lacked reliable depth near an arm base; two clothing trials stopped at 75% completion.

Robot tool revision uses operator judgment and execution traces, including joint speeds, gripper state, and landing position. A quarter-second firmware lag changed toss release from timer-based to measured arm position. One human-throw video prompted an overhand motion with three pitch joints; subsequent revisions corrected a wrist-limit failure, release timing, and near-base grasps.

The revised motion uses a low wind-up followed by a forward thrust. Near- and far-target trajectories adjust release to landing distance; motion continues through release so the gripper opens while the arm is moving. Tilted line-drawn grasps address blocks where straight-down approaches lack reliable depth. Two versions—an untrackably fast reference and an operator-rejected high lob—were rolled back. Each version was tried within the hour and committed. Claude Fable 5.1 programmed the tools, with Claude Opus 5 for part of one session.

The paper also designs longer RoboCasa365 tasks that combine many atomic actions into one instruction, start the robot at a random kitchen location, and require it to search for objects that may be hidden in drawers. Each episode has a 10,000-step limit. These tasks have been designed but no system has yet been evaluated on them; the reported long-horizon results use the benchmark's existing composite tasks.

The current Harness comparison uses one action model and one benchmark, with a 1.6-point overall margin over same-model Codex and Codex ahead on Atomic-Seen. Tool revision uses few seeds and operator-selected pixels; PhyBench uses ten simulated trials per model and task, and the real robot uses three trials per task. Broader testing must separate terminal success, recovery, verifier accuracy, and transfer across viewpoints, layouts, embodiments, and dynamics. Preventing model–verifier co-adaptation, learning physical causality from sparse traces, and combining expert corrections with simulated and physical failures remain open problems.

Conclusion

Physical Coding gives physical execution two coupled, persistent artifacts: a world program that records task-relevant state and a policy program that organizes action, verification, and recovery. Their connection exposes why an action was allowed, locates the condition behind a failure, and lets a proposed correction be tested against recorded evidence. HexaAnything realizes this interface by observing, acting, and receiving feedback through code, preserving experience beyond each episode.

The experiments show several ways that this feedback can improve a system. Keeping XR-1 fixed while adding the Harness raises RoboCasa365 Composite-Unseen success from 34.3% to 38.3%. Training HexaModel on the mixed corpus improves over its base model in the same Harness. Revising tools against an unchanged evaluator improves three RoboDojo tasks. Beyond manipulation, the same interface supports autonomous experiment design, measurement, and analysis in PhyBench; the three stronger planners achieve mean errors below 5% on each task. On the AgileX robot, five of seven tabletop tasks succeed in all three trials. These results support initial Harness, data, model, and tool evolution, while the sample sizes and evaluation boundaries define the current scope of that evidence.

The next steps connect four targets: the Harness, data and environments, the model, and embodiment and compute. Better world schemas, progress predicates, tool interfaces, and executable representations can contribute across all four. Testing whether learned capabilities persist with less Harness assistance and on held-out task compositions will help distinguish internalized model capability from stronger external support. Versioned parents, independent evaluation, protected regression suites, provenance, and rollback are needed to attribute each gain. Simulation offers controlled resets and counterexamples; constrained physical workflows expose sensing, latency, contact, and embodiment effects. Together, they form a path toward reliable self-evolution in manufacturing and science. Autonomous architecture search, unrestricted self-rewriting, and unconstrained physical deployment remain future work.

Research figure