Pantograph evaluated GPT-6 Astra and Claude Fable 5.1 on eight manipulation tasks using our Pandroid bimanual mobile robot, with an identical prompt, tool set, and budget for both models.
Astra completed 36% of attempts and Fable completed 15%. The difference is statistically significant (p ≈ 0.001, stratified by task).
#Setup
We use six bimanual mobile robots to perform 160 total rollouts across eight simple manipulation tasks, ten per model per task. We report success rates for Claude Fable 5.1 and GPT-6 Astra, both on low effort. Grading was blind to which model produced the attempt, and models were randomly interleaved across all robots. We impose a limit of 20 minutes or 96 turns, whichever comes first. Our harness is intentionally simple and generic. We explain the situation to the models and allow them to control the robots via toolcalls.
#Results
We find that Astra significantly outperforms Fable (p=0.001). Much of the gap comes from two tasks. Both models remain far from human performance, as teleoperators easily score 100% in much less time.
Full success by task
Time to completion
GPT-6 Astra on low effort takes faster turns with fewer tokens than Claude Fable on low effort. This means that within our time constraints, it is able to view and respond to more images. We suspect that this is part of the reason it is so dominant.
#Observations
VLMs are robust. They do not seem to mind dramatic glare and variations in lighting, or distracting background. They notice and recover well from failures.
The models also display some common LLM failures, like claiming that they have accomplished the task when they have not. For example, multiple runs end with the model confidently claiming that the towel is in the hamper, when it is clearly visible on the floor.
They struggle with vision, and depth in particular. In order to pick up a wooden block from the floor, they frequently make several attempts at descending heights, each time noticing that the grippers have not yet closed on the block. They also struggle with object permanence. When their own arm occludes an object, they frequently lose track of it.
They will sometimes command the hand into impossible positions, like clipping into the floor. With a traditional policy, you would use forceful safety constraints. With VLMs, the same can be accomplished with prompting. We added a warning when they were pushing into the floor, which fixed the issue while allowing them to skim the ground as needed to pick up things like spoons.
Small robots with handles made a big operational difference, since they can be quickly reset by hand between trials.
#Selected clips
#Methodology
#Harness
Each turn, the model receives images from three of the robot's cameras at 480p and may request an additional frame from a secondary head cam. It also gets a history of its recent actions and a summary of the results. It responds with one or more tool calls from a fixed set of parameterized actions, such as moving a hand a specified distance along an axis, rotating the base by a specified angle, or raising the vertical stage. The harness translates these actions into joint commands, warns the model if a commanded motion is driving into the floor, and returns a new set of images. The prompt is identical for all tasks and contains no task-specific guidance. We wanted to measure model capability separately from the gains you get from task-specific tailoring and harness engineering.
We made two small changes to the prompt during data collection. To prevent robots from interfering with each other, we added a line to the prompt telling them not to do that. We also changed part of the prompt that was too specific to balls and was possibly hurting the block tasks. These changes affected both models equally.
#Embodiment
All trials were run on the Pandroid, our bimanual robot with a mobile base and a vertical stage. Six robots were used in parallel. The vertical stage was particularly nice for VLMs because it is powerful (many tasks require grabbing something and then lifting it) and simple to reason about (one dimensional, no complex IK).
#Budget
For practical reasons and to simulate a realistic deployment constraint, we give each model a wall-clock time limit of 20 minutes to complete their task. These are simple tasks that take almost no time for human teleoperators to complete, so 20 minutes is intended to be generous to account for the limitations inherent to running a full big frontier llm in this way. Note that a fixed time budget means that models that think for longer have fewer turns in which to solve the task.
#Effort
A preliminary experiment on one task found that the high-effort setting reduced success rates for both models under our time budget. Given our time constraints, any improvement in quality from more thinking was dwarfed by the increase in response time. We also suspect that it is most important to get more images and more attempts, rather than more thinking. We therefore decided to use low effort for both models.
#Human baseline
A human teleoperator remotely performed each task five times using the robot's standard teleoperation interface and through the robot's cameras. This baseline establishes that every task is feasible on this hardware and provides a reference point for the models' completion times.
#Future Work
In this work, we showed that a simple, generic harness around a frontier language model can competently, albeit slowly, control a general-purpose bimanual robot.
We are excited to test many more models going forward, including future frontier releases and robot-specific models, and to greatly expand the range of tasks tested. We are also excited to test smaller models that may be able to act more quickly and make more decisions within the time limit. Lastly, we look forward to more carefully testing the trade-offs that come from higher reasoning levels.