Research

Fully general robotics will require new, scalable ideas.

Pan 1: A foundation model for Minecraft

A model trained on videos of Minecraft gameplay that can accomplish goals without training on them.

July 15, 2026

Read the report

Our approach

Fully general robotics models will require new, scalable ideas. We draw from the literature on generative models, self-supervised learning, goal-conditioned reinforcement learning, and the emerging science of deep learning, to develop new approaches for learning and acting in unseen environments.

We believe that simple, scalable methods ultimately perform best, at least once they have been given sufficient time and resources to mature.

The path to general models won't be linear, and there are many problems to solve along the way. Such problems include: modality agnostic reasoning, ultra-long context, in-context reinforcement learning, and many others.

Research directions

Chain-of-thought is currently the dominant approach to applying extra test-time compute to produce better outputs, but it's very specific to text. Are there methods of generative modeling that are independent of modality but still allow for scaling test-time compute? Iterative refinement models like diffusion and flow matching allow for increasing the number of generation steps, but quality stops improving after a certain point, perhaps because they don't allow for backtracking.

We're interested in general formalisms for generative modeling that could support reasoning at test-time without being reliant on text.

Typically, reinforcement learning is done using gradients at training time, either on one problem at a time or on a fixed distribution of tasks. Handling out of distribution problems at inference time happens through generalization, which requires the training task distribution to be broad enough so as to encompass any possible inference time task.

Generative modeling is different: by pretraining on varied sequences, in-context learning can be used at test time to predict completions of even extremely out-of-distribution inputs, even arbitrary sequences of digits.

Is it possible to set up reinforcement learning in a similar way? We want to use sequence models at inference time to sample-efficiently adapt to any reinforcement learning task using in-context information.

Robots will operate in the physical world for many hours, and will benefit from strong long-term memory potentially even more than language models do. We want robots to remember where things are, how users like for things to be done, all things that will be difficult to encode in text memories. Moreover, video contains much more information than text, and generally requires many more tokens to represent.

We're interested in methods for training models with ultra-long context that could naturally allow models to access weeks or months of visual memories.

Pretraining for robotics will require modeling both video and text. Currently, the best video models use continuous methods, e.g. latent diffusion and flow matching. The best text models, on the other hand, are usually discrete autoregressive sequence models. Eventually, the best multimodal models will probably be trained with a unified framework that can encompass both modalities. We're interested in ways of training language models in continuous spaces, towards that end.

Humans learn to read through vision, and how to speak by listening. Somehow we are able to process this high-dimensional data, identify the symbolic, meaningful parts of it, and spend our limited compute modeling that information. Image models trained on pixels of text, and audio models trained on speech, typically don't learn a meaningful language model.

How can we develop new generative modeling methods that learn to read from pixels and speak from audio without special-casing them? This problem is a microcosm of the more general problem of learning how the world works from high-dimensional data.

Join us

We're a small, fast-moving team working together in person in San Francisco. If you're excited about scaling general models that learn from and act in the real world, we'd love to talk.

We're looking for research scientists who want to scale simple methods across the largest sources available.

View the posting

We're looking for someone to own the systems our research runs on, from training infrastructure to the pipelines that move internet-scale video, keeping a small team moving fast on big machines.

View the posting