LLMs Revolutionize Robotics: Code-as-Policies In‑Context Learning

•

 29 min video

•

 6 min read

YouTube video ID: Jv5B5CEaPJI

Source: YouTube video by Y Combinator — Watch original video

PDF

The field of robotics is on the cusp of a major transformation, driven by the generalizability of coding agents and large language models (LLMs). This shift, highlighted by MIT Professor Philip Isola, suggests an era where general-purpose models could significantly enhance the capabilities of various robots. Companies like Wadd Labs and Robocurve are at the forefront of this innovation, leveraging LLMs to make robots more capable.

The Rise of LLMs in Robot Control

Early successful approaches to integrating AI with robots, such as the RT2 paper, utilized pre-trained language models on web text and images to control robots. These models were fine-tuned to output effector poses (coordinates for joint commands) instead of natural language. This marked a significant step, demonstrating that LLMs, even in their earlier forms, could be adapted for robotic control.

The evolution from these early models to current LLMs like Astra mirrors the progression seen in language models themselves. Initially, models were limited to direct action outputs, lacking the ability for complex reasoning or "chain-of-thought" processes. Just as language models evolved to perform multi-step reasoning (e.g., solving math problems by showing intermediate steps), current LLMs can now engage in more sophisticated planning for robotic tasks. This allows for the allocation of more computational resources for complex tasks, moving beyond simple action outputs.

The "Bitter Lesson" and Data Modalities

A key insight in this domain is the "bitter lesson," which suggests that architectural innovations often yield less progress than simply scaling up computation and data. For robotics, this means that instead of focusing solely on robot-specific architectures, leveraging the vast amounts of data used to train general-purpose LLMs can be more effective.

The success of RT2, for instance, stemmed from its ability to tap into the rich modality of web images and text, which significantly improved its performance compared to models trained only on robotics data. This highlights the power of transferring knowledge across different data modalities. The challenge lies in making traditionally "out-of-distribution" data, like robotics data, compatible with "in-distribution" data that LLMs are trained on.

Code as Policies: A Game Changer

The concept of "code as policies" has been instrumental in advancing robot control. This approach involves LLMs writing code to define complex robotic behaviors. Early research, particularly from Google's DeepMind team, demonstrated that coding agents could generate Python functions to control robots for intricate tasks like picking up objects or moving to specific poses.

What was particularly surprising was the one-shot learning capability of these coding agents. Because they were trained on extensive coding data, they inherently understood sequences of actions and logical steps, allowing them to perform new tasks without requiring additional robot-specific training data. This "in-context exploration" ability has been a major motivator for further research in applying LLMs to robotics.

In-Context Learning vs. Weight Updates

The discussion around how learning occurs in these systems often revolves around two main paradigms: in-context learning (ICL) and weight updates (e.g., fine-tuning, full SFT RL).

  • In-Context Learning (ICL): This involves providing examples or instructions within the model's context window, allowing it to adapt its behavior without changing its underlying weights. While ICL is computationally cheap and can lead to rapid improvements in low-data regimes, it has limitations. It often exhibits non-monotonic improvement, caps out quickly after a certain number of examples (e.g., 20-40), and is constrained by the model's context window length. Beyond this limit, performance can degrade.
  • Weight Updates: This involves modifying the model's parameters through methods like LoRA (Low-Rank Adaptation) or full Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This approach is more computationally intensive but allows for deeper, more permanent learning. For scenarios with abundant data, like self-driving cars, weight updates are considered the optimal strategy.

The challenge lies in finding the right balance. While ICL is effective for quick adaptation, consolidating these learnings into more robust, faster-executing skills often requires some form of weight update or "distillation" of experience. This is where the concept of a "harness" comes into play, allowing for the packaging of learned skills into reusable programs.

The Role of Wadd Labs and Robocurve

Wadd Labs focuses on building LLMs that control robots by developing a harness for effective control and collecting data to train better LLMs. Robocurve, on the other hand, specializes in evaluating physical AI, measuring the performance of various robotic embodiments and control approaches, including LLMs and vision-language-action models.

When controlling a robot, directly commanding it with an LLM like Astra can be slow due to latency. Wadd Labs' harness aims to address this by allowing the LLM to write code for repetitive tasks or to call pre-compiled skills, significantly speeding up execution and handling edge cases more effectively. This involves a blend of deterministic code for common actions and flexible VLM (Vision-Language Model) integration for handling variations or failures.

The Platonic Representation Hypothesis and Spatial Intelligence

The "Platonic Representation Hypothesis" suggests that as AI systems are trained on increasingly vast amounts of data, they converge on consistent, underlying representations of the world. This implies that powerful language models, by virtue of their extensive training, might develop representations similar to those needed for robotics, making them inherently capable of controlling robots.

Astra's remarkable spatial intelligence, evident in its ability to control Blender or perform complex robotic tasks, is likely due to its extensive pre-training on diverse data, including "computer use data." This might seem counterintuitive, but the argument is that interacting with graphical user interfaces, such as dragging cursors to orbit CAD objects in Blender, teaches the model about spatial reasoning, concepts like "top-down," "left," and "right," which are crucial for robot control. The historical design of GUIs to mimic the physical world inadvertently created a rich training ground for AI to learn about physical interaction.

The Future: General Purpose Robots

There is a growing consensus among frontier labs and robotics companies that general-purpose robots capable of performing tasks with human-like competence will emerge within the next two years. This would be a "ChatGPT moment" for robotics, where robots can generalize to unseen tasks and environments based on natural language instructions.

For companies like Wadd Labs, this future entails addressing challenges like latency by consolidating in-context learning into faster, reusable skills. The long-term vision involves a system that can learn quickly (via ICL), then "sleep" to compress these learnings into updated weights or refined skills, much like how biological intelligence consolidates memories. This "dream coder-esque" approach, where a growing library of skills is managed and refined, is central to building increasingly capable robotic agents.

The advancements in LLMs and their application to robotics are poised to bring about a transformative period, making robots far more versatile and integrated into daily life than ever before.

  Takeaways

  • The integration of large language models (LLMs) enables robots to execute complex tasks by generating code that serves as policies, allowing one‑shot learning of new behaviors without extensive robot‑specific data.
  • Early approaches like RT2 showed that fine‑tuning pre‑trained language models on web images and text can produce effector poses, proving that cross‑modal data dramatically improves robot control performance.
  • In‑context learning (ICL) lets LLMs adapt quickly using prompts but is limited by context size and plateaus after a few dozen examples, whereas weight‑update methods such as LoRA or full fine‑tuning provide deeper, more permanent skill acquisition.
  • Companies such as Wadd Labs and Robocurve are building harnesses that translate LLM‑generated code into fast, deterministic robot actions, mitigating latency and enabling scalable evaluation of physical AI across diverse embodiments.
  • The “Platonic Representation Hypothesis” suggests that massive pre‑training on varied data, including GUI interactions, gives LLMs innate spatial reasoning, positioning them to become general‑purpose robotic agents within the next two years.

Frequently Asked Questions

What is the 'code as policies' approach in robot control?

Code-as-policies means the LLM writes executable code—typically Python functions—that directly encode robot actions, turning the generated script into a control policy for tasks such as picking objects or moving to poses. This lets the model apply its coding knowledge to produce one‑shot solutions without additional robot‑specific training data.

How does the Platonic Representation Hypothesis explain LLMs' spatial intelligence for robotics?

The Platonic Representation Hypothesis posits that as LLMs are trained on ever larger, diverse datasets they converge toward a universal, abstract representation of the world, which incidentally captures spatial concepts needed for robotics. Consequently, models like Astra inherit spatial reasoning from pre‑training on GUI interactions, enabling them to manipulate 3‑D environments despite never being explicitly taught robotics.

Who is Y Combinator on YouTube?

Y Combinator is a YouTube channel that publishes videos on a range of topics. Browse more summaries from this channel below.

Does this page include the full transcript of the video?

Yes, the full transcript for this video is available on this page. Click 'Show transcript' in the sidebar to read it.

Helpful resources related to this video

If you want to practice or explore the concepts discussed in the video, these commonly used tools may help.

Links may be affiliate links. We only include resources that are genuinely relevant to the topic.

Full transcript is not shown on this page

This page focuses on the summary and original notes. For full verification, refer to the original YouTube video.

PDF