World's Top Researcher on AI, LLMs, and Robot Intelligence

Robotics is entering a new era where general-purpose foundation models—similar to how ChatGPT handles any language task—will enable any robot to perform any physical task in any environment. The key insight: rather than building specialized robots for specific jobs (like dishwashing), the path forwa

1h 15m
Invest Like The Best

Key Takeaway

Robotics is entering a new era where general-purpose foundation models—similar to how ChatGPT handles any language task—will enable any robot to perform any physical task in any environment. The key insight: rather than building specialized robots for specific jobs (like dishwashing), the path forward is training one intelligent system that understands physical interaction fundamentally. This approach leverages diverse data sources and prior knowledge to handle both routine tasks and unexpected situations with common sense, much like humans do.

Episode Overview

Sergey Levine, co-founder of Physical Intelligence, discusses the company's mission to develop foundation models that can control any robotic system to perform any task. The conversation explores why building general-purpose robotic intelligence is actually easier than creating narrow, task-specific robots—mirroring how language models evolved. Key topics include: the "scarecrow problem" (robots need brains, not just bodies), how vision-language-action models bring web-scale knowledge to physical tasks, the role of reinforcement learning in exceeding human performance, and the path to a "Cambrian explosion" of robotic applications once the intelligence platform exists.

Key Insights

General-Purpose Beats Specialized in Robotics

Just as language models became more effective by solving natural language in its full generality rather than targeting narrow tasks like translation, robotic intelligence should tackle physical interaction broadly. Training models across many tasks and robots builds fundamental understanding of physics, causality, and object interaction—making it faster to master new skills than starting from scratch for each application.

The Missing Piece: Common Sense from Language Models

Historically, robotic learning struggled with "long-tail" scenarios—unusual situations the robot never experienced. The breakthrough: multimodal language models contain vast world knowledge from web-scale training. By using chain-of-thought reasoning (the robot literally "thinks" about what it should do before acting), systems can apply semantic knowledge to novel physical situations, handling edge cases with common sense.

Data Efficiency Through Foundation Models

Unlike narrow systems requiring massive data collection for each new task, foundation models pre-trained on diverse robotic data need far less task-specific training. The model develops "physical understanding"—intuitive grasp of what will happen in unfamiliar situations—enabling rapid skill acquisition similar to how humans quickly master new physical tasks.

Moravec's Paradox and Machine Learning

Things easy for humans (picking up a cup) are hard for robots, while human-difficult tasks (calculus) are machine-easy. However, machine learning shifts this equation: domains where collecting data is straightforward become easier over time, even if physically intricate. The remaining challenges are tasks requiring multi-level reasoning and connecting physical skills to web knowledge.

Surprising Generalization Across Embodiments

The same model works across radically different robot bodies—multi-fingered hands, different degrees of freedom, various form factors—without being explicitly told what robot it's controlling. This suggests the core challenge is understanding physical interaction, not adapting to specific hardware configurations.

Notable Quotes

"Fundamentally, the goal of physical intelligence is to develop robotic foundation models that can control basically any embodied system to do any task."

— Sergey Levine

"We believe that doing it at the full level of generality might actually in the long run be easier than trying to special case very specific narrow application domains."

— Sergey Levine

"People can master new skills very very rapidly because we understand physical interaction—we can intuitively grasp what's going to happen in this new unfamiliar situation and let us bootstrap things really really quickly."

— Sergey Levine

"Effective robotic learning, effective generalization isn't actually the optimal way to have like a really exciting demo. The way to have a really exciting demo is to pick a really cool task, control everything else in the environment, and just make it work in that one setting."

— Sergey Levine

"I think something like that might happen in the world of robotics but it can't happen today because if you want to put together some cool new robotics application, you kind of have to build this monstrous stack and you need to basically solve the intelligence problem."

— Sergey Levine

Action Items

  • 1
    Prioritize Learning Over Immediate Results

    When approaching complex problems, focus on building fundamental understanding across diverse scenarios rather than optimizing for quick wins in narrow domains. Like foundation models, invest in broad knowledge that makes future challenges easier to tackle.

  • 2
    Leverage Existing Knowledge Reservoirs

    Before collecting massive task-specific data, identify and incorporate relevant knowledge from adjacent domains. Chain-of-thought approaches—explicitly reasoning through problems using prior knowledge—can handle novel situations more effectively than pure experience.

  • 3
    Design for Data Collection from Day One

    Build systems that can gather useful data while being deployed, even if initially imperfect. Like Tesla's fleet learning approach, create feedback loops where real-world usage continuously improves the model rather than waiting for complete training datasets.

  • 4
    Question Moravec's Paradox in Your Domain

    Identify what seems "obviously easy" but is actually hard, and vice versa. In any field, human cognitive biases about difficulty can mislead strategic decisions—systematically test assumptions about what will be challenging to automate or scale.

Full Transcript

Transcript of World's Top Researcher on AI, LLMs, and Robot Intelligence from Invest Like The Best. Auto-generated from episode audio; may contain minor errors.

My guest today is Sergey Lavine, one of the co-founders and researchers at physical intelligence. As a disclaimer, I'm an investor in physical intelligence because I believe it's one of the most important companies tackling the problem of robotics. As you hear us discuss today, robotics has what I would call a scarecrow problem. All of these amazing physical devices are becoming ever more possible in all sorts of cool permutations, but what they all really need is an intelligence, a brain, and that is what they're developing at physical intelligence.

are trying to develop foundation models that can make any physical robot do any task in any environment. That challenge is daunting and has required many of the world's best researchers, Sergey as one of those leaders, coming together to try to solve this problem. The nature of our conversation today is all of the problems facing robotics and all of the promise of solving these problems across the world. I hope you enjoy this great conversation with Sergey Lavine. This is going to be a real treat and a blast to learn about possibly the most exciting impactful area of technology being developed.

Just to set the stage before we go back in time, maybe you could just define physical intelligence as you see it. Fundamentally, the goal of physical intelligence is to develop robotic foundation models that can control basically any embodied system to do any task. But broadly speaking, you could imagine that in the same way that a language model is kind of rapidly evolving towards a system that can do any task that can be expressed in language, what we would like is to build a new class of models that can do any task that can be done by a physical actuated device.

actuated device. actuated device. And part of the thesis of this company is that we believe that doing it at the full level of generality might actually in the long run be easier than trying to special case very specific narrow application domains. Again, in much the same way that for language models, it turned out to be easier in some ways to solve natural language tasks in their full generality than to narrowly target like machine translation or sentiment analysis or whatever that may not be obvious why you would make that bet versus a robot that just does your dishes or something.

Um, so what are the key trade-offs to understand and why make the decision that you made? Maybe I I can give you like a two-part answer for this. First is how it relate kind of the analogy to language models and the second one is what that means in the robotics world. So uh and the first one is kind of a little bit more informed by evidence right uh in the world of natural language we saw that there were a lot of efforts to develop domain specific solutions that tackled specific problems like you know somebody would spend a lot of time thinking about how like English differs from French and then build a machine translation system.

The reason that language models actually uh kind of uh took over for all of those different application domains is because they can leverage much broader sources of data. And it's not even as simple as saying like, oh, we had this data for this application, this data for this application to like merge everything. But it it's actually more than that. It's when you can leverage weekly label data like data that you like, you know, in the case of language models, you just mine from the web, you actually learn more about the world.

So you establish like foundation of world understanding and then on top of that foundation turns out to be much more effective to build out different applications. So uh to bring this into robotics obviously the the calculus doesn't look quite the same because in in robotics we don't have like a internet size data set that we can just draw on. But this notion of understanding the world if anything is actually uh more important in robotics because if you have many different tasks maybe even many different physical systems then you can go from training individual like you know uh dishes dishwashing specialists or laundry folding specialists and instead train a model that actually understands physical interaction right like people can master new skills very very rapidly uh because we understand physical interaction we can you know intuitively grasp like what's going to happen in this new unfamiliar situation and let us like bootstrap things really really quickly.

So if we can draw on data from many sources, many applications, many robots, then we can have a model that has a physical understanding and then it'll be much much easier to put new applications on top of that platform. What is the hardest part about building in this way for you for you for you when you get when you see other approaches that are more maybe legible to the average person? Oh, there's a robot moving around doing this one specific thing. it looks a certain way.

Yeah. Yeah. Yeah. What's what's the hardest part about this approach as you're doing it? I think this has actually been kind of an issue in my whole career because when you work on robotic learning and the more general the more that this becomes important is important is important is effective robotic learning, effective generalization isn't actually the optimal way to have like a really exciting demo. Like the way to have a really exciting demo is to pick a really cool task, control everything else in the environment.

like set it up so that it's perfectly clean, perfectly, you know, pristine and just make it work in that one setting. Like that's the way you make a robot demo. And generalization, generalization, generalization, you can't just show it in like in one spot, right? Like the point of generalization is that it does something relatively mundane that any human could do, but it does it in any situation. So, we had uh some demos that we released last uh April where we showed our robot cleaning kitchens. And like you know I think it's kind of cool but if you watch an individual video out of context it's just like okay it's like picking up plates like anybody can pick up plates except that we just put it into that home just for that demo and it never had training data from that uh setting.

So obviously you kind of have to like understand what's going on to appreciate why this is actually pushing the frontier. the frontier. the frontier. What is your model for the stakes of what you're doing? Like if you are successful, successful, successful, yeah, yeah, yeah, I'm curious for you to define what that would mean successful other than just like we we cross this chasm of like general physical intelligence or something. But if you cross that line, then what? then what? then what? One of the things that I think would be really really exciting uh that would be enabled by a general purpose embodied foundation model is the ability to unlock people's imagination in how they build robots and other embodied systems.

So like personal computers were a really big deal in my mind because it made it possible for lots of people to hack together all sorts of like really cool stuff and there was this like kind of Cambrian explosion of like amazing applications that started in like the '9s and so on uh and then was further accelerated by the internet and I think something like that might happen in the world of robotics but it can't happen today because if you want to put together some cool new robotics application some cool new robotics idea you kind of have to build this like monstrous stack and you need to basically solve the intelligence problem.

problem. problem. But if there is a solution that someone can build on top of, there's a a foundation model that you can prompt that'll provide like basic functionality and then you can maybe fine-tune it a little bit or adjust it in some way to your application. Now, it actually makes it a lot more attractable for lots of people, lots of companies, lots of individuals to try out all sorts of different things. different things. different things. And I think, you know, we sometimes we think that like robots are going to be like one thing like just like, you know, yeah, there's there's people and now we're going to make like metal people and that'll be like robots.

But I don't think that's how it's going to be because like no technology has been like that. It's going to be more like kind of a toolkit where you can put together all sorts of like really cool applications. Get really creative with it. You know, maybe I'm going to make a robot with like five arms and this one's going to hang from the ceiling and figure out kind of the right thing to tackle your domain. Maybe also experiment with software. But you need the right platform on top of which to do that.

And I think the foundation model can be that thing. that thing. that thing. What are the in your mind the pros and cons of the humanoid approach to robotics? robotics? robotics? One pro of it is that it's really cool, right? Like you can show it to somebody and they're like, "Yeah, I get it." Like Yeah. You hear a lot about the Optimus hand. hand. hand. Yeah. But but it is cool. I I don't like I think there's there's a lot of value to that. There's a lot of value to capturing the imagination and there's a lot of value to like getting people to think about what the future might look like in, you know, in a way that's understandable.

Um but I think it's, you know, in my mind, it's one of many possible kinds of robots that we're likely to have. And I fundamentally the intell the challenge of intelligence looks very similar for all these different robots. And uh I don't think we should be tackling intelligence in the context of one specific body. I think we should handle it in a general way because otherwise it's just really hard to like to get a handle on this. We need lots of data. The cool thing about being able to build robots is that ultimately they don't have to be constrained to look like humans at all.

Right? Like you can build the right tool for the job. You could imagine that you're building a house with a robot that is a swarm of 10,000 quadcopters. Uh, and I think that in the future we'll have a robotic foundation model which can then be adapted to all sorts of applications and and it might really r run the gamut from like, you know, like bulldozers or something to humanoids to robotic arms like this thing. Uh, and maybe it would need to be adapted to each one.

Maybe it would need to be fine-tuned. Maybe we would need something in context to understand how that body works. But the fundamentals of how you interact with objects, how things move in the world, how causality works, like that's all conserved for all of these different systems. Do you have a favorite example of what might be possible with true general intelligence that might not be possible with say like a humanoid only intelligence or something? There are a few things that that I think are are worth thinking about.

Uh, one is that we can make machines that are very big and machines that are very small. Uh I think that in the long run, this is not by any means a short-term thing, but in the long run, I think there's lots of really exciting applications in medicine, in surgery, where we not only might in the long run not be limited to robots that look like humans, we might not be limited to robots that can even be controlled by humans, right? Because currently, for example, in robotic surgery, uh it's done entirely through tele operations.

So you need something that a person can control in real time with the right level of dexterity. And of course that limitation holds for current learning enabled systems too. But in the long run we could imagine addressing that. addressing that. addressing that. If you think about the most important hash marks on the timeline of robotics research that have gotten us to here. I always think it's super helpful to set the historical context before we talk about like what state is today and where we're going. What could you walk us through that like the the relevant hashmarks on the history timeline?

At some level doing endto-end control for robotic systems is a very very old idea, right? Like the first for example autonomous driving systems that used end-to-end learning they existed in the 1980s right like Alvin was uh I think it's 1986 or 87 and that was a driving system that was demonstrated to drive on highways controlled by a neural network and then from a camera the neural network was like tiny right so I I I think that there are some very venerable concepts but historically what has been really difficult in robotic learning is that you need a system that that handles the application you want to address that is cost effective to train for that application meaning that you don't need like a huge amount of data for every single application you want to tackle handles longtail scenarios with common sense so if something weird takes place in the world it needs to like have a reasonable response to it and then also for the thing that it's actually supposed to do it needs to be robust fast and reliable and getting all those things together is very very hard because with machine learning like it works best when there's a lot of data so if you sort of naively approach a robotic problem and say like I want to do like you know washing dishes, right?

Uh the obvious thing to do is to collect like an enormous amount of data of washing dishes, but that's not cost- effective because then you go on to the next application and you go through that process all over again. So being able to train general purpose models that can handle many tasks is essential to this because now you need a lot less data for each new task. But then even further, and this is the thing that has probably changed the most in the last few years, you also then need to handle the unusual scenarios.

And for the unusual scenarios, you are probably not going to have experience. What you need to rely on is knowledge that you've acquired from other sources that you can ground in that new situation. And people are extremely good at this. So if if you're driving a car and there's like, you know, something going on in the in the middle of the road and someone put up a sign saying like don't go here, like, you know, there's the gas leak or something. You've probably never experienced that before, but you can kind of like put these things together and figure out what you're supposed to do in that unusual situation because you have common sense.

And this has been like a huge mystery in um in robotic learning world. Like where do you get that common sense? And this is what's changed in the last few years because turns out that multimodal language models are really good at pulling in knowledge and um trying to articulate that knowledge. They're not very good at like grounding that knowledge in physical situations, but they know stuff. So now it actually that there is a path to get that kind of common sense by essentially leveraging the knowledge that is contained in multimodal LLMs but there's also a challenge because you have to somehow plug into that knowledge in the right way way way right you can't just like show it a picture and say what would you do here because it doesn't have the context like it doesn't know that you are a robot this is what you look like this is what's going on so that's a technological challenge and uh you know we've made some headway in addressing that technological challenge the research community in general has but most important it's kind of that light at the end of the tunnel of like okay Now we have this way of pulling in lots of knowledge which can help us handle those longtail scenarios.

Are there hashmark equivalents on the timeline of like AlexNet or the transformer? Are there big like major events that you think everyone will point to when writing the history books about this? about this? about this? Uh that's a good question. I think it's very early on right now um to like answer that definitively. I mean certainly uh I would say that there's like like like you know kind of have to look back like at least 10 years to have something like that but you know probably the first uh end to-end learning systems which were in the 80s that's definitely a milestone.

I think that the first deep reinforcement learning systems which were in the early 2010s like those are probably a milestone because deep reinforcement learning gives us a way to go beyond human level performance which I think will be essential for robotic systems. And then there's the more recent stuff but you know that's like in the last few years. I don't know how that's going to shake out as far as like whether that's something that people will point to, but I do think that the advent of um multimodal LLMs that can be adapted to robotic control to bring in that common sense.

I do think that's a really important advance. Yeah, Yeah, Yeah, I think we're probably going to see quite a few important advances in the next few years and you know, maybe those will be the things people point to. C can you tell us your own personal history of approaching the problem? So yeah, maybe the origin of when you first became interested and why and then how you've managed how you've decided what to spend your personal time and attention on ever since then. So I started working in robotics in um 2014 and that was actually after I finished my um graduate degree and uh started a posttock with professor Peter reveal at at UC Berkeley.

I actually hadn't worked on on robots before, but I figured I I I should get a little bit more education after finishing my degree. And uh his lab worked on robots, so I I tried to apply what I had learned previously to robotics. Before that, I worked on computer graphics actually. Um, and um, I think that the the the thing that I've always wanted to really figure out is how to get AI systems that get better and better the more they do things because I think that's tremendously powerful.

If you can have a system that gets better and better the more it does something and it just keeps getting better, then there's sort of like no limit. Then it can master all the skills you wanted to do. And initially I tried to approach it in a very kind of a blank slate way meaning like you start with nothing you practice a particular skill and you and you get better at that skill and that was like kind of okay like you can you can do that in a limited setting and you get something that works but it's very hard to turn that into a general system that can work in open world settings because if I practice something over here and then it goes over there now something something is different and it needs to practice all over again.

The next thing I I tried and this is I did this when I worked at um at Google uh afterwards I tried to see if we can do that but now paralyze it across many robots. So collective learning can you put like 20 robots in a room and have them all learn together and that works um and it generalizes but it's very hard for that to handle these like tail cases these edge cases right because now it becomes this kind of like uh kind of uh uh soant of this particular task and that's all it knows in the world.

So the what I think is is the next step is what I mentioned before is combining this ability to practice skills with lots of prior knowledge. And that's actually a really really hard problem. You know, it's not just in robotics where it's a hard problem. I think it's a hard problem in all of AI because arguably the two big impressive results in AI over the last few decades have been generative AI and deep reinforcement learning. Like if you want to like a single example to epitomize this generative AI that's like LLM's deep reinforcement learning, AlphaGo, right?

They're both very very impressive and they're very impressive for very different reasons. The generative AI is impressive because it can reproduce some of the things that humans can do like it can draw pictures that look like you know human pictures uh write text. DRL is impressive for the opposite reason. It does things that humans hadn't thought of like it move 37. Exactly. So I think the the big challenge and this is kind of what I'm leading up to and what I what I hope to uh uh that we'll figure out here at physical intelligence is has to combine those threads how to bring in all of that knowledge that you get with generative AI but also go beyond just human level performance with reinforcement learning reinforcement learning reinforcement learning and that I I haven't figured out yet but I think we've made some good progress on that.

that. that. So what literally have you done and are you doing to to make that happen? So in the past few years we started off first by uh developing the basic foundations. So uh the basic foundation is what's called a vision language action model. A vision language action model you can think of as an LLM that has been adapted for robotic control. So the way these things are trained is they're first trained on text data. Then they're adapted with lots of image data from the web to understand images and then they're adapted to robots with lots of very diverse robot data.

Now that that's a starting point. That's a way to take all of that web knowledge, get it into a model that can control robots um and get some interesting behaviors out of it. And then from there, we studied uh two threads. How to get this thing to handle uh unusual situations with common sense and how to get it to improve with reinforcement learning. So the way you get common sense is by essentially using chain of thought. So the robot enters a scene uh and instead of directly starting to move, it thinks about what it was asked to do.

So if it was told clean up the kitchen, looks at the scene and says like okay based on this I should pick up the the plate. Like it literally talks to itself. It says pick up the plate and then it goes and does it. So that unlocks all this prior knowledge because those intermediate inferences benefit from the webcale pre-train. So that's good that that handles uh edge cases and then the reinforcement learning part comes in after you've practiced it a few times and you can keep getting better and better at the task directly through your experience.

And for example, we had this demo on uh making espresso. That system practiced making those espressos many many times and use that to improve robustness, improve speed, improve throughput. And we're not done with that. Like I think there's a lot more to do there, but we have kind of like the the starting point. Most software companies try to maximize your time on their app to juice engagement. Ramp does the exact opposite. RAMP understands that no one wants to spend hours chasing receipts, reviewing expense reports, and checking for policy violations.

So, they built their tools to give that time back using AI to automate 85% of expense reviews with 99% accuracy. And since RAMP saves companies 5%, it's no wonder that Shopify runs on RAM, Stripe runs on RAM, and my business does, too. To see what happens when you eliminate the busy work, check out ramp.com/invest. Every investor should know about RO because ROGAI's platform is not just another generic chatbot. Instead, it was designed to support how Wall Street bankers and investors actually work. From sourcing, diligence, and modeling to turning analysis into deliverables.

For me, three key things differentiate Robo. First, it connects directly to your systems, so it can work with your actual data. Second, it understands your workflows, how work really happens across a deal or an investment. And third, it runs endtoend and produces real outputs the way the best people do, auditable spreadsheets, investment memos, diligence materials, and slide decks that match your standards. This all comes from the fact that ROGO is built by finance professionals for finance professionals and it's already being adopted by some of the most demanding institutions in the world.

To learn more, visit rogo.ai/invest. OpenAI, Cursor, Anthropic, Perplexity, and Verscell all have something in common. They all use work OS. And here's why. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, skim, arbback, and audit logs. That's where work OS comes in. Instead of spending months building these mission critical capabilities yourself, you can just use work OS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on work OS.

Work OS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit works.com to get started. Sorry to be think about it, but the robot data itself is the right way to think about it. I'm looking at the, you know, the gen one of these things. I see a camera here. Maybe there's some sensors somewhere else. is effectively the data being gathered by various sensors strategically placed on the robot at different parts. different parts. different parts. Yeah. And um something I'll say about sensors is that I think you can actually get away with less than one might think and still do quite a lot.

So this platform here has three cameras, one on each wrist and a base camera. It doesn't have touch sensing. It doesn't have force sensing. It's very bare bones and very low cost. Um I'm sure that more sensors could make it better, but a good learning method can actually like compensate for deficient sensing fairly well. Like the wrist cameras are essentially a touch sensor in disguise because you can see local deformations when you touch something. If I think about the analogy to like the expert systems of the 80s and 90s in in basic AI to the lesson that scale is all you need and the sort of counterintuitive nature of that that you're not teaching it any specific thing just like blasting it with data and there's this reservoir of internet data.

data. data. Talk about the reservoir how to create the reservoir of data needed for this. Yeah. So, I don't think anybody really knows how much robot data is needed to have truly generalizable and powerful embodied AI, but my sense is that we actually don't need to know. What we need to do is get to the point where these systems are useful enough that they can go out into the world and gather more data themselves. Yeah. Yeah. Yeah. Like basically, to put it bluntly, Tesla doesn't worry about how much data their cars can collect, right?

If anything, it's the other way around. That's a little too much data, right? So I think that the key is not so much to quantify like here is exactly the the price tag of getting the ultimate robot data set. The key is to get a system that can go into the world that's useful enough that does a wide variety of different things and they can keep pulling in more data. You brought up the example of Tesla, the the beautiful system of a thing that's useful without the AI to begin with because the human drives it and it gathers data.

Why then not start with your best guess at something that's useful as a single robot to have the same sort of flywheel thing happen? I think it's a good idea. Yeah. Yeah. Yeah. And and do you think that's an approach that you'll pursue? I don't think that there's like one right answer, right? So I think there are some domains where deploying a system under human control makes a lot of sense. There are some domains where deploying a partially autonomous system uh is very reasonable. It it's kind of domain dependent, right?

Because robots aren't just one thing. Yeah. Yeah, like maybe some some people might not want a robot in their home that is constantly being controlled by a person offsite, but maybe for some applications that doesn't matter. Yeah. Yeah. Yeah. What has been so if you if you mark the start of physical intelligence through today, what has been the most surprising thing to you that you've discovered or the nature of how the research has gone? One of the things that's been surprising to me is that I think we've made a lot more progress on dexterity than I thought we would.

Um, so like I kind of had an expectation based on my prior work that you know generalization being able to handle all sorts of different scenes, all sorts of different objects like we had good reason to believe that if we just like collect more and more data that just steadily gets better. But what was surprising is that we could also get these systems to perform very dextrous behaviors dextrous behaviors dextrous behaviors without really doing anything particularly special for that. The same, by the way, also applied to getting systems to work on different embodiment where we could get our models to work on all sorts of other robots, including robots with multi-fingered hands, uh robots with different numbers of degrees of freedom.

And obviously, we needed to get data and we needed to fine-tune the model, but the model itself didn't need to change. It didn't it didn't even need to be told uh through any kind of prompt what the robot was. And that was also surprising to me because I would have thought that we would need some fancy techniques to sort of like adapt the system to faster, more dextrous, more complex tasks and also to different kinds of embodiment. But it actually seems to generalize pretty well across those.

those. those. I'm always interested in like the spectrum of capabilities and especially where the systems today are more advanced than you think people would probably expect and where they're less advanced than people might expect. This is something that's always been very um tricky to understand in robotics. Um, you know, there's there's this idea that roboticists always talk about called Morovix's paradox. Um, it's actually true in all areas of AI, but especially in robotics, this is a big deal. Um, we kind of have a cognitive bias to think that things that are easy for us will be easy for the machine, right?

Like, so solving calculus problems is difficult for most people, but uh, picking up a cup is easy for most people. So we think, oh, machines should be able to do this, but it's actually the other way around that there are things that are easy for us because they have to be, otherwise we wouldn't survive. So we're like, you know, we're very good at like spotting the tiger in the jungle because the people that weren't so good at it got eaten by the tiger and they're not around anymore.

So um because of that, we have this cognitive bias and we think that there are uh things that should be very easy, but they're actually very difficult engineering challenges. However, something that is changing is that machine learning slightly changes that equation. So programming something by hand to pick up any cup anywhere, that's difficult. Getting a machine learning system to do it, if you have data for it, it's actually not that difficult. And I think increasingly what we'll see is um a shift where domains where collecting data is straightforward.

They actually end up falling into the easy bucket over time, even if they are physically intricate. But there will be domains where collecting data is difficult, where you need to use more common sense, where you need to reason at multiple levels of abstraction, connect physical skills that you've learned in other areas to knowledge that you got from the web. And those will will be tough and that that's where we'll we'll need more uh technology advances. advances. advances. What is the science of common sense? You mentioned common sense before like when we say that.

Yeah. Yeah. Yeah. What does that mean? For the purpose of robotic learning, uh we can think of it as as as applying semantic inferences using knowledge learned from other domains to a to the current physical task at hand. Right? So you can think of common sense as sort of the opposite of muscle memory. So muscle memory like if you play a sport, you practice something a lot, you hardly think about it. You just kind of do it on autopilot. Common sense I in my mind I don't know if this is the conventional definition but I think it's a reasonable definition is when you know something to be true because you saw it or you read about it or you heard it and now you are in a situation where that fact is highly pertinent to what you need to do and you are able to make that connection apply it to your situation grounded in the environment that you're in and make the right decision.

right decision. right decision. One of the other differences that's so interesting to me is people that have used chat, everyone's used chatbot now. You query it, you get an answer, query it, get an answer. We're now seeing what happens with cloud code and other things where you give it something complicated and it's able to do a very long, you know, the measure of how long it can go without failing. What's the similar thing long range thing in robotics? It's something that we're working on quite a bit right now.

And uh in fact, the methodology is not that different at some level. So the way that our models work now as I mentioned is they use this kind of chain of thought process to reason about the task. And when you have that you can actually do very long horizon tasks. Um you can have a robot that goes and you know takes out all the dishes from the dishwasher, puts them in the correct cabinets, wipes down the counter, all that kind of stuff. The interesting thing here is that we found maybe about 6 months ago that our models had gotten to the point where they could be improved just from supervising them with highle instructions.

So what does that mean? You you take a robot, you put it in a new kitchen, you ask it to clean the kitchen, it gets to work and then it fails somewhere. So now okay, what do you do? Well, you you add more data. Traditionally what you what we would do in that situation is add more teleyoperation data to cover a wider range of kitchens. But what we tried kind of on a whim is to see okay well what if we don't add more telly operation data.

What if we just add more data labeled with the semantic command. So basically just take whatever the robot experienced and just you know label it with some semantic commands but don't add any more low-level actions. And that actually helps that actually improves it ability to generalize. So what that means is that the bottleneck had actually shifted from the lowest level meaning the the robot's ability to physically do the task to this like middle level where now the system is more bottlenecked by its ability to interpret the scene and select the correct next step which can be supervised with language.

And that's a big deal because now that means that someone can literally talk to the robot coaching basically. Yeah. Exactly. And make it better just by talking to it. If we're in 2050 and there's no robot in my kitchen doing my dishes for me, what do you think the most likely explanation is for it not having gotten there by that point? My suspicion is that um there is a a long tale of challenges that has to do with the interaction of technology and and people like you know in some ways autonomous cars aren't that different in this regard where the getting to a level of comfort with deploying autonomous vehicles on the road was you know that that was a significant challenge that ran in parallel with getting the technology to that level.

And you know for example like the early Tesla self-driving was like a bit controversial because it wasn't perfect and there was a question like are people comfortable with this level of imperfection. So probably there are some tasks for robots uh where people will be comfortable with something that's not perfect something that needs to learn from its mistakes. There are some areas where people will not be comfortable. Are you comfortable with you know occasionally breaking your dishes? Maybe in a few years it will stop breaking those dishes but maybe in the meantime it's not quite there.

Are you comfortable with a robot like that um in a home where there's like small children? Maybe not. That's okay, right? So, I think that figuring out how those factors interact and what that means for um the timeline and for how these systems get better with experience. I think that's a tricky question and I think it needs to be approached very carefully with a lot of sensitivity. There may be some domains where it makes a lot more sense for these systems to be deployed and bootstrap and collect more data and maybe other domains require more care.

Could you imagine a purely technical explanation for why something might not work? I think the place where I would see the biggest technical risk is um dealing with the breadth of different situations. So I think if we were talking about a well-defined but you know but slightly chaotic environment like cleaning hotel rooms or working uh you know assisting human uh cooks in a in a restaurant, right? Like that that sort of thing. I think I have like a very good sense for how to get that under control.

If you're imagining a robot going into a home, you know, one place where I can anticipate a challenge is that there are a lot of other unexpected things that can happen and you need a system that's very good at inferring what's going on and adapting to it or reacting intelligently. And I think we have a lot of ideas for how we can approach it. But that is kind of the hardest part of the problem because when you're in a situation where just about anything could happen and uh you're controlling like a physical device that you know affects the world around it.

Uh then you really need to uh get things right at least at some level pretty much in every case. Like it doesn't mean that you always have to succeed but it doesn't mean that you always have to do something kind of sensible that people are okay with and I think there are a lot of really good ideas for how to do that but that is probably the most challenging part of the equation. If I go back to thinking about uh the right model to think about the physical intelligence approach to doing this whole exercise, help me like make it as simple as possible.

So one might be we're going to build a whole variety of different kinds of things to different kinds of form factors to do a whole variety of different kinds of things and mash all this data together and start to you know experiment with how we can you know on evals make it better. better. better. Is that just the simplest way of doing it? Is there an even simpler way? And I'm asking because I'd love to then contrast it with some other approaches that you're interested in that you're not doing that others are doing.

Uh I see. So in my mind the most important thing to get right is um to get the system to be general and in particular to get it to be general with respect to how it can be improved. Right? So you know for example handdesigned robotic controllers are not very general aspect of how they can be improved because it requires like a human engineer to go in and improve it. A learning based system a learning based perception system is more general because all it requires is human labelers to go in and label more data.

A system that learns autonomously from data that it gathers through its own experience is even more general because you don't even need the human labelers. So the key is this generality particularly with respect to improvement. Um, and the decisions we make are to a very large extent centered around that. around that. around that. So, I don't know if the the correct design for a robot is to have three cameras. I don't know if it needs like a touch sensor. I think we're very agnostic to that.

I think we'll try a lot of those uh different uh choices. I'm not even sure if in the long run it's going to have a language model. Maybe it'll have some other kind of model that's trained on very diverse data. Uh but the key is this level of generality. generality. generality. What other approaches are the most interesting to you? I think one thing that's like a very important question in this area and something that I think the research community and the tech community has not fully answered is the dichotomy between different data sources particularly with respect to real data and simulation.

It's a very controversial topic. I have a very strong opinion about it. But I think that it's worth acknowledging that if we look for example at humanoids, if you you know if you've seen videos of humanoids doing all these acrobatics, right? There's a particular pipeline that makes that work which is very heavily reliant on simulation and actually very light on real world data often almost often actually zero real world data. And then there are the approaches that work well for robotic manipulation that often are the opposite.

They often use very little simulated data, often use large amounts of real world data and very large foundation models. And it is kind of surprising that in these two robotic domains, the dominant approaches look so different. Now, it may be that one will win out and there's a particular approach that can handle everything in the long run or maybe there's some sort of synthesis of these ideas uh that's important. I I don't know the answer and I have I have my own subjective opinions. I think the approach we're taking is a very good one, but I think that it's interesting to look at that and see why is it that these things are so different.

so different. so different. Can you talk about the contrast between cool and useful? Like the Boston Dynamics robot is very cool. Like the the backflip is super cool. The I don't know what I need that requires a robot do a backflip. Um so I'm curious how you think about optimizing around cool versus useful. I think the strategy we've taken uh I don't know if it's it's like the right strategy but uh the strategy we've taken is subject to the constraint that it's useful make it as cool as possible uh and uh that that's kind of reflected in in you know in our blog posts in our videos is that we make decisions first and foremost based on uh our assessment of what will drive the tech forward towards this truly general broadly applicable uh robotic foundation model.

But in doing that, we try to stress test it against the kind of the toughest challenges we can throw at it. And that's, you know, the toughest challenges are the ones that look cool. So, we didn't set out, for example, to build a robot that uh can make espresso or can uh, you know, fold uh laundry, but in the process of building these general systems, we figure like these would be particularly challenging, particularly exciting things to try with them to see how far we can push them.

Can you talk about the robot Olympics? Yeah. Yeah. So, um, there was, um, uh, a gentleman named Benji Hson who used to work at, um, Everyday Robots, part of Alphaba before it dissolved, and he spents a lot of time thinking about tasks that robots could do. So, he wrote a really interesting blog post, uh, a while back, uh, where he basically said like, hey, um, there was this like, uh, robot Olympics uh, that was held in China where robots would like run around on a on a track and jump and so on, but maybe these aren't like the real challenges you should worry about.

How about a robot Olympics centered around essentially everyday tasks that people do that's kind of more of a paradox thing where tasks that people find really easy but that robots struggle with. And he had things like, you know, opening a door, uh washing a frying pan uh with with grease on it, using a plastic bag to pick up dog poop, right? That's like things that people don't find particularly challenging, but that like no current robotic system can do. And he listed, you know, maybe a dozen of these things.

Uh and we wanted to give this a shot. We weren't this wasn't actually like part of a concerted like research project. It was more like you know we had developed processes and systems for just like ingesting new tasks that we wanted to use for all sorts of tasks and we figured okay like a good way to test this is to say like hey here's like a big list of tasks let's just go through this process that we've developed and see if it works basically. So it's almost like a test of like our internal operations and model training system.

Then we tried these things and actually turned out that we could solve almost all of them. We didn't get um there's one we couldn't do which was turning a dress shirt inside out because the grippers on this thing wouldn't fit inside the sleeve. So, we probably need to change the gripper. Uh and I think on a technicality we didn't succeed at peeling an orange because he said do it with the fingers and our fingers weren't strong enough. So, we had to use like a little tool like a little um knife basically.

Uh but every everything else we could do. Uh, and what was actually interesting to me is I mean obviously it's cool like the videos are nice but in you know if anybody watches those videos one thing that I think is important to keep in mind is we didn't like develop anything special for this. We literally use this as a test of our like task on boarding process. And I think that's that's like there's something interesting there because it suggests the power of generality that when you have this this kind of general system you can really just like on board all these crazy tasks without really doing anything particularly sophisticated.

I was curious before when you said superhuman ability like on dexterity or something like that where we're limited by what we can do or maybe by what we can control even if it's gets gets smaller. What are some of the other dimensions like that we might surpass human ability on in terms of like physical you know physical ability? What are the other uh uh trend lines that are most interesting to you? So here's a fun one. We were uh working on a task where a robot had to plug in um cables like things like power cables or Ethernet cables or something like that.

And when a person does this uh I mean obviously if you practice it a lot you'll get really good at it. But when a person does this without having practiced a lot you pause frequently right because it's it's not a physical thing. You just have you have to cognitively process what's going on. You have to make sure that it's like all aligned and all that stuff. So you do it very slowly. And if you're teleoperating a robot you do it even more slowly because there's this level of indirection.

it turns out to be like pretty straightforward to go in and like find all those pauses and remove them. Uh, and you can speed things up further. So you can get to a task where a person demonstrates what it means to succeed and then you can have the robot practice the task and succeed in the same way but a lot more quickly, a lot more efficiently. And uh the most general way to do this is with reinforcement learning, but there are also like some simple tricks you can do that to if you just want speed.

So that's like one example of something where uh you can have a machine that does it a lot better and you know at some level you have like a processing bottleneck like that's why the person does it slowly because they have to process what's going on but speeding up processing is something that people understand quite well in computer science. There was this amazing Michael Kiteon novel called Prey where there's a question about form factor where it seems like for a given problem there there may be an optimal or set of optimal shapes of the robot to perform the task and that what you should do is analyze the problem then have something that can almost like morph or transform into the right form factor.

How do you think about that the innovation on the form factor side rather than the data and model side? I think that in general in robotics the ability to innovate on form factors has been very constrained because of the AI challenge right so uh if you have a a traditional um AI pipeline like you know you're doing some motion planning and stuff like that it's hard to just like go and cobble together some new robot because when you do that you have to like characterize the dynamics of the system you have to do CIS ID you have to build up all this stuff if you could just put together a robot in your garage load up a robotic foundation model and tell it to do a bunch of stuff.

Like maybe it won't be perfect at it. Maybe it needs more data to really perfect it, but you can at least like get the thing moving. moving. moving. I think that can be a really powerful engine just to get everybody to experiment with this stuff. So I don't I don't think that like I'm like the right person to design the perfect robot. There are people here of course who are a lot better at that, but in general I think that it's just like with personal computers that I think the key is to let people experiment and play around with it and just radically lower the barrier to entry for that.

And I think then we'll see a lot of a lot more creativity like you know when uh when we first when people first started using computers uh personal computers there was a limited number of form factors. Now you can have a computer in your phone a computer in your car and embed a computer in your refrigerator. They're everywhere and they're very different and generality good software good foundation on top of which you can build applications. Those are key to enabling that. that. that. Your co-founder Locky once described to me the feeling of physical intelligence for a human is like learning how to ride a bike.

Like there's that moment when like you didn't know how to do it and then you do know how to do it and that feeling is physical intelligence that like snap of understanding. You know there's actually a physiological explanation for this. There were studies that were done in monkeys using tools and you can actually find where in the brain uh like which neurons activate for the monkey to figure out where its hand is. It turns out that if it's using a tool they activate based on the location of the tool tip, not based on location of the hand.

M so like the tool being an extension of your body is like that is a real physiological thing like your brain literally does that. So what does that knowing that what does that do to impact the approach to your research? research? research? Well I I think to me it says that physical intelligence should be uh at at some level agnostic to embodiment that that that a good foundation model should figure out how to manipulate whatever body uh it's controlling whatever tools it has at hand. That there's basically one problem.

not many different problems. There wasn't like a humanoid problem and a car problem and a bulldozer problem and a robot bolted to the table problem. There was one problem and if you solve it at full level of generality, that's really, really powerful. powerful. powerful. As your business scales up, everything gets more complex, especially your compliance and security needs. With so many tools offering band-aids and patches, it's unfortunately far too easy for something to slip through the cracks. Fortunately, Vanta is a powerful tool designed to simplify and automate your security work and deliver a single source of truth for compliance and risk.

There's a reason that Ramp, Cursor, and Snowflake all use Vanta. It frees them to focus on building amazing differentiated products, knowing that compliance and security are under control. Learn more at vanta.com/invest. I know firsthand how complex the tech stack is for asset management firms. And seemingly every new tool and data source makes the problem even worse, adding more complexity, more headcount, and more risk. Ridgeline offers a better way forward. One unified platform that automates away the complexity across portfolio accounting, reconciliation, reporting, trading, compliance, and more, all at scale.

Ridgeline is revolutionizing investment management, helping ambitious firms scale faster, operate smarter, and stay ahead of the curve. See what Ridgeline can unlock for your firm. Schedule a demo at ridgeline.ai. We're start we're in the early stages of seeing some of the job and other sorts of transformation in businesses in the economy etc that LLMs make possible. Uh certainly we've seen it in in engineering. How do you think about what might happen or what you hope will happen when we're at a similar stage whenever that happens to be for robotics where all of a sudden we have this thing that's general that's useful like where do you think the world's very efficient at deploying these things?

People are creative. Where do you expect to see the world start to change most in the early days post physical? Yeah. Yeah, that's a really interesting question. Um, I tr I I really don't know, right? Like I don't think anybody would have been able to predict how the LM stuff evolves and people would have guessed, but this is why I keep coming back to this idea that maybe the key is to let people try lots of things. Like one of the really amazing things about applications of LLMs is that they are like really accessible and uh somebody could put together a really cool new prototype that under the hood is just like prompting like you know chat GPT or something but they can experiment with it they can try it out see what it does and there's an amazing kind of a power to having lots of people lots of smart people rapidly iterating and prototyping lots of things um and and that's why I you know that's that's a lot of why physical intelligence has really put a premium on uh on engagement like we've open sourced uh our models.

Uh we would like to uh engage with lots of other companies that are uh building robots because uh we all see a lot of power in this effect of having many people trying out lots of things. things. things. What are the major controversies in the robotics community? Obviously I'm an academic so uh to me a controversy is like someone gets in in an argument with me at a conference but but I can tell you that the kind of arguments that I that I found myself in and it's kind of an interesting trajectory that in the early days the main argument I would have with people is does learning have a place in robotic AI and and and I think part of why that was often a controversial point is that in a traditional engineering pipeline Robots do look very different than software artifacts.

Like, you know, they're they're physical that that that they can affect stuff around them. They can, you know, there are safety considerations. There are a lot of weird situations they can get into. And it's it it took a really long time for the robotics research community to really internalize that you don't necessarily need to program in things like knowledge of physics. Like you don't necessarily need a physics simulator inside your robot when it's planning, but you can actually have a learning system figure all that stuff out.

And that that was a very controversial thing for a very long time. I think at this point there's a lot of acceptance that learning is a really important part of robotics. But I don't think there's still universal acceptance that endto-end learning is the right way to go. That basically I don't think there's universal acceptance of the bitter lesson. The bitter lesson says that you should not program the machine to think the way you think it should think, but you should let it learn from data. And that is not a universally accepted idea.

I I think there's good arguments against it, but I think that in the long run, if we want that generality, especially generality in the machine's ability to improve, then we need it to primarily be learning from data. What is the good argument against? My best attempt at steel manning this is that if you want something reliable in a really complicated open world setting, then you can't afford not to use what you already know about the physical world. And we've got like textbooks full of this stuff, so why don't we just like plug in what we know from the textbooks?

What is compositional learning? Can you describe that? There's an example I can give you uh which maybe is like uh the more vivid way to communicate. This is this is an example that is due to um one of my students actually had this idea where he uh asked uh a language model to provide a recipe for how to make uh a sandwich in international phonetic alphabet. alphabet. alphabet. International phonetic alphabet is these symbols that that they use in a dictionary to explain how to pronounce a word.

And it's very peculiar because it only ever appears for individual words in a dictionary. like you never see tech free form text written in international phonetic alphabet but if you ask a good language model it will write paragraphs in IPA for you and that is compositional generalization that means that you have never seen this particular language this particular uh alphabet used to write paragraphs but you understand paragraphs you understand that it's compositional with different alphabets so you can solve the problem and you can imagine the same thing coming up in robotics that you've learned a repertoire of skills and now you can combine and mix those skills and apply them to solve new problems problems problems it makes me wonder what the last type of tasks you think will be possible for a robotic system robotic system robotic system to achieve.

I think changing a child's diaper will be really really hard. Say more. Say more. Say more. Well, um I think that there is a I I think this this really is just Morix's paradox all over again that people are extremely good at certain things. We're very good at physical things. We're also very good at interacting with other people. And it makes sense like we have to be like that's a lot a lot of our existence. So things that involve behaviors that interact with other people uh where you have to like uh you know actually help somebody like you have to help somebody get out of bed or something like that.

I think that's a lot harder than people appreciate. So I think like elderly care, taking care of small children, I think those things are going to be hard and they're probably going to be harder than people think. And the stakes are very high. I want my baby to be last. It's not just that. The stakes are high in many places. It's just that it's probably like the pinnacle of something that fools us into thinking that it's easier than than that it's easier than it really is, right?

Because just like we are so evolved for interacting with people and doing things physically and um you know if you're if you're helping somebody get up the stairs or get out of bed or something like that like you don't have to think very carefully about how you're going to do that like you you kind of know. So I think it's really the pinnacle of Morgan's paradox. If I think about an LLM as a brain and now it's effectively studied everything, I don't know how else to put it, and then I think about a robotics brain, a robotics model's brain instead, what are like the dark parts of the brain?

Like what what has it not been able to study or penetrate or learn? Like what are the areas that have just been really difficult that matter but have been hard to to to for us to get into? One of the things that people are remarkably good at is using physical analogies to understand other situations. I don't know whether this is something that LLMs can or can't do, but um it it is something that people use a lot. They use it in everyday life and they also use it for very sophisticated problems.

So for example, you could say that company has a lot of momentum. That's a physical analogy. You know exactly what it means. Like I don't have to explain that statement to you. But if you actually think about that, it is like quite a complex thing. There's like a lot writing on that word momentum. There is a an interview with u Richard Fineman where he talks about analogies that he makes in regard to subatomic particles and he says like okay we use like the word spin okay like the thing is not really spinning like it's not like a spinning top but all those kind of analogies really help us make sense of it and not just in a way that allows explaining concepts but it actually leads to conclusions that actually leads to inferences and those inferences actually make sense right um so that's kind of remarkable that we we are so primed to interact with the physical world.

So primed to have physical intelligence that you can use it in everyday speech by saying that company has a lot of momentum and you can use it when advancing fundamental theoretical physics. physics. physics. That's kind of remarkable. I don't know if LLMs can do that. Uh maybe they can, but I think that really understanding physical interactions, causal structures, all that kind of stuff. Uh there is something kind of special about that and it's clearly something that people get a lot of mileage out of. I'd love to talk about the the role of researchers and the actual people doing the research in LLM world.

It's fairly shocking how few people are in the great at the global scale responsible for basically all the progress in in LLMs, someone like Ilia as an example. What is that like in robotics? Like how many people in the world are truly impacting this trajectory? And then I want to ask about like what good research means. I think those kinds of questions are often very hard to answer about science because I think that you know we sometimes have a tendency when we especially when we look at at history to underline like particular milestones um and certainly in machine learning this is the case like you can say okay like well like Alex net was a was a big step forward this was a big that's true but I think it's also important to remember that these uh advances um they happen because lots of people are trying lots of things and even some of the failures are very instructive.

Like I complained before a little bit in a low-key way about the controversy around endtoend robotic learning, but I don't know if robotic learning would have advanced the same way if it were not for the for the controversy, so to speak. It is true that you can kind of look through the list of successes and mark down that like, oh, like these folks were like, you know, have a a history of repeatedly hitting home runs. But I think in reality in the scientific community it's not just the home runs that are responsible for progress and even and even some of the failures and even some of the bad ideas are very instructive and pushing towards the good ideas.

Yeah, it's fascinating to think about. Um the example you gave before is so interesting where the research insight was like just give it some coaching and it gets better. Seems like that sort of insight is is can be very powerful and and high leverage which makes me wonder like what have you learned about what makes for a great researcher? Research is definitely different from engineering because in research the important thing is to get to an answer to a question which often requires cutting some corners. And one of the most delicate decisions in research is when do you try new things versus when do you stick with what you're already trying?

And that's very very delicate. It's very very hard to figure that out. And if you get it wrong then you can miss something really remarkable. So if you get it wrong and you don't stick with something for long enough you might be like right there. You might be about to get to the answer and then you stop just short of it. That's terrible. Or you could get stuck like just hammering against something that's never going to give way for years. Uh so deciding when to sort of turn a little bit and look this way and that to open yourself up to more opportunities versus when do should you keep hammering on the thing because you're about to get the solution.

That's often the most important decision and people some people have an instinct for getting that right and that counts for a lot. You've obviously been in and around and are one, you know, great researchers. What are these people like like as people? How do they tend to be distinctive from, you know, the average person? person? person? I think they're just the same. I'm thinking about the the the people that that I deeply respect that are really good at this stuff. I have a very hard time thinking of a single set of personality traits.

the one the one constant I that there is no constant basically just you know there might be a commonality in that to do effective science you have to be very passionate about that but even that passion can come from many different places like I've worked with people that were remarkably effective that are just driven purely by like the desire for novelty like they don't give a damn about what their technology does they don't give a damn about whether it's useful they just want like cool new ideas I've also worked with other people that just like really want to solve a particular thing and they're just as happy like you know uh uh building stuff as they are testing out experiments as they are just like hammering away at things like whatever whatever it takes and all those types can be very effective.

effective. effective. You mentioned uh the difference between research and engineering which also makes me think of manufacturing like Elon would be fond of saying that the factory is the product like the hardest part of this whole equation is actually the scale up of you know whatever that whatever this thing ends up looking like of making you know 100 million of those. How do you think about that part of the equation or is it too remote at this stage to spend too much of your time? No, I think it's an important part of the equation.

I'm not sure it's like the part of the equation that we most need to figure out right now, but it's certainly part of it. Um, and as you might have guessed from my answers to the other questions that a lot of how I prefer to think about this is to is to is to figure out that the hard part and then enable a lot of experimentation on the other parts. Right? So yes, making a robot at scale is difficult. Making a robot at scale is even more difficult if you don't know what kind of software is going to run it afterwards and you're not even sure whether it's the right kind of robot.

So I think one of one of the really valuable things we can get out of general purpose uh AI tools like robotic foundation models is the ability to like get a lot of the other stuff figured out so that at least like some of the uncertainty goes away so that when when you scale things up you have some confidence that this is like really going to work. A lot of people that listen to this are entrepreneurs, people that run companies. A very popular question has become how should a traditional company begin to think about using LLMs or preparing itself for you know the the ongoing improvement of these models.

How would you answer the same question for robotics? robotics? robotics? It's a very good question. It's also a very difficult question because the technology is changing so rapidly. So um and I want to illustrate why this question is difficult with an example. So here is a particular uncertainty about the tech. This is going to be a little bit specific, but it's a it's an example. example. example. Will the robots rely more on demonstrations or on reinforcement learning from autonomous data? We're working on both of those things, and they're clearly both important.

But how somebody should prepare for the technology will be pretty different if they're expecting that they need lots of teley operation to produce lots of demonstrations and like a little bit of autonomous experience versus the opposite like a tiny number of demonstrations that huge amounts of autonomous experience. Like is it 9010 or 10 1090? uh and that's something we're hopefully going to learn about over the next few years, but it does change the uh correct approach pretty dramatically. dramatically. dramatically. So that's kind of a case a case study of how changes in technology will dramatically alter this from a business standpoint.

Is the right way to think about it just like get really clear on the economics of the labor in your business or something? And I'm curious how you think about that like the way that this will change the nature of labor itself. I think um coding tools are like a really nice example to look at for a template of how this might work. Like it it's not like coding tools came on the scene and suddenly uh we don't need software engineers anymore. It's that the coding tools increase the productivity individual software engineers.

There's some amount of work that needs to be done to make sure that people are able to use them. uh there's some amount of technology development that needs to be done to make them useful for the the appropriate use cases and these things are co-evolving and they're also still changing like you know coding agents are different than like you know code completion tools and so on but I think it's like a nice template for us to look at to see how AI tools combine with people doing a job increase their productivity and also you know raise new challenges and I think we'll actually see something like that with robotics too that a more realistic template is not like you know the the humanoid like goes in and the people just leave.

I think it'll be more like there are some aspects of the of the job that can be done by a robot, some that can be done with a robot working together with a person. Uh some that can be uh you know where the person needs to like do something special to make the robot more productive, some where it's the other way around where the robot does something that makes the human more productive and it'll be this kind of dance that we've seen with coding tools. Do you have a favorite robot that you It's not part of what physical intelligence is doing.

I do really like the uh the Boston Dynamics robot, the the new especially the new version of the Atlas because it is in some ways very humanlike and in some ways very not humanlike, right? Like they they they made some interesting decisions about how they want more range of motion on the joints, so it can do some like pretty cool things. Um it's also a very agile robot which is really cool. It makes those awesome demos. So I'm a big fan of that. I'm generally a big fan of like everything that Boston Dynamics has done.

Should or could anything be read into the fact that Boston Dynamics has been doing very cool demos for a very long time and don't actually do anything useful for PE for customers? Yeah, it's a fair question. I think it's also a fair question for lots of robotics companies to be fair. Um I I think that you know what I'll say in general terms is that I think it there is a lot of value in demos that serve to illustrate challenges on the road to something useful and productive.

And uh you know obviously you can also do a demo without being on the road to something useful and productive. But I think that there is value in demos. I think that demos that are used correctly in service to a mission can provide people with an illustration of what to expect and they also provide a challenge. You just have to be like honest in setting up that challenge. How much do you think about the the business endpoints? Like I think to this point Roomba is like the bestselling robot of all time in in the consumer category which is kind of surprising.

Um, and of course we might be on the edge of some sort of Cambrian explosion, but how much of your cycles do you spend thinking about like this is the shape of a product that might result from this that maybe is the way we bootstrap our way to all this data? data? data? Yeah, Yeah, Yeah, I mean I certainly spend some time thinking about it. Um, I think it's just something that's very hard to reduce to like a very concrete answer right now. Uh, but it's not too bad to like think about a space of possibilities.

And you know a lot of what we're doing when we develop uh our models when we experiment with different tasks when we do demos like the robot Olympics underneath we're kind of prototyping what does it look like when we try to do something real with this obviously to different degrees of real and what goes wrong. So it is something we think about a lot. It's not something that I have even close to like a concrete answer to but you know there's a space of possibilities and a lot of what we actually are planning to do in 2026 is also experiment with different things in that space.

that space. that space. When you study the history of general purpose technologies which certainly you know this would be a major one if it comes to fruition you often find this constellation of things happening around that thing that enable it. Obviously like LLMs are a direct complement to what you're doing. Are there any other surprising technology areas or trends that help you do what you do but are different? different? different? So, one interesting thing is that robotics hardware has become dramatically more affordable over the last few years.

So, um when I started working in robotics about a decade ago, I worked with a robot called a PR2, which I believe had a cost of about $400,000. $400,000. $400,000. Um when I started my lab at UC Berkeley, I uh used a robot that uh was in the ballpark of $30,000. Now each arm on this thing is maybe a tenth of that. Uh we think that can be even less. And that's not due to like any one single technology. It actually involves both hardware and software.

So the kind of um uh lowcost arms that we have here, they wouldn't be useful in an industrial setting because traditional control methods that rely on a great deal of precision wouldn't be able to use them. So there's kind of a a c a wide range of different advance a cluster a constellation as you said uh that have pushed down the price point of these things and I think that does make it a lot more practical to think about general purpose robotics today for people that would want to be fairly technical about following uh major milestones that are happening in this field.

Where does that information show up? up? up? A lot of it shows up in research papers. Uh research papers unfortunately are not a very accessible source of information because uh it takes uh a bit of care to like kind of sort through everything and figure out what is the signal and what is what does something really mean like you know research results are sort of intended for uh an audience that already understands the starting point from all the past research results. Um but that that's a big one.

I think that robotics and I think technology in general is one of those things where the public facing artifacts like the demos and the videos that somebody might post on social media are often actually not very good for providing a sense for the true underlying state of things because they're sort of meant more as like kind of a a demonstration at the edge of capability and grounding that like what what what does the demo really mean requires digging deeper. Um so yeah pro probably research papers are the way to go.

Sometimes even worse than that, you have to actually go talk to the individual people and find out what the inside story really is. And uh you know, maybe that's not a great uh situation to be in, but that's kind of how science works. works. works. As we look forward to the future in your mission, what feels the most uncertain? I do think the timeline is uncertain. Um I'm, you know, if anything, my sense of the timeline has gotten more optimistic since we started. But it's uncertain because of the nature of the technology that this is something that is where there's a bootstrap challenge like getting to a particular level of usefulness so that uh robots can be deployed so they can do useful tasks so they can start collecting data from open world settings at scale and because that's such a like a a sudden kind of event with uh getting past the activation energy I think there is a lot of uncertainty about uh the timing of that and that's exacerbated by the fact that The timeline looks different depending on what kind of technology is deployed.

So like the example I gave before about whether uh it should be data collection through teleoperation or data collection with autonomous systems or something in between maybe shared autonomy maybe like this this coaching kind of thing like those all sort of change the picture in terms of how deployments work and how in the wild data collection works. So because of that I do think there's like quite a bit of uncertainty about time. You're in such an interesting position because um you're at the center of research.

Lots of different kinds of people are talking to you, asking you questions. What What are questions that you're surprised people don't ask you? But what are the things that people don't ask you about that they should? Well, I think the question you asked earlier actually about uh how how somebody should prepare. I think there's a variant of that question which would be something like okay if I want to start using autonomous robots for a thing like what should I start setting up? Yeah. Yeah. Yeah. Should I should I set up operation?

Should I set up uh you know should should I modify my task in some way so it's more accessible? Should I design new hardware? Like maybe I should design new hardware so I can plug your software into it. And I think people make a lot of assumptions about that. Like for example uh one assumption is like well machine learning requires data so let me just like figure out something that will collect data. That's not often the best assumption because you need the right kind of data.

Like maybe some data is easy like it's easy to get like videos of people doing something but that doesn't mean that's the right kind of data. And it might be domain dependent. might be dependent on the on the your thesis about the technology that will succeed. So, I think that people do make a lot of assumptions about that. Not that I necessarily have a better answer for them even if they ask me, but it's something where there's a big space of possibilities. possibilities. possibilities. We talked about these like big uncertain long-term timelines.

What is like the very next thing you are trying to solve that's like extremely visible? So without like giving too much away uh what I can say is that uh a big focus for us right now is actually better understanding this kind of like mid-level reasoning part of the problem because we think that we have a pretty good sense for how to acquire the low-level physical behaviors but getting those low low-level physical behaviors to generalize requires bringing to bear a lot of this kind of common sense knowledge and knowledge and knowledge and the representation of that might be really important.

So LLMs make certain kinds of representations very convenient. They make it very convenient to basically turn text into other text, but that's not necessarily the best representation for what an embodied system needs to do. Like sometimes it needs to think about things more spatially, sometimes semantically, sometimes other representations. And trying to figure out exactly how to structure that kind of internal thinking process might be a very important question and the answer to that question might be different in the world of embodied foundation models than it is in the world of LMS.

So that's like a concrete thing that we're working on. Now, Now, Now, where do you think you fall? Like if I if I could somehow get the hundred most uh informed and active robotics researchers in the room at once and pull them on how certain they are that things will have, you know, unlimited capabilities and how soon that might happen. Where where do you fall in that distribution? I'm on the optimistic end when it comes to um established robotics researchers and on the pessimistic end relative to um uh robotics entrepreneurs.

entrepreneurs. entrepreneurs. Interesting. And and I understand the entrepreneur part for sure. Uh they're optimistic by nature. Why are you on the optimistic end of the researcher community? community? community? Robotics has a very long history uh which has precious few successes. Let me put it this way. Especially when it comes to robotic AI. So, I think if we're being honest about it, like you know, most robots that are out there doing useful work are still running, you know, state-of-the-art technology from the 1980s. Uh, and there that's because the robotics problem is hard.

Maybe not our fault. It's just a difficult problem. Um, and because of that, I do think that there is like good reason for caution to say that well, uh, okay, maybe we've like made a lot of headway on this part of the problem, but there's like many other problems that still remain. remain. remain. Right? Now, part of why I'm optimistic about this is that I kind of have like a sense of what has proven tough for me before and I can see a lot of the puzzle pieces that I'm imagining could be slotted in to address many of those things.

things. things. But, um, you know, as my co-founder Carol likes to say, like when you when you've climbed the mountain, only then do you see if there's another mountain after it, right? Uh, and in robotics, there's been a lot of experience of lots of mountains. of mountains. of mountains. Given that endurance is required, who or what most inspires you? I am actually quite inspired by Boston Dynamics. Um, again, I think there's like a lot of things that we can debate on the technology side, but I think that there is a lot of value in repeatedly showing something that people wouldn't have thought possible even if there's all sorts of caveats and assumptions and so on.

uh and certainly in robotics like you know what whatever we might say about demos and whatnot like I think it's very fair to say that people have revised their thoughts about what's possible from seeing uh some of that stuff. stuff. stuff. Uh so I think that's that's definitely one. I think I'm also inspired by organizations that create an atmosphere for experimentation and I think that you know there are some research labs that have done a very good job of this. I think actually OpenAI has has historically done a great job of this of creating an atmosphere where individual researchers can experiment with things and be empowered to see those things through like you know chat GPT was you know basically John Schulman's pet experiment for a while like it wasn't a concerted corporate strategy with lots of you know spreadsheets and uh pie charts it was like a pet project so I think there's something pretty inspiring about organizations that empower people to have pet projects turn into world changing ing uh successes successes successes and you know certainly one of the aspirations that I and my co-founders have here at physical intelligence is to provide some of that to the best of our ability.

Uh and it's it's it's hard to do like you kind it's very hard to have an organization where that has that kind of capability. of capability. of capability. I feel like Google used to have that one day you can do whatever you want thing. Is is that the spirit of it or or what? I I I was absolutely shocked when I started working at Google at the level of um I guess leverage that I felt I could have. So um one of the projects that I did in um with with many of my colleagues there in 2015 was uh colloquially referred to as the arm farm.

So we we took a couple dozen robots, put them uh in a in a lab and had them collect data. And that was like a very bottom-up thing. It's just like I found out from somebody that they had a warehouse full of robots that that nobody was using. Uh and I asked uh Jeff Dean and Man Hook if we could like stick them in a lab. And I was just thinking like, okay, they're not going to take me seriously. I was, you know, I just started.

I was like, you know, uh level four research scientists. Uh and Jeff was like, yeah, let's do it. Like, you know, what do you need? And I think that I I just remember feeling like, wow, I had never in my life thought that, you know, I'd have that that kind of uh leverage. I mean, obviously, you know, I was very young at the time. Uh and I think that that's that's very special. And I think getting to a place where people can unlock their creativity and have that kind of agency make can make for a very remarkable place.

My friend uh Jesse has this great question, which is um for companies that you're not involved with, which one do you most hope succeeds and why? Um people used to say boom. a lot because they want to fly places faster. Increasingly, as I've asked this question, people have said pi because just the the sheer impact that it might have have have if you're successful is massive on such a global scale. And it's been really fun just to hear about all the ins and outs of how you're thinking about the problem and attacking it.

When I do these interviews, I have the same traditional last question for everyone. What is the kindest thing that anyone's ever done for you? It's a tough question to answer because I do I do think there are many moments in my career where I felt like I I got a leg up on something. I think I have the kind of personality where I sometimes don't appreciate in the moment and only reflect on it afterwards. I don't think like I have a single answer to that, but probably like the three moments in my career that stand out.

Actually, one of them I had already mentioned to you, which was the the arm farm thing. I I'm especially grateful to to Jeff and to Vincent for kind of willing to take that that bet on on me and my colleagues. And there are a couple other moments like certainly I think when I started my uh posttock with Peter Beal at Berkeley I had like zero robotics experience. I had done like virtual character animation and computer graphics right uh and I felt like that was sort of a bet on my potential more so than my actual accomplishments.

And there was another moment even earlier on that maybe is like even more minor where when I was in college, I got an internship at Nvidia that uh really got me to like experience some cool uh stuff when when I was just like a sophomore and uh I think the hiring manager for that also took a bet on me and I think that these kinds of things I think they really matter in a person's career and I think that uh uh maybe at the moment I should have been more grateful but certainly in hindsight it's something that made a big difference and hopefully I can make that difference in other people's careers as well.

Well, I've learned so much from you and your co-founders and so much today. Thank you so much for your time. Thank you. Most software companies try to maximize your time on their app to juice engagement. RAMP does the exact opposite. RAMP understands that no one wants to spend hours chasing receipts, reviewing expense reports, and checking for policy violations. So, they built their tools to give that time back using AI to automate 85% of expense reviews with 99% accuracy. And since RAMP saves companies 5%, it's no wonder that Shopify runs on RAM, Stripe runs on RAMP, and my business does, too.

To see what happens when you eliminate the busy work, check out ramp.com/invest. As your business grows, Vant scales with you, automating compliance and giving you a single source of truth for security and risk. Learn more at vanta.com/invest. vanta.com/invest. vanta.com/invest. Ridgeline is redefining asset management technology as a true partner, not just a software vendor. They've helped firms 5x and scale, enabling faster growth, smarter operations, and a competitive edge. Visit ridgelineapps.com to see what they can unlock for your firm. OpenAI, Cursor, Anthropic, Perplexity, and Verscell all have something in common.

They all use work OS. And here's why. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, skim, arbback, and audit logs. That's where work OS comes in. Instead of spending months building these mission critical capabilities yourself, you can just use work OS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on work OS. Work OS is the fastest way to become enterprise ready and stay focused on what matters most, your product.

Visit works.com to get started. Every investor should know about Rogo because ROGAI's platform is not just another generic chatbot. Instead, it was designed to support how Wall Street bankers and investors actually work. From sourcing, diligence, and modeling to turning analysis into deliverables. For me, three key things differentiate Rogo. First, it connects directly to your systems, so it can work with your actual data. Second, it understands your workflows, how work really happens across a deal or an investment. And third, it runs endtoend and produces real outputs the way the best people do.

Auditable spreadsheets, investment memos, diligence materials, and slide decks that match your standards. This all comes from the fact that ROGO is built by finance professionals for finance professionals and it's already being adopted by some of the most demanding institutions in the world. To learn more, visit rogo.ai/invest. AI/ Invest