The Brain Behind the Body: Foundation Models and Embodied Intelligence

Articles
August 28, 2026

In our last article we said the second great bet in robotics was on the brain rather than the body. Here we go and check whether that brain exists yet - and what the answer leaves for a market that builds things.




In our last article we argued that robotics had grown into its own vertical inside applied AI, and we drew a map of it: five layers, running from the motors at the bottom to the business models at the top. We also made a claim in passing that deserves rather more scrutiny than we gave it. The second great bet in this category, we wrote, is on the brain rather than the body - hardware-agnostic foundation models that aim to control any robot regardless of its shape.

Something like thirty billion dollars of private company value currently rests on that sentence being true. If those models work as advertised, this layer is winner-take-most, a handful of very well-funded laboratories take it, and there is very little in it for anyone else. If they do not, the layer fragments into a hundred specific problems - and fragmented layers are where funds like ours have always made a living. So we went to check.

What the brain layer is, and how it arrived

For fifty years, teaching a robot a new task meant writing that task out by hand: every waypoint, every gripper angle, every exception for the one time the box arrives upside down. A robot foundation model replaces all of it with something closer to instruction. The model takes in what the robot can see, plus a sentence in ordinary language, and produces the movements of its joints. The distance between those two approaches is the distance between programming a machine and training one, and it is why this layer suddenly matters.

The breakthrough is recent enough to date precisely. In 2023, Google published a model called RT-2 and showed something nobody had shown before: a model already trained on the internet — one that knows what a banana is, and that bananas turn up in kitchens — could be taught to produce robot movements the same way it produces words. The robot inherited an understanding of the world that nobody had programmed into it. Alongside it came Open X-Embodiment, a shared training set pooled from 21 institutions covering 22 robot types, 527 skills and more than 160,000 tasks, on the theory that data from many different robots would make every one of them better.

What happened next set the pattern for everything since. Within a year, an open model called OpenVLA - built largely by academics on that shared data - beat Google's own version by more than sixteen percentage points across twenty-nine tasks while being roughly seven times smaller. The frontier was overtaken by a free download before the frontier had a product. Over the following eighteen months the design settled: Physical Intelligence, NVIDIA and Google all converged on variants of the same recipe. And the openness held. NVIDIA publishes its GR00T models for anyone to use commercially at no cost, and this May the Shenzhen-based X Square Robot gave away Wall-OSS-0.5 in full, reporting that it beat Physical Intelligence's π0.5 by 17.5 percentage points on a fifteen-task real-robot suite.

Sources: Google DeepMind, the Open X-Embodiment collaboration, OpenVLA, Physical Intelligence, NVIDIA and Figure AI publications, 2022–2026.

Three years, then, from a research result to something a graduate student can download on a Tuesday. Which reorganises the map we drew in that first article.

The map needs a correction

We described this layer as a contest between the companies building brains - Physical Intelligence with its π-series, Skild AI with its omni-bodied model - set against the companies building bodies. That was true as far as it went. It was also incomplete, in a way that matters more than the part we got right.

The two largest participants in this layer are not trying to win it. NVIDIA gives its robot brains away because NVIDIA sells the computer underneath them, and every robot that ships with one is a sale regardless of whose model it runs. The pricing tells the story better than any strategy deck. In July, without announcement or explanation, NVIDIA raised prices across the whole Jetson line by as much as 101 percent: the Thor production module went from $2,999 to $4,999 in thousand-unit volumes, the developer kit from $3,499 to $5,499. The brain got cheaper. The hardware it runs on got considerably more expensive. Google is playing the same game from the software side. Gemini Robotics 2, announced on 30 July, ships as three models, demonstrates on Apptronik's Apollo 2 humanoid and comes with an AI partnership covering the next Boston Dynamics Atlas. Google does not want to build the robot. It wants to be the layer underneath every robot - a position it has occupied before, in a different industry, to considerable effect.

For a startup whose entire product is the model, this is the single most important fact in the category. You are not competing against two well-funded rivals. You are competing against two of the largest technology companies on earth, who are working to make the thing you sell worth as close to nothing as they can manage, because they are paid somewhere else.

But read that fact the other way round, because we think the second reading is the more important one. The giants are not closing this category to founders. They are subsidising its foundation. What they are giving away free is the expensive part - the years of research, the pooled data, the pretraining runs nobody outside a handful of laboratories could afford. What is left over is the part that requires a specific customer, a specific machine and a specific room - and that was never something a research laboratory could take from you anyway. Everyone now starts from roughly the same model. That has not been true in any previous wave of this industry.

What thirty billion dollars bought

The capital has not been deterred. Crunchbase counts $47.4 billion of global venture funding into physical AI in the first half of 2026 across 521 deals - nearly four times the $12 billion raised in the second half of 2025, and more than the $41.9 billion invested in the three years from 2022 to 2024 combined. But the more revealing number is the deal count. Dollars rose roughly fourfold while the number of financings rose about eleven percent. The cheques got enormous; the number of companies getting funded barely moved.

Source: Crunchbase News, physical-AI venture funding, H1 2026. Crunchbase's definition covers robotics, autonomous vehicles, aerospace, drones, industrial automation and sensors.

One transaction captures what is being priced. In January, SoftBank led Skild AI's $1.4 billion round at more than $14 billion, against roughly $30 million of revenue. Three months earlier, the same firm had bought ABB's robotics division - a real business with real factories - for $5.375 billion against $2.3 billion of revenue. The same buyer paid roughly two times revenue for the body and several hundred times revenue for the brain. That gap is not carelessness. It is a specific bet, made in cash, that a general robot brain is coming and that whoever owns it owns the industry.

The promise is not yet proven - and nobody can check it

So it is worth being precise about what is actually being bought. The promise is generality: one model, any robot, any task, transferring to buildings and objects it has never seen. That is what “omni-bodied” means, and it is what justifies paying several hundred times revenue for software rather than two times revenue for a factory.

The best evidence available says the promise is not yet kept. RobotArena∞, accepted at ICLR 2026, took models from laboratories around the world and tested them across hundreds of settings. Its finding was blunt: performance falls away as soon as a model is moved somewhere it was not trained, which suggests these systems are not general-purpose at all, but unusually good at recognising the particular places they were taught in. The laboratories' own numbers point the same way. Physical Intelligence reports its π0.7 model succeeding more than nine times in ten at tasks it was trained on, and six to eight times in ten at tasks it meets for the first time. Google reports a single Gemini Robotics 2 model driving three different robot bodies succeeding at fine finger work anywhere between thirty-two and ninety-two percent of the time, depending on the task.

Sources: Physical Intelligence π0.7 technical report (April 2026); Google DeepMind Gemini Robotics 2 disclosures (July 2026). Figures are as reported by each developer and are not directly comparable across models.

Those are honest numbers from serious teams, and they are nowhere near what a paying customer needs. But here is the part that should govern how an investor behaves: there is no way to verify any of it from outside. Chatbots are graded against standard exams that anyone can run and everyone accepts. Robotics has no equivalent - no agreed test, no independent league table, no fixed set of tasks on fixed hardware. Every number in the paragraph above was produced by the company reporting it, on its own robots, in its own building. So the position today is this: the most valuable claim in this layer is unproven, and there is no instrument with which anyone outside these companies can test it.

We argued previously that a robot company should be held to the same standard as any applied-AI company: a deployed workflow, a paying customer, and a defensible answer to why an incumbent cannot simply copy it within a year. In this layer the third condition is brutal, because the incumbent is NVIDIA and it copies for free. That is not a reason to avoid the category. It is a reason to be extremely specific about where inside it you stand.

And there is a second way to read the RobotArena∞ result, which we think matters more than the first. If these models really are specialists at the environments they were trained in, then the environment is the asset. A founder with genuine access to an unusual one - a line, a greenhouse, a hangar, a fleet - holds something the frontier laboratories cannot buy their way into, precisely because their models do not transfer into it. The same finding that makes a $14 billion valuation look fragile makes a small company with a strange building look durable.

What this means for Türkiye

Start with the arithmetic, because it shapes everything that follows. The median disclosed round in this layer is larger than most Turkish funds, and Türkiye's entire startup ecosystem raised about $1.4 billion across 360 deals in 2025, three-quarters of them at seed. Building a frontier foundation model from Türkiye would be an outlier achievement rather than a plan, and we would be glad to be proved wrong by a team that does it. For most founders the more useful question is not how to win that contest, but where else in this layer the value sits.

But look again at what the evidence actually said. If generality is not real - if these models are specialists at the environments they were trained in - then there is no single brain to buy. There is only a model meeting a particular process, in a particular building, with particular machines and particular ways of going wrong. And in that world the scarce input is not compute or model talent. It is proprietary industrial data, and physical access to the machines that produce it.

Which is the same argument we made in that first article, extended one layer up. We wrote that this wave, unlike the last one, cannot be won with a browser tab and a GPU cluster alone — that it needs steel, motors and a factory floor, and that this is what makes it interesting to investors who have spent their careers in markets that build things. The point we did not make is that the software layer runs on the same fuel. Türkiye's 1.42 million vehicles a year, its $10.56 billion of defence and aerospace exports, its farms and textile plants and drone fleets are not only manufacturing credentials. They are training data that does not exist in any shared dataset, attached to operators who can be persuaded to let a small team instrument their line.

The openings follow from that directly. The shared datasets everyone trains on are full of American warehouses and laboratory kitchens; they contain almost nothing about textile finishing, hazelnut sorting, greenhouse harvesting or airframe inspection. A model that is worse than π0.7 at everything and better at the one task a customer pays for is a perfectly good company. Denied and degraded connectivity, which is an inconvenience to most research teams, is the design assumption in Türkiye's largest robotics sector. And policy has moved the same way: on 18 August, Türkiye's 2026–2030 AI Action Plan entered into force by Presidential Circular, naming physical AI and robotics as priority areas and committing to dual-use incentives designed to move robotics and autonomous-systems capability out of the defence industry and into civilian applications, with three priority pilots, two university-industry robotics centres and three defence-to-civilian transfer projects targeted by the end of 2027.

What we will and will not back

We are not the right first cheque for a company setting out to pretrain a foundation model. That race is run at a capital scale where we could not be useful to a founder, and being useful is most of what an early investor is for. We are also cautious about companies whose only asset is the model itself, for the reason set out above: two of the largest firms in the world are working to make that asset free.

What we are actively looking for sits downstream of the model and next to the work. The adaptation layer, where a general model is turned into a reliable one for a specific process, is the clearest case - closer to the customer, far cheaper than pretraining, and the place where most of the remaining engineering actually lives. It is also not hypothetical. Qualia, founded in Copenhagen last year, does exactly this: its engineers walk a customer's production line, decide which flows are ready for a foundation model and which are better left to classical automation, then fine-tune an open model - π0.5, SmolVLA, GR00T - and hand back the weights on the customer's own hardware. A pre-seed team from the Nordics, running on models it did not build, selected this year for Google DeepMind's robotics programme. That is what a company looks like on the right side of this argument, and there is no reason the next one cannot come from Istanbul.

Three more positions interest us for the same reason. Vertical brains built on data nobody else can gather. On-board inference, where the model has to run on the machine rather than in a data centre. Evaluation and verification tooling, which on the evidence above is an unclaimed position rather than a crowded one. And the integration and robotics-as-a-service models that sit at the top of the map we drew earlier in this series, where a customer pays monthly for a working robot instead of buying one.

What we ask founders is that same test, made specific. Show us the number that gets worse when conditions change, not the one that looks best. Tell us who chose the evaluation, and let us choose a different one. Be honest about which parts of your system you built and which you took off the shelf - starting from a free model is good engineering, not a weakness. And show us the loop: which customer, doing which task, generating which data, that makes your system better next quarter in a way nobody without that access can replicate.

None of that is a high bar because we expect little. It is a high bar because the tools are now genuinely on the table. The most capable robot brains ever built are free to download, the compute to run them fits on a machine, and the one thing that cannot be downloaded is a real process with a real customer attached. That combination has never existed in this industry before, and it favours the people closest to the work.

What this comes down to

The robot brain turned out to be the fastest-moving and least defensible part of this stack. It went from research result to free download in three years, the two biggest players are giving it away on purpose, and the generality that justifies its valuations has not been demonstrated and cannot currently be independently tested. None of that makes the category less exciting. It makes it a different category than it first appeared - one where the durable value sits not in the model but in the specific, unglamorous, hard-won fit between a model and a real place of work.

That is a genuinely good outcome, and not only for us. It means this layer is not being quietly settled somewhere else while the rest of us watch. The most capable robot brains ever built can be downloaded tonight by anyone with the curiosity to try them, and what decides who wins from here is not who has the largest cluster but who can get a machine to work, reliably and repeatedly, in a room that has never appeared in any dataset. There are a great many such rooms in this country. We would like to hear from the people who know them best.