Osama F Rama
← All writings

Physical AIEmbodied IntelligenceMoonshot

Solving Physical AI

Towards Embodied Cognitive Intelligence

· 8 min read

1. A decade of progress

AI has progressed in three stages over the past decade. Specific models: one model per task. General models: one model for many tasks and modalities. Agents: models that take a world state and an instruction, then plan and act until the goal is met. Agency changes the unit of value, from an answer to a completed task, which can be verified. Verification closes the loop: the outcome of every action becomes a training signal.

The next decade belongs to physical agents. Embodied intelligence incentivizes the architectures and learning paradigms behind human-level cognition: continual learning, sample-efficient learning, and self-improvement, each demanded by a world that changes, charges for every trial, and verifies every action by itself.

This manifesto sets out our goal: a System 1 for embodied cognitive intelligence (ECI), the fast layer that perceives and acts. ECI is a system with the complete capabilities of human cognition, operating in the physical world. It perceives through its own sensors, plans, acts through its own body, and continually learns from its experience.

2. The case for Embodied Cognitive Intelligence

2.1 The scale of the physical economy

The global economy is about $110 trillion a year, and most of it is atoms. Of the world's ~3.5 billion workers, most do work that needs a body, from farms and factories to hospitals and labs.

Every economy in history has been a function of its people and their productivity. Physical agents break that function twice. Supply: physical agency becomes a manufactured good and scales like production. Demand: a person with a team of assistants gets whatever they want done, most of it work nobody hires for today.

Every person ends up with a team of agents, digital and physical, and robots outnumber people. Growth has always compounded at the scale of human labor. With agency manufactured, economic growth becomes a function of physics.

2.2 Why human-level cognition needs a body

Grounding in the physical world supplies inductive biases that language can only describe. The world is continuous, causal, and made of persistent objects, and it enforces those rules on every action. Infants show object permanence and intuitive contact physics months before they have words: the priors come from experience, and they make learning efficient.

The physical world runs on consistent rules, so every action yields a verifiable outcome, and with it a self-supervised learning signal.

Physical intelligence forces addressing the open problems in AI. Continual learning, sample-efficient learning, and energy-efficient inference come naturally to animals and remain unsolved for machines. Each is optional for a language model and mandatory for a body in the physical world. Embodiment is the most efficient path to human-level cognition and recursive self-improvement, and very likely a necessary one.

3. Why physical AI is unsolved

Each stage of AI began with a new training signal: labels, then the corpus of human text, then verifiers such as tests, leading to reinforcement learning with verifiable rewards (RLVR). Physical agents learn from embodied experience, the most informative signal and the most expensive: the world is a perfect verifier, but every verification costs a trial in time, hardware, and risk. Physical AI is the problem of getting the world's verification at the price of a test.

Current approaches for physical AI run on human data. Teleoperation yields one robot hour per human hour, at less than the operator's own speed. Scale is capped by fleet size times paid humans; the policy is capped by the demonstrator's speed.

VLAs rode the success of LLMs. But bolting an action head onto a language model inherits the language-modeling objective in place of objectives optimized for the physical world, and falls short on precision, speed, and contact.

World Action Models (WAMs) pretrain a world model on video, then teach it how to act by fine-tuning on human trajectories or preferences, the way RLHF tuned language models. The fine-tuning reintroduces the human bottleneck, and none of it runs at control rate.

The world already provides a better reward than any human. RLHF exists because the only check on language is a person saying which answer is better. Physics checks itself. Did the task complete, how fast, at what energy cost, did anything fall.

3.1 Scaling demonstrations doesn't work

Passive data answers the wrong question. A demonstration shows what happened when someone else acted. Control needs what will happen when this policy acts, in the states this policy reaches. The states that matter are the ones the data systematically lacks.

The objective isn't in the data. A text corpus carries its own supervision: predicting the next token is the task. A trajectory's label comes from outside it: the same motion is a success or a failure depending on the goal and on what the world did. Scaling trajectories scales behavior without scaling the objective. Reward has to come from experience.

Animals are the only existence proof of physical intelligence, and they did not scale data. Every animal reaches motor competence from an innate prior plus self-generated practice, verified by the world. Even where humans learn from demonstration, in sport or craft, the demonstration is a hint and the hours of practice are the training. The recipe is the same across species: a prior, practice, and the world as verifier. It should be the same for robots.

3.2 Classical physics is complete

Everything a robot touches obeys classical mechanics: rigid-body dynamics, contact, friction, elasticity, fluids at everyday scale, and the electromagnetics and thermodynamics of its own actuators. At robot scale the laws are complete. What is unknown is the state.

Games are the precedent. The one domain where AI became superhuman from self-generated experience is the one whose rules were fully known: AlphaZero learned from self-play alone. Physics is the game whose rules are classical mechanics; the simulator is the board. Closing the distance between a simulator and the world is the path to superhuman physical intelligence.

That turns the corpus problem into a compute problem. The corpus of human text is finite and has to be found; experience can be synthesized from first principles at whatever scale compute allows. Physical intelligence then follows the curve that carried language models, with a manufactured corpus in place of a crawled one. One GPU runs physics at 10⁵–10⁶ steps per second, about a decade of embodied experience per day and the output of a fleet of several thousand robots; one rack outproduces any fleet ever built.

3.3 Inductive biases and learning objectives grounded in physics

Inductive biases grounded in physics can take physical intelligence beyond the human level: better agility, dexterity, and compliance. Every constraint physics imposes is structure the model gets for free. In the same terms, a model can design task-specific constraints and objectives of its own and search for the best policy against them.

Agility is the one objective imitation cannot reach, even in principle. A policy optimized for agility is also a more intelligent one. A slow policy finishes most tasks by waiting for feedback at every step. At speed, feedback arrives too late; the policy has to predict what the world will do before it does it, and that takes a better model of the world's dynamics. Speed measures how much of the world a policy can predict. Making agility the explicit objective is both a product goal and a bet on method: agility can only be optimized through experience.

Dexterity is agility at the point of contact. Every make and break of contact is a discontinuity in the dynamics, which is where demonstrations are thinnest and prediction matters most.

Compliance is an objective and an insurance policy. Biological muscle is compliant; stiff position control is what makes robots dangerous and brittle at contact. As an objective, compliance penalizes impact forces and torque discontinuities. As insurance, a compliant policy is robust by construction: it absorbs the contacts the simulator got wrong, the disturbances of an unstructured environment, and contact with people, which is what makes a robot safe to work beside.

Energy efficiency is the universal regularizer, and it is why animals move the way they do. Every animal minimizes metabolic cost subject to the task; that optimization produces natural gait, smooth reaching, and passive stability.

Prediction is how a policy learns dynamics, and what it predicts decides what it learns. The state control cares about is low-dimensional and predictable: pose, velocity, contact. Pixels are high-dimensional and full of what is unpredictable and irrelevant: lighting, texture. A latent predictive objective spends capacity on dynamics and produces representations that encode what control needs. The same latent model is what planning runs on.

3.4 Embodied intelligence is a full-stack problem

Embodied intelligence is a full-stack problem: a model is only as agile as the weakest link beneath it, whether sensor rate, compute latency, or mechanical bandwidth. A poor choice at the hardware or control layer is a hard limit on the model and application layers.

The constraints start with the body. Actuator bandwidth and limb inertia bound agility; the hand's degrees of freedom and tactile sensing bound dexterity; backdrivable joints decide whether compliance is possible at all. In most cases, those decisions are made before training starts, and the model inherits them.

Latency is the clearest case. Robots have a ~20-fold advantage over biology: an electronic control loop closes in 1 ms; a spinal reflex takes 20–50 ms. Most stacks give it away to a 100 ms model or a round trip to a datacenter, and the robot ends up slower than a reflex.

4. A System 1 model for Embodied Cognitive Intelligence

A note on where we are. We are building the System 1 layer for physical agents, and the platform it runs on. We are working with a small group of partners, the research labs, developers, and businesses at the frontier of physical AI building for humanoid robots, to put useful agents into the physical world.

Coming soon. For early access: cogence.company