Last week we hosted Deepak Pathak, the founder and CEO of SkildAI, building a general brain for robots.
Here are our main takeaways from our dinner:
1/ Intelligence is decision-making. Everything else is clock speed.
Body control runs at 500–1,000 Hz. Vision at ~10 Hz. Speech at 2–5 Hz. Math at ~0.001 Hz. It’s the same machinery all the way down.
Sit with that for a second: the way our entire industry splits itself into vision, language, and controls teams may be an artifact of our tooling rather than the problem itself.
2/ We’re building intelligence top-down. Evolution built it bottom-up.
Evolution spent billions of years on movement and hundreds of millions on vision and hands. Language appeared only 10,000–50,000 years ago, and the brain barely changed to accommodate it.
That makes language a product of intelligence, rather than its source.
Deepak’s question to the room: Where is the evidence that scaling the last sliver of that timeline will fill in everything underneath it?
Nobody had an answer.
3/ “General” only means something if it’s falsifiable.
Anyone can put the word in a deck. His definition: any robot, any task, one brain.
They cut a quadruped in half, and both halves keep walking. An arbitrary humanoid walks within seconds.
That’s a claim you can break in public - which is why I trust it more than the adjective.
4/ Omni-bodied is a data strategy wearing a research costume.
Robotics has no internet to scrape. The only large-scale data will come from deployed fleets, and hardware will never converge on a single shape.
There are dozens of TV makers, phone makers, and car makers. Bet on one form factor, and you quietly cap your data ceiling forever.
5/ There are four data sources, and each one’s weakness is another’s strength.
Teleoperation: pristine, tiny.
Gloves: middling on both.
Simulation: infinitely scalable, with a reality gap.
Human video: diverse, free, and never robot data.
Anyone selling you on one source is selling you their own strength.
6/ Use video and simulation for pre-training. Use teleoperation for post-training.
It’s the same shape as LLM training.
His read on the field is that most teams run this backward, burning their scarce, clean data on pre-training - where scale matters more than polish.
7/ Watching Federer for 100 hours will not make you play like Federer.
Video gives you the what. It never gives you the forces, the momentum, or the forehand-to-backhand transition.
Those have to be practiced - and practice happens millions of times in simulation.
This is the cleanest argument I’ve heard against treating egocentric video as a standalone answer.
8/ Deployment is the third pillar, and it’s the one nobody pays enough attention to.
Data. Architecture. Deployment.
Deepak named the third as the single biggest reason this wave could fail.
On the objection that “deployment isn’t venture-scale,” his response was blunt: “A very convenient argument to give.”
Skip deployment, and you lose the only thing that tells you you’re wrong.
9/ Robotics has no working evaluations, so revenue is the scoreboard.
You cannot find two robotics papers evaluated on the same benchmark. There’s no arena, shared setup, or honest baseline.
That leaves money paid as the only real signal - and across the field, that number still rounds to zero.
He called it a convenient shared incentive to remain unmeasured. Hard to argue.
10/ Customers don’t care about accuracy. Zero.
They assume accuracy. That’s what they’re paying for.
What they buy is cycle time and cost.
The metric that earns you a paper and the metric that earns you a purchase order barely overlap. Closing that gap took his team a year of actual deployment.
11/ Evolution did the pre-training. Childhood is post-training.
His curiosity-driven exploration work solved Atari without rewards - and never worked on a real robot.
His theory now: children reach the same milestones whether they grow up in a mansion in LA or a village in India, so most of that learning was already paid for.
That gives robotics permission to be spectacularly data-inefficient because we’re paying the evolution bill from scratch.
12/ Language stops working at about 30 seconds.
“Pass the sandwich” - perfect.
A three-minute unfamiliar task - useless, because you’d have to dictate joint angles.
Beyond that horizon, the interface becomes demonstration.
As he put it: If words were enough, handymen would be out of work.
One more thought on focus, from his advisor:
The world max-pools you to your single highest achievement.
Decide what your max pool is. Spend your time there, and let everything else fall into place.
Discussion about this post
No posts

