We’ve been promised the “iPhone moment” for robotics so often it’s become the industry’s longest-running gag. But every once in a while, something comes along that forces even the most cynical tech-watcher to sit up and take notice. This time, it involves a robot, a spatula, and a pancake. Skild AI, the absurdly well-funded startup now valued at over $14 billion, has just unveiled S1: a foundation model capable of learning a complex, 10-minute task simply by watching a single video. No fine-tuning, no months of painstaking data collection, and no army of teleoperators required. Just watch and learn.
When their robot first flipped a pancake without a hitch, the Skild team did what any self-respecting engineer would: they assumed it was a bug. They spent two days trawling through their million-hour pre-training dataset, certain that some rogue breakfast-making data had slipped through the cracks. They found nothing. The robot had generalised, inferring the “out-of-distribution” task of flipping from a single video prompt. This wasn’t mere mimicry; it was a genuine spark of physical common sense.

The Death of the Fine-Tuning Slog
For years, teaching a robot a new trick has been a soul-destroying exercise in data collection. A company might spend 50 to 100 hours gathering data and tweaking a model just to get it to perform one new job reliably. This is the expensive, time-consuming bottleneck that has kept general-purpose robots trapped in research labs. Skild’s S1 proposes a radical alternative rooted in the same principle that made ChatGPT a global phenomenon: in-context learning.
Instead of laboriously re-training a model, you simply show it an example in its prompt window. For a Large Language Model (LLM), that’s a text example. For S1, it’s an egocentric video of a human performing the task. The model is pre-trained on a massive diet of internet videos, simulations, and real-world data, allowing it to build a generalised understanding of physical interaction. From there, a single demonstration is enough to unlock a new skill, whether it’s potting a plant, assembling a hygiene kit, or brewing a pour-over coffee.
“This concept is called in-context learning, and it’s what defined the leap from research projects like GPT-1 and GPT-2 to revolutionary products like ChatGPT,” the company stated in its announcement video. “You place a demonstration in the context window… and it emits robot actions to complete the task.”
This approach doesn’t just work; it leaves the old way in the dust. For tasks the model already knows, S1 performs on par with traditional Vision-Language-Action (VLA) models. But for novel, never-before-seen tasks, the difference is night and day.
The Scaling Laws of Physical Intelligence
The real magic happens at scale. Skild released a chart that should be plastered on the wall of every robotics lab in the country. It shows that as pre-training data increases, S1’s ability to handle novel tasks via in-context learning skyrockets, while language-prompted VLAs hit a ceiling of diminishing returns.

After 100,000 hours of pre-training, S1 achieves a 66% success rate on out-of-distribution tasks. The language-prompted VLA? A measly 9%. That’s not an incremental improvement; it’s a phase change. It suggests that, just as with LLMs, there are predictable scaling laws for robotics. Pour in enough data, and entirely new capabilities emerge. The model doesn’t just get better; it gets smarter.
This intelligence isn’t brittle, either. Skild has demonstrated S1 adapting to disruptions, improvising after a mistake, and even correcting for errors made by the human in the demo video. It can handle objects being shifted or swapped, suggesting a robust understanding that goes far beyond simple path replay.
From Lab Demo to Industrial Reality
If this were just another university project, it would be a neat curiosity. But Skild AI is a commercial juggernaut backed by the likes of SoftBank, NVIDIA, and Jeff Bezos’s Bezos Expeditions. With a $14 billion valuation and a reported $30 million in revenue generated in just a few months last year, this isn’t a demo—it’s a product announcement from a company with existing OEM partnerships and hundreds of robots already deployed in the wild.
This is the next logical step in their mission to build a universal Translation not available (en-gb) . We’ve seen them teach robots to Translation not available (en-gb) before, but S1 represents a fundamental shift in how these machines are deployed. The goal is to change the economics of automation, making it accessible to small businesses, hospitals, and supermarkets—organisations that can’t afford a dedicated robotics PhD to retrain their machines.
Of course, there are caveats. The benchmarks are internal, and Skild has yet to release a technical paper, API, or the model weights. But the footage is compelling, and the business momentum is undeniable. The era of the “one-robot, one-task” specialist may finally be drawing to a close. With S1, Skild is betting that the future of robotics is less about programming and more about showing. And it all started with a pancake.
