Skild's S1 AI Learns to Flip Pancakes From a Single Video Demo

We’ve been promised the “iPhone moment” for robotics so many times it’s become a running gag in the industry. But every once in a while, something comes along that forces even the most jaded among us to sit up and pay attention. This time, it involves a robot, a spatula, and a pancake. Skild AI, the absurdly well-funded startup now valued at over $14 billion, just dropped S1, a foundation model that can learn a complex, 10-minute task from watching a single video. No fine-tuning, no months of data collection, no army of teleoperators. Just watch and do.

The first time their robot flawlessly flipped a pancake, the Skild team did what any self-respecting engineer would: they assumed it was a bug. They spent two days scouring their million-hour pre-training dataset, convinced that some rogue breakfast-making data had snuck in. They found nothing. The robot had generalized, inferring the out-of-distribution task of “flipping” from a single video prompt. This wasn’t just mimicry; it was a spark of physical common sense.

Video thumbnail

The Tyranny of Fine-Tuning is Over

For years, deploying a robot on a new task has been a soul-crushing exercise in data collection and fine-tuning. A company might spend 50 to 100 hours gathering data and tweaking a model just to get it to perform one new job reliably. This is the expensive, time-consuming bottleneck that has kept general-purpose robots confined to research labs. Skild’s S1 proposes a radical alternative rooted in the same principle that made ChatGPT a global phenomenon: in-context learning.

Instead of laboriously re-training a model, you simply show it an example in its prompt window. For an LLM, that’s a text example. For S1, it’s an egocentric video of a human performing the task. The model is pre-trained on a massive diet of internet videos, simulations, and real-world teleoperation data, allowing it to build a generalized understanding of physical interaction. From there, a single demonstration is enough to unlock a new skill, whether it’s potting a plant, assembling a hygiene kit, or making pour-over coffee.

“This concept is called in-context learning, and it’s what defined the leap from research projects like GPT-1 and GPT-2 to revolutionary products like ChatGPT,” the company stated in its announcement video. “You place a demonstration in the context window… and it emits robot actions to complete the task.”

This approach doesn’t just work; it dramatically outperforms the old way. For tasks already familiar to the model, S1 performs on par with traditional Vision-Language-Action (VLA) models. But for novel, never-before-seen tasks, the difference is staggering.

The Scaling Laws of Physical Intelligence

The real magic happens at scale. Skild released a chart that should be pinned to the wall of every robotics lab. It shows that as pre-training data increases, S1’s ability to handle novel tasks via in-context learning skyrockets, while language-prompted VLAs show diminishing returns.

A chart comparing the success rate of Skild's S1 in-context learning against language-prompted VLA models on unseen tasks, showing S1's performance scaling exponentially with more pre-training hours.

After 100,000 hours of pre-training, S1 achieves a 66% success rate on out-of-distribution tasks. The language-prompted VLA? A paltry 9%. That’s not an incremental improvement; it’s a phase change. It suggests that, just as with large language models, there are predictable scaling laws for robotics. Pour in enough data, and entirely new capabilities emerge. The model doesn’t just get better; it gets smarter.

This intelligence isn’t brittle, either. Skild demonstrates S1 adapting to perturbations, improvising after making a mistake, and even correcting for errors made by the human in the demonstration video. It can handle objects being shifted or swapped, suggesting a robust understanding that goes beyond simple path replay.

From Lab Demo to Industrial Reality

If this were just another university project, it would be interesting. But Skild AI is a commercial juggernaut backed by the likes of SoftBank, NVIDIA, and Jeff Bezos’s Bezos Expeditions, with a $14 billion valuation and a reported $30 million in revenue generated in just a few months last year. This isn’t a demo; it’s a product announcement from a company with existing OEM partnerships and hundreds of robots already in the wild.

This is the next logical step in their mission to build a universal Skild's AI Brain Controls Any Robot Body . We’ve seen them teach robots to Skild AI Teaches Robots to Cook by Making Them Watch YouTube before, but S1 represents a fundamental shift in the deployment paradigm. The goal is to change the economics of automation, making it accessible to small businesses, hospitals, and grocery stores—organizations that can’t afford a dedicated robotics PhD to retrain their machines.

Of course, there are caveats. The benchmarks are internal, and Skild has yet to release a technical paper, API, or the model weights. But the videos are compelling, and the business momentum is undeniable. The era of one-robot, one-task may finally be drawing to a close. With S1, Skild is arguing that the future of robotics is less about programming and more about showing. And it all started with a pancake.