Skild AI's S1 Robot Learns Factory Tasks From Single Video Demo
The foundation model uses in-context learning to execute complex, multistep work without retraining, reaching $100M revenue run rate in under a year.

One Video, No Retraining Required
Industrial robots typically require extensive reprogramming when production lines change or new products arrive. Skild AI's S1 foundation model takes a fundamentally different approach: operators record a single video demonstration of a desired task, and the robot executes it without updating its weights or undergoing task-specific training.
The model, launched last week, uses a technique called in-context learning to interpret demonstrated intent, objects and sequences, then maps them into actions for the physical robot. According to details first reported by NVIDIA, S1 can perform unfamiliar tasks lasting up to 10 minutes, including plant potting, pancake making and kit assembly—work that can span dozens of manipulation steps in sequences the robot hasn't previously performed.
In testing on new, multistep tasks, S1 succeeded approximately 66% of the time at each step, compared with 9% for a comparable AI system. Skild AI estimates that one short video demonstration provides roughly the same learning value as 380 hands-on training examples, which could require 50 to 100 hours of manual data collection.
Why it matters
The ability to teach robots new tasks through video demonstration rather than extensive retraining addresses a core bottleneck in industrial automation: adaptability. As manufacturing environments become more dynamic and product cycles shorten, the traditional model of fixed-function robots becomes increasingly expensive to maintain. S1's approach could significantly reduce the time and expertise required to reconfigure production lines, making automation accessible to a broader range of applications and smaller manufacturers.
From Lab to Factory Floor
Skild AI reached a $100 million annual revenue run rate just 10 months after its first commercial deployment. The company has established more than 60 deployment partnerships spanning manufacturing, logistics, inspection, security and food preparation.
One deployment with Foxconn and NVIDIA involves dual-arm manipulators performing high-precision assembly of NVIDIA Blackwell systems. In a demonstrated workflow, robots install busbars and limit blocks, fasten 16 screws and adapt to disturbances across multistep tasks requiring precise motion, contact-aware control and error recovery.
In one plant-potting test, the Skild AI team moved from recording a demonstration to autonomous hardware execution in 11 minutes. The model can adjust when objects move, recover from errors and combine skills in previously unprogrammed sequences.
Technical Foundation
Skild built S1 using NVIDIA AI infrastructure, including Isaac Lab for reinforcement learning and Isaac Sim for physically based virtual environments. The company uses NVIDIA Cosmos world foundation models to diversify training data and convert video into structured descriptions, while Cosmos Curator handles annotation, filtering and organization at scale.
The Newton physics engine within Isaac Lab helps model forces, contact, collision and pressure to reduce the simulation-to-reality gap. Skild and NVIDIA are jointly developing GPU-accelerated simulation solvers for modeling how robots physically touch, grip and manipulate solid objects, which will be made available to developers as part of Newton.
"Learning by experience, and not preprogramming, is the step change that has happened in robotics," said Deepak Pathak, cofounder and CEO of Skild AI.
These details were first reported by NVIDIA.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call