Skild AI Unveils S1 Robot Model That Learns Tasks From Video
The company says S1 uses in-context learning, an approach analogous to prompting in large language models, allowing an operator to demonstrate a task on video and have the robot interpret the demonstration and reproduce the intended behavior.
Skild says:“Show it a video of a task, short or long, seen or unseen, and it executes.”
Rather than specifying tasks primarily through language, S1 takes a video demonstration as its prompt. The model then interprets the demonstrator's intent and translates it into actions appropriate to the robot and its environment.
According to Skild, the same model weights can be used across familiar and previously unseen behaviors without additional fine-tuning. S1 is built using Nvidia AI infrastructure for large-scale training.
The company demonstrated S1 performing previously unseen tasks including plant potting, pancake cooking, pour-over coffee making and kit assembly. The tasks involve dozens of manipulation steps and can run for up to 10 minutes from a single visual demonstration.
In one plant-potting experiment, Skild says only 11 minutes elapsed between recording the human demonstration and S1 beginning to execute the task autonomously on robot hardware.
Skild says:“With conventional workflows, deploying a policy for a new manipulation task begins with hours of teleoperation and a task-specific fine-tuning run. Teaching S1 something new takes minutes.”
The company also compared in-context learning with a conventional language-prompted vision-language-action model. On previously unseen tasks, Skild reports that S1 achieved a 66 percent success rate after pre-training on 100,000 hours of data, compared with 9 percent for the language-prompted model.
Skild says a single video demonstration produced performance roughly equivalent to a conventional model receiving around 380 post-training demonstrations. Collecting that amount of training data for the long-horizon tasks took between 50 and 100 hours of teleoperation.
The company also reports that S1 can respond to changes that were not present in the original demonstration, including objects being moved or substituted, and can sometimes recover from mistakes during execution.
Skild argues that this ability to learn rapidly will become increasingly important as robots move from controlled environments into applications where tasks and conditions continually change.
The company says:“If every change requires repeated iterations of data collection, policy fine-tuning, and validation, robots will never keep pace with the environments in which they operate. Instead, robots should acquire new behaviors the same way people do: by observing a single demonstration.”
S1 is already being used with Skild AI's commercial partners, according to the company.
Legal Disclaimer:
MENAFN provides the
information “as is” without warranty of any kind. We do not accept any
responsibility or liability for the accuracy, content, images, videos,
licenses, completeness, legality, or reliability of the information
contained in this article. If you have any complaints or copyright issues
related to this article, kindly contact the provider above.

Comments
No comment