Papers
arxiv:2609.36691

Video2Skill: From Streaming Experience to Reusable Embodied Skills

Published on Sep 29
ยท Submitted by
Jianshu Zhang
on Oct 6
Authors:
,
,
,
,
,
,
,
,

Abstract

Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a model watches videos in sequence and maintains a persistent skill library that shapes its later decisions. To systematically measure this ability, we introduce Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: (i) locating manipulation events in time, (ii) grouping events of the same transformation, and (iii) deciding when to reuse an existing skill or create a new one. Across 19 open-source VLMs, many models group events at near-chance level, and scale does not consistently help. Their errors depend on how perception and library updates are coupled: joint models merge distinct transformations into one skill, while models that update the library from text descriptions duplicate recurring ones. Supervised fine-tuning, including our counterfactual library-state rebalancing (CLaRe), improves grouping but exposes a deeper bottleneck: trained models consolidate familiar skills yet rarely expand the library. Their libraries stall below half the reference size, and transformations unseen in training are located in time but almost never given a new skill. Recognizing when existing skills are insufficient thus emerges as the central challenge.

Community

Paper submitter

๐Ÿค– Reusing skills is one possible way for embodied agents to generalize. Manipulation varies but shares a few skills: wiping a table or a window is one skill. Many works give agents skills to plan with. But where do these skills, and their data, come from?

๐Ÿ“ผ Today, VLMs mostly label videos one at a time, and nothing carries over. People don't learn that way. As experience streams in, we spot skills we know, add new ones, and reuse them later. This streaming setting matters, but it has received little attention.

๐Ÿง  VLMs may become the brains of embodied agents. So we ask: from streaming experience, can they find reusable skills and keep one consistent skill library?

๐Ÿงต Video2Skill: From Streaming Experience to Reusable Embodied Skills

Why it matters:
๐Ÿ“ˆ It can label much more data for training agents.
๐Ÿ” It is planning in reverse. If a model can't find skills in what it has seen, how can it plan with them for something new?

๐Ÿ“„ Paper: https://arxiv.org/abs/2609.36691
๐ŸŒ Project: https://andyzworks.github.io/video2skill/
๐Ÿค— Data: https://huggingface.co/datasets/Sterzhang/video2skill-bench

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.36691
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.36691 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.36691 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.