Language-guided robot manipulation
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
A compact latent interaction policy that learns action-grounded, future-predictive visual representations.
Abstract
Compact policies grounded in latent dynamics.
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. SLIM is a compact 0.5B-parameter latent interaction policy that focuses this computation on observations, actions, and the transitions induced by actions.
SLIM learns action-grounded predictive latents through self-supervised masked trajectory prediction, alternating between action reconstruction and future-latent prediction. A compact Mixture-of-Transformers backbone then performs language-conditioned flow-matching action generation. Across simulation benchmarks and real-world evaluation, SLIM matches or exceeds representative VLA and world-action-model baselines with fewer parameters, lower deployment latency, and substantially lower GPU memory usage.
Method
Masked trajectory learning for action-grounded latent interaction.
Results
Strong performance with a compact policy backbone.
Benchmark comparisons
Competitive across robustness and long-horizon control.
Representative baselines from the current paper tables. Bar length encodes benchmark performance; the right column reports model parameters.
LIBERO-plus
Zero-shot overall success (%)
CALVIN ABC to D
Average sequence length (max 5)
Real-world evaluation
Robust multi-task manipulation on a physical robot.
Real-world manipulation
See the policy in action.
Explore five tasks under nominal conditions, added distractors, changed lighting, and changed background textures.
01Stacking three plates↗
02Placing a carrot into a bowl↗
03Taking toast out of a toaster↗
04Stacking blocks↗
05Wiping a whiteboard↗These are 20 selected successful executions, not all evaluation trials or an estimate of success rate. All clips show SLIM and are encoded at 3× speed, with no audio. The videos do not indicate real-time execution speed or model inference latency. Lighting clips contain changing colored light patterns.
Task 01
Stacking three plates
Stack three plates into a single stack.
Nominal
Selected success · 3×Distractor
Selected success · 3×Lighting
Selected success · 3×Background
Selected success · 3×Task 02
Placing a carrot into a bowl
Pick up the carrot and place it into the bowl.
Nominal
Selected success · 3×Distractor
Selected success · 3×Lighting
Selected success · 3×Background
Selected success · 3×Task 03
Taking toast out of a toaster
Take the toast out of the toaster and place it onto the plate.
Nominal
Selected success · 3×Distractor
Selected success · 3×Lighting
Selected success · 3×Background
Selected success · 3×Task 04
Stacking blocks
Complete two stacking actions to form the full block stack.
Nominal
Selected success · 3×Distractor
Selected success · 3×Lighting
Selected success · 3×Background
Selected success · 3×Task 05
Wiping a whiteboard
Pick up the eraser and wipe the whiteboard.
Nominal
Selected success · 3×Distractor
Selected success · 3×Lighting
Selected success · 3×Background
Selected success · 3×Ablations
Trajectory objectives and EMA stabilize OOD performance.
Stage-1 masked trajectory learning improves both LIBERO-plus robustness and CALVIN long-horizon composition over Stage-2-only training.
An IDM:FDM loss ratio of 0.125:1 gives the strongest joint result. EMA raises LIBERO-plus success from 66.82% to 77.45% and CALVIN average length from 4.382 to 4.556.
Analysis
Trajectory learning focuses action evidence on manipulation-relevant regions.
Citation
BibTeX
@article{wang2026slim05b,
title = {SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation},
author = {Jingkai Wang and Zihan Tang and Gu Zhang and Mingyu Cao and Jiapeng Chen and Jingjiao Zhao and Xiansheng Chen and Pengwei Wang and Lemao Liu and Dejing Dou},
journal = {arXiv preprint arXiv:2608.09771},
year = {2026},
url = {https://arxiv.org/abs/2608.09771}
}