Vision-language-action models

Established

Models that consume images and instructions and emit robot actions, bringing foundation-model scaling to embodied control.

Trend score: 49

Trend over time

Monthly assigned-paper counts from arXiv (scored: trend_v1, confidence: High). History is still short, so scores are preliminary and sharpen over time.

Evidence papers — 50 most recent of 923