Multimodal foundation models

Established

Single models that jointly understand images, text, and audio, unifying perception and language in one backbone.

Trend score: 46

Trend over time

Monthly assigned-paper counts from arXiv (scored: trend_v1, confidence: High). History is still short, so scores are preliminary and sharpen over time.

Evidence papers — 50 most recent of 7241