Skip to content

Multimodal AI

Peak · emerged January 2021 · high June 2026 · +3% on the month before · 56 mentions on Trending in 14 days

Models that read and write more than text: images, audio and video in the same model. Named by Radford et al., OpenAI: CLIP: Connecting text and images (Jan 5, 2021). Also called Multimodal models, Vision language models, VLM, Multimodal LLMs.

Index 0 to 100, monthly
03673202120222023202420252026Origin

Signals

SignalReads as
Hacker News stories, last 12 months284 stories, -9% on the 12 months beforeAdoption
arXiv papers, last 12 months4,122 papers, +58% on the 12 months beforeAdoption
Wikipedia views, August 20265,847, +47% on a year beforeAdoption
Product releases naming it33 since Sep 1, 2025Adoption
GitHub repositories67,013Adoption

Companies

Releases

33 releases since Sep 1, 2025
Name
Docling 2.130.0IBM · Sep 22, 2026 3 days ago
GLM 5.3 FlashX now available on AI GatewayVercel · Sep 18, 2026 7 days ago
Multimodal AI_SUMMARIZE for automatic theme summarization (Public Preview)Snowflake · Sep 14, 2026 11 days ago
Amazon Bedrock Managed Knowledge Base now supports multimodal embeddings for video, audio, and image content with TwelveLabs Marengo 3.0AWS · Sep 11, 2026 2 weeks ago
OpenAI GPT-6 Astra is now available on Unity GatewayDatabricks · Sep 4, 2026 3 weeks ago
Docling 2.125.0IBM · Sep 3, 2026 3 weeks ago
Qwen 3.8 27B available on public endpoints qwen-3.8-27b is now available on Cerebras public endpointsCerebras · Sep 3, 2026 3 weeks ago
SDK releases: Python SDKElevenLabs · Aug 31, 2026 3 weeks ago

Models

Papers

Name
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesMeta · Aug 6, 2026 7 weeks ago
Gemma 4 Technical ReportGoogle DeepMind · Jul 2, 2026 2 months ago
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of contextGoogle DeepMind · Mar 8, 2024 2 years ago
GPT-4 Technical ReportOpenAI · Mar 15, 2023 3 years ago

GitHub projects

67,013 repositories
NameStars
huggingface/transformers🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, · Oct 29, 2018 7 years ago166,616
Mintplex-Labs/anything-llmStop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience · Jun 4, 2023 3 years ago66,428
bojieli/ai-agent-book《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码 · Sep 9, 2025 1 year ago50,784
bytedance/UI-TARS-desktopThe Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra · Jan 19, 2025 1 year ago39,109
sgl-project/sglangSGLang is a high-performance serving framework for large language models and multimodal models. · Jan 8, 2024 2 years ago36,408
deepset-ai/haystackOpen-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agen · Nov 14, 2019 6 years ago26,591
MovementStage
Physical AIModels that perceive and act in the physical world: robots, self-driving and the simulators that train them.Rising
World modelsModels that learn how an environment behaves and can simulate it, for games, video and training robots.Rising
Reasoning modelsModels trained with reinforcement learning to think step by step, spending more compute at answer time.Plateau

History

DateWhat changed
Sep 24, 2026Tracking started: emerged January 2021, peak, 4 other names recorded

Sources: Curve: Hacker News story titles, Wikipedia pageviews and arXiv papers by month, refreshed Sep 24, 2026. Related entities from the fru.dev sites' public APIs, matched by name. How stages work.

Trending terms by email

Monday mornings: the week's top trending terms in data, tech and AI.

Double opt-in. Unsubscribe any time.