Skip to content

LLMOps and evals

Peak · emerged March 2023 · high July 2026 · +5% on the month before · 5 mentions on Trending in 14 days

Measuring, tracing and monitoring model-backed software: eval suites, LLM observability and prompt management. Named by OpenAI: OpenAI Evals: a framework for evaluating LLMs (Mar 14, 2023). Also called LLMOps, LLM evals, Evals, LLM observability, LLM evaluation, LLM-as-a-judge.

Index 0 to 100, monthly
050100202420252026Origin

Signals

SignalReads as
Hacker News stories, last 12 months227 stories, +67% on the 12 months beforeAdoption
arXiv papers, last 12 months1,374 papers, +85% on the 12 months beforeAdoption
Product releases naming it8 since Sep 1, 2025Adoption
GitHub repositories4,191Adoption

Companies

Releases

8 releases since Sep 1, 2025
Name
Langfuse 4.44.0Langfuse · Sep 24, 2026 yesterday
Langfuse 4.43.0Langfuse · Sep 23, 2026 2 days ago
Run Terminal-Bench and other Harbor evals on Vercel SandboxVercel · Sep 17, 2026 8 days ago
What's new in ClickStack - Aug ’26ClickHouse · Sep 16, 2026 9 days ago
LLM evaluation suites in Pipeline Builder are generally availablePalantir · Sep 10, 2026 2 weeks ago
W&B Weave 0.53.7CoreWeave · Aug 27, 2026 4 weeks ago
🧪 Introducing EvalsHex · Aug 4, 2026 7 weeks ago
Legacy Claude Console Workbench sunsets on August 17, 2026Anthropic · Jul 17, 2026 2 months ago

GitHub projects

4,191 repositories
NameStars
confident-ai/deepevalThe LLM Evaluation Framework · Aug 10, 2023 3 years ago18,433
open-compass/opencompassOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, acros · Jun 15, 2023 3 years ago7,472
langwatch/langwatchThe platform for LLM evaluations and AI agent testing · Sep 9, 2023 3 years ago4,872
huggingface/evaluation-guidebookSharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and · Oct 9, 2024 1 year ago2,147
tjunlp-lab/Awesome-LLMs-Evaluation-PapersThe papers are organized according to our survey: Evaluating Large Language Models: A Comprehensive Survey. · Oct 29, 2023 2 years ago810
allenai/olmesReproducible, flexible LLM evaluations · Nov 2, 2024 1 year ago395
MovementStage
AI agentsModels that plan and act in steps, calling tools and other services until a task is done.Peak
Synthetic dataTraining and evaluation data generated by models or simulators instead of collected from people.Declining
Context engineeringDeciding what goes into a model's context window, and in what order: instructions, retrieved documents, tools and memory.Plateau

History

DateWhat changed
Sep 24, 2026Tracking started: emerged March 2023, peak, 6 other names recorded

Sources: Curve: Hacker News story titles, Wikipedia pageviews and arXiv papers by month, refreshed Sep 24, 2026. Related entities from the fru.dev sites' public APIs, matched by name. How stages work.

Trending terms by email

Monday mornings: the week's top trending terms in data, tech and AI.

Double opt-in. Unsubscribe any time.