LLMOps, Observability & Evals
Once an LLM application is in production, this category is what tells you whether it’s actually working — tracing tools (LangSmith, Helicone), evaluation platforms (Braintrust, Galileo), and monitoring (Arize, Weights & Biases) built specifically for the failure modes unique to LLM systems: hallucination, prompt regressions, and silent quality drift that a normal APM tool won’t catch. As more AI products move from demo to production, this category has grown from a niche into something builders increasingly treat as non-optional. GROUNDING tracks new eval methodologies, benchmark releases, and the LLM-as-judge techniques (used by GROUNDING’s own scoring engine, documented in the Methodology page) that are reshaping how model output gets measured.
At a glance
- 1 tracked companies (1 hand-curated)
- Most recently updated: DagsHub (2026-06-04)
Most actively covered
- DagsHub ★
FAQ
What is the LLMOps, Observability & Evals category?
Builds monitoring, evaluation, or tracing tools for LLM systems in production.
How many companies are in this category?
GROUNDING currently tracks 1 companies in LLMOps, Observability & Evals, 1 of which have a hand-curated profile.
How current is this hub?
The most recently updated entry is DagsHub, last mentioned 2026-06-04; the radar refreshes hourly.