Skip to content

LLMOps, Observability & Evals

Once an LLM application is in production, this category is what tells you whether it’s actually working — tracing tools (LangSmith, Helicone), evaluation platforms (Braintrust, Galileo), and monitoring (Arize, Weights & Biases) built specifically for the failure modes unique to LLM systems: hallucination, prompt regressions, and silent quality drift that a normal APM tool won’t catch. As more AI products move from demo to production, this category has grown from a niche into something builders increasingly treat as non-optional. GROUNDING tracks new eval methodologies, benchmark releases, and the LLM-as-judge techniques (used by GROUNDING’s own scoring engine, documented in the Methodology page) that are reshaping how model output gets measured.

At a glance

  • 1 tracked companies (1 hand-curated)
  • Most recently updated: DagsHub (2026-06-04)

Most actively covered

FAQ

What is the LLMOps, Observability & Evals category?

Builds monitoring, evaluation, or tracing tools for LLM systems in production.

How many companies are in this category?

GROUNDING currently tracks 1 companies in LLMOps, Observability & Evals, 1 of which have a hand-curated profile.

How current is this hub?

The most recently updated entry is DagsHub, last mentioned 2026-06-04; the radar refreshes hourly.

1 page with this tag.