Model Hosting, Inference & API Gateways
These companies don’t train frontier models — they take open-weight and licensed models and serve them as a hosted API, competing on latency, cost-per-token, and uptime rather than capability. Together AI, Fireworks, Replicate, OpenRouter, Baseten, and DeepInfra are the category’s most-tracked names, and OpenRouter in particular has become a de facto price/latency comparison layer across dozens of models at once. For builders this category is often the pragmatic default: you get frontier-adjacent capability without operating GPUs yourself, and can switch providers if one gets slow, expensive, or deprecates a model you depend on. GROUNDING tracks new model availability on these platforms, pricing changes, and outages/deprecations that force a builder to migrate.
At a glance
- 3 tracked companies (3 hand-curated)
- Most recently updated: Groq (2026-06-05)
Most actively covered
- OpenRouter ★
- vLLM ★
- Groq ★
FAQ
What is the Model Hosting, Inference & API Gateways category?
Serves third-party models via a hosted API rather than training its own.
How many companies are in this category?
GROUNDING currently tracks 3 companies in Model Hosting, Inference & API Gateways, 3 of which have a hand-curated profile.
How current is this hub?
The most recently updated entry is Groq, last mentioned 2026-06-05; the radar refreshes hourly.