Comet
ML experiment tracking and LLM observability platform, including Opik for evaluating LLM apps.
Comet is annexing AI cost governance from the observability side.
◆Recent moves
- 1d ago
LLM Model Selection: How to Pick the Right Model for Every Agentic Task
An explainer on matching models to agentic tasks instead of defaulting every tool call to the flagship. It rehearses the cost argument that Cost Intelligence productizes, but ships nothing itself.
View source ↗ - 1d ago
Best LLM Observability Tools of 2026: Top Platforms & Features
A roundup of LLM observability platforms for 2026 — category-defining SEO content aimed at buyers evaluating Opik against alternatives. No product change.
View source ↗ - 11d ago
I Built a RAG Pipeline for F1 Team Radio, Then Made It Grade Itself
A developer-relations build log: a RAG pipeline over F1 team radio, then graded with Opik's own evaluation tooling. It demonstrates evaluation-driven development on a fun dataset rather than reporting a release.
View source ↗ - 26d ago
One Prompt, 24 Versions: How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform
A customer story on Digibee versioning one prompt 24 times through Opik to power its integration platform. Useful proof of the prompt-experimentation workflow, but a case study, not a change.
View source ↗ - 29d ago
Beyond the Single Trace: How We Built Agent Diagnostics for Opik
Agent Diagnostics ships in Opik, moving debugging from reading one trace at a time to surfacing misbehavior across many. It is the clearest recent product move in the feed and fits the shift from passive tracing to active diagnosis.
View source ↗ - 1mo ago
What Is an Agent Harness? The Layer That Makes AI Agents Actually Work
A concept piece naming the retry logic and prompt assembly most teams build ad hoc as the agent harness layer. Category framing that positions Opik, with nothing shipped.
View source ↗