← Back to all sparks
C

Comet

AI-ASSISTANTS
Velocity5.0

ML experiment tracking and LLM observability platform, including Opik for evaluating LLM apps.

Comet is annexing AI cost governance from the observability side.

opikagent-observabilitycost-intelligenceevaluationmodel-selectioncontent-marketing
Current state
Comet's feed mixes real Opik engineering with a steady layer of SEO explainers, and the last two weeks have been almost entirely the latter — model-selection guides and an observability tools roundup. The product substance sits slightly further back: Agent Diagnostics, which reads across traces instead of one at a time, plus Cost Intelligence and an MCP server optimization pass. Bodies arrive as RSS teasers, so direction is readable but scope is not.
Where it's heading
Opik is widening from tracing into two adjacent jobs: telling teams which model to run where, and telling them what that choice costs. Cost Intelligence, the MCP token audit, and now a model-selection guide all point at spend governance as the commercial wedge, with evaluation-driven development as the methodology wrapped around it. The Oracle Open Agent Specification integration adds a portability argument on top — instrument once, keep the framework choice open.
Prediction
Expect model selection to stop being advice and become a product surface — routing or recommendation driven by Opik's own trace and cost data, sitting next to Cost Intelligence.

Recent moves

  1. 1d ago

    LLM Model Selection: How to Pick the Right Model for Every Agentic Task

    An explainer on matching models to agentic tasks instead of defaulting every tool call to the flagship. It rehearses the cost argument that Cost Intelligence productizes, but ships nothing itself.

    View source ↗
  2. 1d ago

    Best LLM Observability Tools of 2026: Top Platforms & Features

    A roundup of LLM observability platforms for 2026 — category-defining SEO content aimed at buyers evaluating Opik against alternatives. No product change.

    View source ↗
  3. 11d ago

    I Built a RAG Pipeline for F1 Team Radio, Then Made It Grade Itself

    A developer-relations build log: a RAG pipeline over F1 team radio, then graded with Opik's own evaluation tooling. It demonstrates evaluation-driven development on a fun dataset rather than reporting a release.

    View source ↗
  4. 26d ago

    One Prompt, 24 Versions: How Digibee Builds Prompts with Opik to Power Their AI-Native Integration Platform

    A customer story on Digibee versioning one prompt 24 times through Opik to power its integration platform. Useful proof of the prompt-experimentation workflow, but a case study, not a change.

    View source ↗
  5. 29d ago

    Beyond the Single Trace: How We Built Agent Diagnostics for Opik

    Agent Diagnostics ships in Opik, moving debugging from reading one trace at a time to surfacing misbehavior across many. It is the clearest recent product move in the feed and fits the shift from passive tracing to active diagnosis.

    View source ↗
  6. 1mo ago

    What Is an Agent Harness? The Layer That Makes AI Agents Actually Work

    A concept piece naming the retry logic and prompt assembly most teams build ad hoc as the agent harness layer. Category framing that positions Opik, with nothing shipped.

    View source ↗