Snorkel AI
AI data development platform for enterprise model fine-tuning, evaluation, and curation.
Snorkel has stopped labeling data and started defining what agent competence means.
◆Recent moves
- 13d ago
Milestone-Based Evaluation and Training for Long-Horizon AI Agents
A research post arguing that long-horizon agents should be scored on intermediate milestones rather than a final answer, since earlier decisions constrain later ones. It extends the same measurement thesis running through the feed, but describes methodology rather than a released artifact.
View source ↗ - 15d ago
Enterprise environments and training AI agents for real-world workflows
Makes the case that enterprise workflows require simulated environments with tools, records, and approval rules rather than single-turn benchmarks. Positioning for Snorkel's environments work, without a shipped artifact attached.
View source ↗ - 22d ago
Claude Opus 5: Performance and Error Analysis on Frontier Coding Tasks
An error analysis of Claude Opus 5 on Senior SWE-bench, noting its standing on the leaderboard and its result in the bug and performance investigation category. Third-party model evaluation that demonstrates the benchmark rather than changing Snorkel's product.
View source ↗ - 1mo ago
Senior SWE-Bench: Evaluating Coding Agents Like Senior Engineers
Introduces Senior SWE-Bench, an open-source benchmark of 100 tasks drawn from real pull requests across 12 production repositories, half held private to limit contamination. A concrete released artifact rather than commentary, and the reference point the model analyses in this window are scored against.
View source ↗ - 1mo ago
Grok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work
Results from running Grok 4.5 against Snorkel's GDPval+ professional-reasoning dataset alongside two other frontier models. Comparative evaluation content that showcases the dataset.
View source ↗ - 1mo ago
Agents’ Last Exam: AI Benchmarking for Real Work
A reading-group writeup of Agents' Last Exam, a long-horizon benchmark built with Berkeley RDI and hundreds of expert contributors. Event coverage of collaborative research rather than a Snorkel release.
View source ↗