Research

Sort by
May 11, 2026

Marginal tool utility in agentic coding

Marginal tool utility and tool efficiency measure whether individual tool calls improve an agent's probability of solving the task. Removing noisy tools preserved accuracy while doubling efficiency.

Marginal tool utility signs across default APEX-SWE Observability trajectories by GPT-5.3-Codex.
September 22, 2026

Root cause accuracy with Claude Code

We reran the RCA benchmark with Claude Code on Fable 5.1. It improved in every tool setup, and Foam MCP took it from 63.6% to 86.4%.

86.4%
Claude Code with Foam MCP
+23pp
from adding Foam MCP
22 production incidents, same judge as April
June 18, 2026

Clustering and noise reduction

How Drain3 streaming template mining compares to Sentry fingerprinting across 11 production environments.

10.9x
median issue reduction
95.5%
fewer issues to triage
Across 11 production environments
May 27, 2026

Structural gaps in error monitoring: evidence from production systems

Evidence, causes, and a path forward for grouping, prioritization, configuration decay, alert noise, and AI-generated fixes.

1Clustering & duplicates
2Error prioritization
3Configuration decay
4Alert noise
5AI-generated fixes
April 4, 2026

Bug fixing accuracy from observability data

A benchmark for the question every debugging agent should answer: what caused the production failure? Evaluated on root cause analysis from telemetry, not log summarization.

81.8%
Foam agent
63.6%
Cursor with Foam MCP
22 production incidents, April 4 run