Repository-level bug fixing: read the failing test, navigate an unfamiliar codebase, and produce a patch that resolves the issue.
What this model claims, and what it has not measured yet.
Measured on real benchmark episodes against the Claude Fable 5 frontier anchor. Accuracy is the task-success gap in points; savings are per run, latency as p50 model time.
How far this model has come.
Everything that changes something, or spends anything, is behind the login.