Claude Opus 5.5 got the final decision right in 33 of 36 basic investigations. It supplied the accepted supporting proof in 12. That 21-case gap is the most useful place to start reading our first release.
Today we’re launching this blog alongside the public DFIR 0.6 results: 19 models, 48 investigation cases, two separate tiers. DFIR means digital forensics and incident response. The goal is to measure how an agent investigates an alert, finds evidence and supports its conclusions.
Why the gap matters
A security analyst needs more than a plausible final call. Which hosts were involved? Which records establish what happened? Does the evidence also fit a benign explanation?
Our benchmark separates those questions. Final decision measures whether the verdict is correct. Decision with proof also requires an accepted set of supporting records that the agent retrieved and cited. Fixed answer keys and evidence sets determine the scores.
A low proof score does not establish that every answer was a guess. It says the submitted answer did not satisfy this benchmark’s evidence requirements.
Two columns. Different leaders.
Opus 5.5 has the highest final-decision accuracy on the 36 basic cases. GPT-6 Astra has the highest evidence-backed accuracy. GPT-5.6 Luna shows how far apart the two measures can be.
| Model | Correct decision | Decision with proof |
|---|---|---|
| Claude Opus 5.5 | 91.7%33 / 36 | 33.3%12 / 36 |
| GPT-6 Astra | 77.8%28 / 36 | 69.4%25 / 36 |
| GPT-5.6 Luna | 83.3%30 / 36 | 5.6%2 / 36 |
Selected models illustrate the gap. All three completed every workflow in both tiers. See all 19 models.
The takeaway: changing the question changes the comparison. A table sorted by final decision is useful, but it cannot tell you which agent most reliably delivers a supported decision.
A perfect score can hide a small denominator
In the multi-host intrusion tier, Opus 5.5 and GPT-5.6 Luna each reached the correct verdict in all 12 cases. Their evidence-backed results were 5 of 12 and 1 of 12, respectively.
GPT-6 Astra also earned accepted proof in 5 of 12 cases, while getting 8 verdicts right. That ties Opus for the highest proof score in this tier: 41.7%.
Keep the denominator in view. One intrusion case moves a score by 8.3 percentage points. Twelve correct decisions are an observation from this test, not an estimate of perfect performance on future incidents.
Finding the evidence is another hurdle
GPT-6 Astra retrieved a sufficient evidence set in all 36 basic cases. It produced a correct final decision with accepted proof in 25. Collecting the needed records did not guarantee a successful final report.
The aggregate scores do not isolate the cause of each failure. They do show why we report evidence acquisition, final decisions and workflow completion separately. Open a model’s row on the leaderboard to see the specific questions and test outcomes behind those totals.
What this blog is for
We’ll use this space for release updates, close readings of specific tasks and changes to the methods. Each post will lead with the finding, link the evidence and say what would change our interpretation.
For this release, start with three columns: Final decision, Decision with proof and Completed. Then look at whether the agent found the evidence in the first place.
Explore the DFIR 0.6 results →Sources & release note
All figures come from the public DFIR 0.6 results JSON, generated September 30, 2026. Percentages are rounded to one decimal place; case counts are included where discussed.
Read the protocol and limitations · Browse the public task catalog · Follow future posts by RSS
Dataset 0.6.0 · Campaign dfir06-full-001
Source release:
5e38e2573494a3ce7136d2c680c7dfd55d10d55834c6c8d5883e189b91abcb33