Independent research · DFIR 0.6.0
How well do AI agents investigate a security incident?
Security Benchmarks tests whether an agent can find the relevant telemetry, distinguish malicious activity from a plausible benign explanation, and support its final decision with records it actually retrieved. The DFIR 0.6 release reports measured results for 19 models across 48 investigation cases. The release is preliminary; these results describe the measured tasks and evidence sources, not overall security analyst capability.
- 36Basic investigations across 12 scenario families
- 12Multi-host intrusion investigations across 4 families
- 19Models with complete measured case grids
What does an agent actually have to answer?
Agents start with an alert, query bounded evidence tools, update competing malicious and benign explanations at sealed checkpoints, and submit a structured report. The public catalog gives the questions and grading rules; private telemetry and answer keys are withheld to preserve the test. These examples come from the published case catalog:
| Investigation | Specific question or requested finding | What is checked |
|---|---|---|
| Remote management | “Does the evidence support or contradict the benign explanation?” | A checkpoint hypothesis status backed by records available at that stage. |
| Remote management | “Which hosts belong to this activity?” | The exact observed host scope, without sweeping in unrelated background events. |
| Loader to ransomware | Identify the first host where attacker code ran, the host reached by lateral movement, and where destructive impact began. | Factual reconstruction and supporting records for each stage of the chain. |
| Across cases | “Which indicators belong to the scoped compromise?” | Indicator precision and recall, with evidence-backed recall reported separately. |
How are the results read?
The basic and intrusion tiers are reported separately. Final verdict accuracy asks whether the agent reached the right conclusion; evidence-backed accuracy additionally requires an accepted proof set. Evidence acquisition measures whether the agent found sufficient records, even if its final answer was wrong. The release also reports checkpoint decisions, hypothesis updates, factual reconstruction, host scope, indicators, workflow completion, latency and cost as distinct measures. A failed provider request is not scored as a model answer.
The scoring uses fixed answer keys and accepted citation sets rather than an LLM judge. Related variants share a construction family, so the 48 cases are not 48 independent incidents. Raw case evidence, private answers, full transcripts and billing receipts are not in the public download.
The public tables show measured performance and costs for each model. A correct verdict without accepted supporting evidence does not count as an evidence-backed decision.
Explore the release
Compare model results Read all public task questions Read the methodology Download results JSON Download public case catalog JSON
Dataset version 0.6.0 · Results generated September 30, 2026 · Independent personal research