AI agent digital forensics benchmark
Security Benchmarks is an independent research project measuring how AI agents investigate security incidents. The public DFIR 0.6 release compares 19 models on 48 cases in two separate tiers: 36 basic investigations and 12 multi-host intrusion investigations.
Results separate final verdict accuracy, evidence-backed decisions, evidence acquisition, factual reconstruction, scope, and cost. Read the DFIR 0.6 benchmark overview and task examples or download the public results.