The Security Benchmarks blog

Read past the score.

Research updates on AI security agents, the evidence behind their answers, and what the scores leave out.

Latest dispatchFollow via RSS ↗

No. 02 / Decision-model analysis

18/18 verdicts. 12/18 next steps.

Jev’s perfect verdict score, the next-step gap, and what six SOC decision models tell us about choosing an action.

Read the decision-model post →
Jev 1.13 · SOC decision suite 1.0.0
18/18

Correct
verdict answers

12/18

Accepted
next-step answers

Six synthetic cases. One pass.
Classification and action are separate tests.

No. 01 / Benchmark analysis

The verdict is right. Where’s the evidence?

Our first 19-model DFIR release shows why a correct answer and a supported answer deserve separate columns.

Read the launch post →
Claude Opus 5.5 · Basic investigations
33/36

Correct final
decisions

12/36

Decisions with
accepted proof

Same model. Same cases.
A different question about performance.

Our approach

Useful findings. Visible evidence. Room to be wrong.

Security Benchmarks is an independent research project studying AI agents in digital forensics and incident response. This blog explains new results, investigates specific tasks and records changes to the methods.

We keep posts brief, show the denominators and distinguish a measured result from an interpretation. Read the benchmark overview or go straight to the results tables.