Same model. Same cases. A different question about performance.
Our approach
Useful findings. Visible evidence. Room to be wrong.
Security Benchmarks is an independent research project studying AI
agents in digital forensics and incident response. This blog
explains new results, investigates specific tasks and records
changes to the methods.
We keep posts brief, show the denominators and distinguish a
measured result from an interpretation. Read the
benchmark overview
or go straight to the results tables.