<?xml version='1.0' encoding='utf-8'?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>Security Benchmarks blog</title>
    <link>https://secbenchmarks.com/blog/</link>
    <description>Research updates and analysis of AI agents in digital forensics and incident response.</description>
    <language>en-us</language>
    <lastBuildDate>Wed, 30 Sep 2026 22:32:40 -0400</lastBuildDate>
    <atom:link href="https://secbenchmarks.com/blog/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>18/18 verdicts. 12/18 next steps.</title>
      <link>https://secbenchmarks.com/blog/2026-09-30-soc-decision-models.html</link>
      <guid isPermaLink="true">https://secbenchmarks.com/blog/2026-09-30-soc-decision-models.html</guid>
      <pubDate>Wed, 30 Sep 2026 22:32:40 -0400</pubDate>
      <description>A close look at six SOC decision models: Jev’s perfect verdict score, the next-step gap, answer-format failures and the limits of a six-case benchmark.</description>
      <category>Decision-model analysis</category>
    </item>
    <item>
      <title>The verdict is right. Where’s the evidence?</title>
      <link>https://secbenchmarks.com/blog/2026-09-30-ai-security-agents-evidence.html</link>
      <guid isPermaLink="true">https://secbenchmarks.com/blog/2026-09-30-ai-security-agents-evidence.html</guid>
      <pubDate>Wed, 30 Sep 2026 21:08:26 -0400</pubDate>
      <description>Our first 19-model DFIR release shows why a correct answer and a supported answer deserve separate columns. Opus 5.5 got 33 of 36 basic verdicts right, with accepted proof in 12. Read the case counts, comparisons and caveats.</description>
      <category>Benchmark analysis</category>
    </item>
  </channel>
</rss>