Jev 1.13 answered all 18 verdict questions correctly. It selected an accepted next step in 12 of 18. GPT-6 Astra reached 18 of 18 in both categories. GPT-5.6 Sol got 15 verdicts right, but all 18 next steps.
These are the published SOC decision suite 1.0.0 results: six models, 272 questions per model, six synthetic cases and two related construction families. SOC means security operations center. This is a separate test from our 48-case DFIR agent evaluation.
The evidence is already on the desk
Each model receives the same prepared telemetry, question and answer choices. It must choose an answer. It does not search for logs or call investigation tools. Jev uses a native decision interface; the chat models return a choice through a JSON wrapper.
The suite asks six kinds of question: classify the alert; update a hypothesis; check a conclusion; judge evidence relevance; choose a next step; and follow an escalation policy. Those are separate scores. Fixed answer keys determine credit, with multiple next actions accepted where appropriate.
Two decisions, different results
| Model | Correct verdict | Accepted next step | Invalid verdict answers |
|---|---|---|---|
| GPT-6 Astra | 100.0%18 / 18 | 100.0%18 / 18 | 0 / 18 |
| Jev 1.13 | 100.0%18 / 18 | 66.7%12 / 18 | 0 / 18 |
| GPT-5.6 Sol | 83.3%15 / 18 | 100.0%18 / 18 | 0 / 18 |
| Claude Sonnet 5 | 77.8%14 / 18 | 61.1%11 / 18 | 2 / 18 |
| GLM-5.3 Flash (FP8) | 50.0%9 / 18 | 55.6%10 / 18 | 9 / 18 |
| Claude Opus 5.5 | 27.8%5 / 18 | 22.2%4 / 18 | 13 / 18 |
All six published models. Each shown category has three questions per case. Scores weight cases equally; these 18 questions are not 18 independent incidents. Primary results · GLM results.
The next-step gap matters. Jev’s 66.7% is below the category’s 72.2% best constant-answer baseline. That baseline always chooses the single option that scores best across the category. It exposes the ease of exploiting an imbalanced test; it is not a useful operating policy.
A verdict score alone would miss this distinction. These results suggest testing the action you want a model to take, rather than treating accurate classification as a proxy for every downstream decision.
The output contract is part of the score
Claude Opus 5.5 had 13 invalid verdict answers out of 18; GLM-5.3 Flash had nine. Invalid or missing choices receive zero. Prose or fenced JSON can fail this protocol even when it contains the correct choice.
Those failures matter to a system expecting a usable decision. They also limit the interpretation: low accuracy here does not isolate reasoning quality from answer-format compliance. Opus’s investigation results come from a different task and response contract.
A cheap decision still needs to be the right one
Jev’s recorded cost for its 272-question pass was about $0.02; Astra’s was about $2.16. These are observed run costs, not current price quotes or production cost forecasts. The models used different interfaces and settings.
The useful question is which task a cheaper model can handle reliably. This small evaluation does not establish a safe routing threshold. The escalation category tests compliance with a stated policy, not whether a larger model will solve the next case better.
Read the category you would actually use
On the leaderboard, select Archived model results, then SOC decisions. Compare the category score with its baseline, invalid answers and critical errors. A perfect verdict column is one useful observation. The next action is another test.
Explore the decision results →Sources & release note
All measured figures come from the public primary decision release and GLM release, bound to the same suite and protocol. Percentages are rounded to one decimal place. Published benchmark data and scoring are unchanged.
Read our investigation-model analysis · Follow future posts by RSS
Suite 1.0.0 · source cases 0.4.1
Primary release:
0ab3d5225e33324d72c05dd4a1f4fad826a2cf2ed8564d94a80dfb7607dfbeae
GLM
release:
5a2f1e011a084fb47ed545e2f1e1b26f6ee469124b8393de8933d758a9464c1b