Once you've confirmed a bad process, you still have to scope the incident. What did it write? What did it launch? Which accounts and systems did it reach? A familiar filename or a shared account tells you very little on its own.

We put Clef, Clef Flash and Jev 1.13 through that task. Each started with one confirmed incident member and received the evidence needed to assess connected processes, files, hosts, accounts and infrastructure. Clef recovered more of the incident in this first eight-case pilot. All three left confirmed incident members outside their final scope.

Follow the evidence, one connection at a time

The harness asks whether an item belongs to the same incident. Accept a connection and the harness follows it, exposing more candidates. Withhold it and that branch can stop there.

That is where the operational tradeoff shows up. Pull in everything sharing a filename or server and the case fills with noise. Stop too early and the next process, file or host never gets assessed.

The cases exercise process lineage, exact file identity, timing, persistence and session pivots. A process name is a lead. The process instance, hash and timestamp are what let you join the activity up. The models get those records; they do not collect logs or run tools against a network.

One bad process. What else is connected?

Follow the evidence from a confirmed malicious process.

From one bad process to the rest of the attack A confirmed malicious process writes a file. That file starts a new process, which makes a network connection. Each link needs supporting evidence. A backup job on the same computer is separate, with no link to the attack. The starting process is supplied and earns no credit. writes runs connects Bad process Dropped file New process Network connection Given to the model Same computer. No attack link. The routine backup stays outside the incident. Backup job From one bad process to the rest of the attack A confirmed malicious process writes a file. That file starts a new process, which makes a network connection. Each link needs supporting evidence. A backup job on the same computer is separate, with no link to the attack. The starting process is supplied and earns no credit. writes runs connects Bad process Dropped file New process Network link Given Same computer. No attack link. The backup stays out.

Every arrow needs evidence. Accept a link to investigate the next item; miss it and that branch may stop there.

What they recovered

We scored the final set of incident members with F1, shown here on a 0–100 scale. Missing a member costs credit. Adding something unsupported also costs credit. The confirmed starting point earns nothing because it was supplied to the model.

Suite 1.0.0 · 8 cases · 3 repeats · family-macro means
Model Scope score (F1) Expected incident pieces found
Clef 70.8 / 100 64.6%
Clef Flash 62.1 / 100 52.1%
Jev 1.13 20.8 / 100 18.1%

There were eight cases, one from each of eight scenario families, with three runs per model. Each case has equal weight in these averages. The repeat runs use the same evidence; they are not fresh incidents.

No model added an item the key marked unrelated or unresolved in this run. The separation came from missed members. That is a narrow observation from this case set, not evidence that any model is safe from false connections.

Find the attack. Keep unrelated activity out.

An illustrative example: the overlap is what the model got right.

The real attack and the model’s case overlap, but do not match Illustrative example, not a measured model result. The actual attack contains a dropped file, new process and network connection. The model includes the file and process, shown in the overlap, but misses the connection in the actual-attack-only region and wrongly adds a backup job in the model-only region. Two correct finds, one miss and one wrong addition give an F1 score of 66.7 out of 100. The supplied starting process is excluded. Actual attack Model’s case Network connection 1 missed Dropped file New process 2 found Backup job 1 wrong addition The real attack and the model’s case overlap, but do not match Illustrative example, not a measured model result. The actual attack contains a dropped file, new process and network connection. The model includes the file and process, shown in the overlap, but misses the connection in the actual-attack-only region and wrongly adds a backup job in the model-only region. Two correct finds, one miss and one wrong addition give an F1 score of 66.7 out of 100. The supplied starting process is excluded. Actual attack Model’s case Network connection 1 missed Dropped file New process 2 found Backup job 1 wrong addition
Example scope score (F1)66.7 / 100

2 correct finds
1 missed item
1 wrong addition

Missed attack activity and unrelated additions both lower the score. This example is not a result from any of the three models.

How the score is calculated

F1 balances how much of the attack the model found with how much of its case actually belongs. Here, it found 2 of 3 attack items, and 2 of its 3 additions were correct.

F1 = (2 × correct finds) ÷ (2 × correct finds + missed items + wrong additions).
Here: 4 ÷ (4 + 1 + 1) = 0.667, or 66.7 out of 100.

The supplied starting process earns no credit. Keeping unrelated activity out avoids a penalty, but earns no points on its own.

Where the trail stopped

Jev's answers landed in the review bucket more often under our common rules. Its expansion reached 51.3% of the available candidates on average. Clef reached 77.7%; Flash reached 71.5%.

Those unseen candidates matter. If the model withholds the first link in a chain, a correct call on the next link cannot help because the harness never asks it. The final score therefore measures the model together with the policy used to follow its answers. We froze that policy before the run.

We also ran a separate check using the same 29 individual membership questions for every model. Clef got 56.5% right, Flash 44.0% and Jev 32.9%. Always answering “related” would score 67.9% on this mix. The question set leans toward related items, and the raw accuracy needs to be read with that baseline beside it.

What this result supports

Clef is the stronger scoping result in this pilot. It still missed enough of the expected scope that an analyst would have more work to do.

We made 464 paid calls for about 1.4 cents, then reconstructed the requests, expansion paths and scores from the saved responses. Eight development cases are too few to set a production routing policy. We have not tested a separate held-out set, and three passes over the same evidence do not fix that.

The next run needs held-out cases and a closer look at the missed links: weak evidence, a poor membership call, or a threshold that stopped the branch too early. Those are different failure modes and need different fixes.

Source and release note

All figures come from the incident-scope 1.0.0 measured release. Percentages are rounded to one decimal place. Native endpoint information: Clef, Clef Flash, Jev 1.13.

Measurement: a2532d6e20414efb9e37e01536e8ff69187cbb34cf3a4bcb219093370ae5db13
Public release: 48c8550a372e37e21f65023cfbd56cfe2899124c818e861add49a3f779fcb3d9

Technical settings and verification

Incident-scope suite 1.0.0 uses full native decision packets. Membership requires a score of at least 0.8; rejection requires at most 0.2. Intermediate answers withhold membership. Review also flags scores within 0.05 of the link threshold. Settings were frozen before execution. Fixed decisions use the same reference frontier; expansion follows each model’s accepted links. Scores average within families, equally across families, then across three repeats.

Clef and Clef Flash used PrimeIntellect. Jev used TypeSafe and resolved to typesafe/jev-1.13-20260917. Provider-reported charges totaled $0.013817727. Independent replay reconstructed every request and frontier, checked exact model/provider identity, matched the persistent budget ledger and recomputed the scores. Probabilities have not been validated as calibrated confidence.