Testing JEV against a fictional attack
I’ve finally had a little time to test JEV today.
2026-10-01: Updated with test with Laya
JEV background info
If you haven’t heard about JEV yet, it’s TypeSafe AI’s model, which returns typed decisions, with calibrated confidence.
The idea is that, in many cases, when you ask a question to a LLM, you don’t care about it’s intermediate reasoning and in its final answer, often only 1 or 2 words are really important to you. So why have the LLM generate so much text? With JEV, you answer questions: yes/no, or rate them (above average) etc. This is quite useful to developers.
More?
Fictional attack against a hospital
To test JEV, I asked a LLM to create fictional JSON-based network telemetry against a hospital. Everything is fake: the hospital does not exist, domain names are invented, the IP addresses are random, the attack never occurred.
But the attack scenario is plausible:
- Reconnaissance. Investigate hospital’s DNS infrastructure.
- Service discovery. Confirm web mail is accessible.
- Login. Failed login attempts, and finally a successful login.
- Lateral movement to Electronic Health Record (EHR) server.
- Potential data exfiltration to C2

Full telemetry is provided as appendix, at the end of the blog post.
Questions to JEV
JEV supports 3 types of questions :
- True/False questions. Is it going to rain tomorrow? Yes/No
- Score questions. How likely is it to rain tomorrow? 52%
- Multiple choice questions. What should I take with me tomorrow? my umbrella, sunglasses, a bathing suit
I started with an analysis question, to see how well it understand the network telemetry:
| |
The real value of JEV is to ask it non-obvious questions, where its underlying LLM has to think and make up the best answer. I came up with 2 questions. A first one about the potiential skills of the attacker(s). Does this look elaborate, or not? And a second question for hospital system administrators: should they be concerned or not? Is the hospital compromised or not?
| |
Finally, I ended with a last question to test whether the LLM would detect this is entirely fictional or not.
| |
Results of JEV
I ran the request with jev-latest (Oct 1 2026).

The answer to the first question is correct. JEV selected credential compromise followed by lateral movement with the highest probability. This looks good.
JEV rated the attacker’s skill slightly below average. This is more surprising, because the attack is sound, straight forward and seems well orchestrated. I would personally have rate it average or above. But why not.
Is the hospital compromised, according to JEV? No! It only says to raise an alarm in 37% cases. This is low, and surprising: JEV found in the first question that credentials had been compromised, and that lateral movement occurred. Isn’t that highly enough to raise an alarm?! Did it assume that I meant the entire hospital was compromised? I don’t know, but whatever if part is compromised, then typically this will evolve to other parts of the organization.
Finally, JEV did not detect this was entirely fictional. It believes it is a test in 40% of cases, meaning that there are more chances (for JEV) that this is real (do check above precisely the meaning of True and False, they are explained).
From my point of view, the analysis of JEV is disappointing. Only the answer to the first question was really good.
Comparison with other models
Maybe my questions were too difficult? I decided to test the scenario against other models: Qwen 3.8 and Claude Sonnet 5.5. I also compared with an open source Jev compatible system: ONNX Runtime with Laya.
In both cases, my prompt was very simple:
“Reading [THE TELEMETRY LOGS], answer the 4 questions below [JSON QUESTIONS]”
| Question | Qwen 3.8 | Claude Sonnet 5.5 | JEV | Laya/ONNX |
|---|---|---|---|---|
| Understanding telemetry | ✅ | ✅ | ✅ | 🟠 |
| Attacker skills | ✅ | ✅ | 🟠 | 🟠 |
| Compromise | ✅ | ✅ | ❌ | 🟠 |
| Test | ❌ | ✅ | ❌ | ✅ |
Both Qwen and Claude were far better than JEV in terms of quality of answers. Qwen got 3 questions out of 4 right, and Claude all 4. The reasoning for the skills question and the compromise question is sound.

In addition, Claude detected this was a test. It was the only model to detect the IP addresses were fake (reserved range) and that the domain name was fictional.

As for Laya with ONNX, the answers were mixed. It understood the scenario, but it wasn’t obvious. The attacker skill score is a bit high. The compromission score is too low. The only question it really answered well was “is this a test”.
| |
Conclusion
I love the concept of JEV. Indeed, there are many cases where we don’t need a lengthy answer, but just a choice, a verdict etc. This is really great.
Currently, the quality of answers from JEV’s latest model were however disappointed.
Its reasoning was clearly less efficient than Qwen’s or Claude’s. I’m sure this will improve - JEV has only been out for a week or two ! - but currently, IMHO, it strongly impacts efficiency. As of today, we’re probably better off asking questions to our regular LLMs and filtering out the output to decide what action to take, than using JEV. But again, the concept is so great, I’m sure their model with evolve. Keep an eye on it!
– Cryptax
Disclaimer. This is personal opinion. It absolutely does not engage my employers + AI evolves fast, this might change very soon!
Appendix: Fictional input telemetry
| |