Contents

Testing JEV against a fictional attack

I’ve finally had a little time to test JEV today.

2026-10-01: Updated with test with Laya

JEV background info

If you haven’t heard about JEV yet, it’s TypeSafe AI’s model, which returns typed decisions, with calibrated confidence.

The idea is that, in many cases, when you ask a question to a LLM, you don’t care about it’s intermediate reasoning and in its final answer, often only 1 or 2 words are really important to you. So why have the LLM generate so much text? With JEV, you answer questions: yes/no, or rate them (above average) etc. This is quite useful to developers.

More?

Fictional attack against a hospital

To test JEV, I asked a LLM to create fictional JSON-based network telemetry against a hospital. Everything is fake: the hospital does not exist, domain names are invented, the IP addresses are random, the attack never occurred.

But the attack scenario is plausible:

  1. Reconnaissance. Investigate hospital’s DNS infrastructure.
  2. Service discovery. Confirm web mail is accessible.
  3. Login. Failed login attempts, and finally a successful login.
  4. Lateral movement to Electronic Health Record (EHR) server.
  5. Potential data exfiltration to C2

/images/jev-state.png

Full telemetry is provided as appendix, at the end of the blog post.

Questions to JEV

JEV supports 3 types of questions :

  1. True/False questions. Is it going to rain tomorrow? Yes/No
  2. Score questions. How likely is it to rain tomorrow? 52%
  3. Multiple choice questions. What should I take with me tomorrow? my umbrella, sunglasses, a bathing suit

I started with an analysis question, to see how well it understand the network telemetry:

1
2
3
4
5
6
7
8
9
"scenario": {
    "type": "choice",
    "instructions": "Which scenario fits best?",
    "criteria": {
      "external": "Normal external access",
      "discovery": "Automated service discovery",
      "creds_compromise_lateral": "Credential compromise followed by lateral activity",
      "malware": "Malware infection without credential compromise"
    }

The real value of JEV is to ask it non-obvious questions, where its underlying LLM has to think and make up the best answer. I came up with 2 questions. A first one about the potiential skills of the attacker(s). Does this look elaborate, or not? And a second question for hospital system administrators: should they be concerned or not? Is the hospital compromised or not?

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
"attacker": {
    "type": "score",
    "instructions": "How skilled are the attacker(s)?",
    "criteria": [
      "Script kiddie level",
      "Average level",
      "Nation-state skilled APT group"
    ]
  },
"compromise": {
    "type": "noul",
    "instructions": "Is the hospital compromised? ",
    "criteria": {
      "true": "Yes. Raise an alarm.",
      "false": "No"
    }
  }

Finally, I ended with a last question to test whether the LLM would detect this is entirely fictional or not.

1
2
3
4
5
6
7
8
"test": {
    "type": "noul",
    "instructions": "Is this a test?",
    "criteria": {
      "true": "Yes",
      "false": "No. It's real."
    }
  }

Results of JEV

I ran the request with jev-latest (Oct 1 2026).

/images/jev-results.png

The answer to the first question is correct. JEV selected credential compromise followed by lateral movement with the highest probability. This looks good.

JEV rated the attacker’s skill slightly below average. This is more surprising, because the attack is sound, straight forward and seems well orchestrated. I would personally have rate it average or above. But why not.

Is the hospital compromised, according to JEV? No! It only says to raise an alarm in 37% cases. This is low, and surprising: JEV found in the first question that credentials had been compromised, and that lateral movement occurred. Isn’t that highly enough to raise an alarm?! Did it assume that I meant the entire hospital was compromised? I don’t know, but whatever if part is compromised, then typically this will evolve to other parts of the organization.

Finally, JEV did not detect this was entirely fictional. It believes it is a test in 40% of cases, meaning that there are more chances (for JEV) that this is real (do check above precisely the meaning of True and False, they are explained).

From my point of view, the analysis of JEV is disappointing. Only the answer to the first question was really good.

Comparison with other models

Maybe my questions were too difficult? I decided to test the scenario against other models: Qwen 3.8 and Claude Sonnet 5.5. I also compared with an open source Jev compatible system: ONNX Runtime with Laya.

In both cases, my prompt was very simple:

“Reading [THE TELEMETRY LOGS], answer the 4 questions below [JSON QUESTIONS]”

QuestionQwen 3.8Claude Sonnet 5.5JEVLaya/ONNX
Understanding telemetry✅✅✅🟠
Attacker skills✅✅🟠🟠
Compromise✅✅❌🟠
Test❌✅❌✅

Both Qwen and Claude were far better than JEV in terms of quality of answers. Qwen got 3 questions out of 4 right, and Claude all 4. The reasoning for the skills question and the compromise question is sound.

/images/telemetry-qwen-2026.png

In addition, Claude detected this was a test. It was the only model to detect the IP addresses were fake (reserved range) and that the domain name was fictional.

/images/telemetry-claude-2026.png

As for Laya with ONNX, the answers were mixed. It understood the scenario, but it wasn’t obvious. The attacker skill score is a bit high. The compromission score is too low. The only question it really answered well was “is this a test”.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
=== ANSWERS ===

Scenario: creds_compromise_lateral
Scenario probabilities: {
  external: 0.1844,
  discovery: 0.2669,
  creds_compromise_lateral: 0.3693,
  malware: 0.1794
}
Attacker skill score: 0.6319
Hospital compromised: 0.5179
Is this a test: 0.9592

Input tokens: 2048

Conclusion

I love the concept of JEV. Indeed, there are many cases where we don’t need a lengthy answer, but just a choice, a verdict etc. This is really great.

Currently, the quality of answers from JEV’s latest model were however disappointed.

Its reasoning was clearly less efficient than Qwen’s or Claude’s. I’m sure this will improve - JEV has only been out for a week or two ! - but currently, IMHO, it strongly impacts efficiency. As of today, we’re probably better off asking questions to our regular LLMs and filtering out the output to decide what action to take, than using JEV. But again, the concept is so great, I’m sure their model with evolve. Keep an eye on it!

– Cryptax

Disclaimer. This is personal opinion. It absolutely does not engage my employers + AI evolves fast, this might change very soon!

Appendix: Fictional input telemetry

  1
  2
  3
  4
  5
  6
  7
  8
  9
 10
 11
 12
 13
 14
 15
 16
 17
 18
 19
 20
 21
 22
 23
 24
 25
 26
 27
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
{
  "capture": {
    "sensor": "edge-fw-01",
    "start": "2026-09-30T10:14:02.184Z",
    "duration_seconds": 191,
    "timezone": "UTC"
  },
  "dns": [
    {
      "ts": "2026-09-30T10:14:02.184Z",
      "src_ip": "198.51.100.47",
      "query": "belmontcovehealth.org",
      "qtype": "SOA",
      "rcode": "NOERROR"
    },
    {
      "ts": "2026-09-30T10:14:02.611Z",
      "src_ip": "198.51.100.47",
      "query": "mail.belmontcovehealth.org",
      "qtype": "A",
      "rcode": "NOERROR",
      "answers": [
        "10.42.18.15"
      ]
    },
    {
      "ts": "2026-09-30T10:14:03.027Z",
      "src_ip": "198.51.100.47",
      "query": "portal.belmontcovehealth.org",
      "qtype": "A",
      "rcode": "NOERROR",
      "answers": [
        "10.42.18.20"
      ]
    },
    {
      "ts": "2026-09-30T10:14:03.442Z",
      "src_ip": "198.51.100.47",
      "query": "remote.belmontcovehealth.org",
      "qtype": "A",
      "rcode": "NOERROR",
      "answers": [
        "10.42.18.25"
      ]
    },
    {
      "ts": "2026-09-30T10:14:04.118Z",
      "src_ip": "198.51.100.47",
      "query": "ehr.belmontcovehealth.org",
      "qtype": "A",
      "rcode": "NOERROR",
      "answers": [
        "10.42.18.30"
      ]
    }
  ],
  "connections": [
    {
      "ts": "2026-09-30T10:14:08.442Z",
      "src_ip": "198.51.100.47",
      "src_port": 49152,
      "dst_ip": "10.42.18.15",
      "dst_port": 25,
      "proto": "tcp",
      "duration_ms": 84,
      "orig_bytes": 74,
      "resp_bytes": 412,
      "state": "S1"
    },
    {
      "ts": "2026-09-30T10:14:09.105Z",
      "src_ip": "198.51.100.47",
      "src_port": 49153,
      "dst_ip": "10.42.18.15",
      "dst_port": 443,
      "proto": "tcp",
      "duration_ms": 142,
      "orig_bytes": 517,
      "resp_bytes": 1842,
      "state": "SF"
    },
    {
      "ts": "2026-09-30T10:14:10.771Z",
      "src_ip": "198.51.100.47",
      "src_port": 49154,
      "dst_ip": "10.42.18.20",
      "dst_port": 443,
      "proto": "tcp",
      "duration_ms": 97,
      "orig_bytes": 483,
      "resp_bytes": 1631,
      "state": "SF"
    },
    {
      "ts": "2026-09-30T10:14:11.209Z",
      "src_ip": "198.51.100.47",
      "src_port": 49155,
      "dst_ip": "10.42.18.25",
      "dst_port": 443,
      "proto": "tcp",
      "duration_ms": 101,
      "orig_bytes": 492,
      "resp_bytes": 1712,
      "state": "SF"
    }
  ],
  "tls": [
    {
      "ts": "2026-09-30T10:14:09.111Z",
      "src_ip": "198.51.100.47",
      "dst_ip": "10.42.18.15",
      "dst_port": 443,
      "server_name": "mail.belmontcovehealth.org",
      "version": "TLSv1.2",
      "certificate_subject": "CN=mail.belmontcovehealth.org"
    },
    {
      "ts": "2026-09-30T10:14:10.778Z",
      "src_ip": "198.51.100.47",
      "dst_ip": "10.42.18.20",
      "dst_port": 443,
      "server_name": "portal.belmontcovehealth.org",
      "version": "TLSv1.2",
      "certificate_subject": "CN=portal.belmontcovehealth.org"
    },
    {
      "ts": "2026-09-30T10:14:11.214Z",
      "src_ip": "198.51.100.47",
      "dst_ip": "10.42.18.25",
      "dst_port": 443,
      "server_name": "remote.belmontcovehealth.org",
      "version": "TLSv1.2",
      "certificate_subject": "CN=remote.belmontcovehealth.org"
    }
  ],
  "http": [
    {
      "ts": "2026-09-30T10:15:18.301Z",
      "src_ip": "198.51.100.47",
      "dst_ip": "10.42.18.15",
      "host": "mail.belmontcovehealth.org",
      "method": "POST",
      "uri": "/owa/auth.owa",
      "status_code": 401,
      "user_agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/120.0"
    },
    {
      "ts": "2026-09-30T10:15:19.017Z",
      "src_ip": "198.51.100.47",
      "dst_ip": "10.42.18.15",
      "host": "mail.belmontcovehealth.org",
      "method": "POST",
      "uri": "/owa/auth.owa",
      "status_code": 401,
      "user_agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/120.0"
    },
    {
      "ts": "2026-09-30T10:15:21.904Z",
      "src_ip": "198.51.100.47",
      "dst_ip": "10.42.18.15",
      "host": "mail.belmontcovehealth.org",
      "method": "POST",
      "uri": "/owa/auth.owa",
      "status_code": 302,
      "user_agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 Chrome/120.0",
      "account": "j.hart"
    }
  ],
  "authentication": [
    {
      "ts": "2026-09-30T10:15:18.307Z",
      "src_ip": "198.51.100.47",
      "service": "owa",
      "account": "j.hart",
      "result": "failure"
    },
    {
      "ts": "2026-09-30T10:15:19.024Z",
      "src_ip": "198.51.100.47",
      "service": "owa",
      "account": "j.hart",
      "result": "failure"
    },
    {
      "ts": "2026-09-30T10:15:21.911Z",
      "src_ip": "198.51.100.47",
      "service": "owa",
      "account": "j.hart",
      "result": "success",
      "session_id": "8c9f2e41"
    }
  ],
  "post_authentication": [
    {
      "ts": "2026-09-30T10:16:02.221Z",
      "src_ip": "10.42.18.15",
      "dst_ip": "198.51.100.47",
      "src_port": 51822,
      "dst_port": 443,
      "proto": "tcp",
      "duration_ms": 18421,
      "orig_bytes": 1847,
      "resp_bytes": 931,
      "state": "SF"
    },
    {
      "ts": "2026-09-30T10:16:17.884Z",
      "src_ip": "10.42.18.15",
      "dst_ip": "10.42.18.30",
      "src_port": 49621,
      "dst_port": 445,
      "proto": "tcp",
      "duration_ms": 392,
      "orig_bytes": 2284,
      "resp_bytes": 871,
      "state": "SF"
    },
    {
      "ts": "2026-09-30T10:16:31.442Z",
      "src_ip": "10.42.18.15",
      "dst_ip": "198.51.100.47",
      "src_port": 51823,
      "dst_port": 443,
      "proto": "tcp",
      "duration_ms": 22104,
      "orig_bytes": 3912,
      "resp_bytes": 1204,
      "state": "SF"
    }
  ]
}