Jev picks. It doesn't write.
A visual walk through TypeSafe's Jev: what it is, what's actually known about how it's built, what "zero hallucination" really means, and where I plan to use it: my blog scout and my self-improving harness.

This is the deck for a talk I'm giving my team about Jev. I kept the words short on purpose. Scroll it like slides.
Nothing here is writing text
Jevpilot, a 3D driving sim where Jev makes the calls. Clip: @jpschroeder
Every frame, the model picks one of a handful of moves. No sentences, nothing to parse. That's Jev.
What it is
typesafe.ai, Sep 23, 2026
Jev is TypeSafe AI's first "System One" model. You give it some state and a few typed questions. It gives back a pick for each question, with probabilities. It never writes a sentence.
Where the name comes from

Kahneman's split from Thinking, Fast and Slow. System 1 is the fast gut call: is that a face, should I brake. System 2 is the slow, effortful kind: long division, writing an essay.
We built LLMs for System 2 work, then use them for System 1 calls all day long. Jev only does the gut call.
Writing vs picking

An LLM builds its answer one token at a time, and every token waits for the one before it. Jev scores every option at once, in one pass.
Same 27 questions, Jev vs an LLM, from TypeSafe's launch. Clip: @typesafeai
TypeSafe's own side-by-side: the same 27 questions to Jev and to an LLM. Jev is done in 0.114 s. The LLM takes 8.6 s.
What goes in, what comes out

State goes in. Questions go in, each with options you define. Typed answers come out, all at the same time.
Three question types: Choice (up to 255 options), Score (2 to 10 ordered levels) and Noul (one probability). Every answer comes with a confidence.
Under the hood

Here's what is actually known about the network, sorted by how sure we can be.
- TypeSafe says: it's transformer-based but not a language model. They call it "a new model architecture" with a parallel sampler. All answers come out in one pass, each question is answered on its own against the shared state, and it's trained only on synthetic data.
- Outsiders measured it by poking at the API (not confirmed): a decoder LLM backbone with the text output swapped for a probability readout, a tokenizer closest to Qwen's, probably mixture-of-experts at roughly 10B active parameters. The state gets processed once and shared. Options inside one question can see each other; separate questions can't.
- Nobody outside knows: the size, the base model, the training data, or even encoder vs decoder. There's no paper.
Outside measurements: archerhume, "Jev's Architecture Unmasked" and the 6.6M-parameter rebuild in jev-from-scratch.
How it's trained
From TypeSafe's launch post
LLM chat models get tuned with RLHF: the model is rewarded when a human rater likes the answer. Jev gets tuned with what TypeSafe calls RLCD, reinforcement learning for calibrated decisions. It's rewarded when its probabilities match what actually happened.

The funny part: Diogo Almeida, TypeSafe's CEO, spent about four years at OpenAI working on RLHF, InstructGPT and ChatGPT.
What "calibrated" buys you

When a calibrated model says 70%, it's right about 70% of the time. That's what lets you put a threshold on it and trust the threshold.
One outside test measured a calibration error of 0.031 on 1,200 MMLU questions. Good, not perfect.
"Zero hallucination"
From TypeSafe's launch post. Read the last bullet.
TypeSafe puts Jev at 0% on these charts, and then tells you right underneath: "Our number is not empirical. Schema matching is guaranteed."

So here's what "can't hallucinate" really means. It can't make up an answer that isn't on your list. It can still pick the wrong thing from the list. That second part is the one you have to measure yourself.
How you call it
docs.typesafe.ai
from typesafe_sdk import Choice, Noul, TypeSafeClient
client = TypeSafeClient() # reads TYPESAFE_API_KEY
r = client.system_one(
state="car at 42 km/h, obstacle 12 m ahead, lane 2, raining",
questions={
"action": Choice(instructions="What should the car do now?",
criteria={"brake": "slow down now",
"maintain": "keep speed",
"accelerate": "speed up"}),
"is_safe": Noul(instructions="The current speed is safe"),
},
)
r.answers["action"].choice # "brake", plus probabilities and a confidence
r.answers["is_safe"].noul # e.g. 0.12
TypeSafe quotes 70 to 500 ms end to end, $0.042 per million input tokens, and free output. Running it over 176 of my blog candidates cost me about a cent.
What TypeSafe's evals actually show
evals.typesafe.ai
TypeSafe took four automation jobs: triaging security incidents, watching agent traces, processing invoices and answering customer service tickets. They ran each job two ways. Circles are one big prompt that does the whole job. Diamonds are a workflow: the job is broken into many small typed questions, and plain code makes the final call from the answers.
Every dot is one model, averaged over the four jobs. Higher means more accurate. Further left means cheaper per case, and each gridline is 10x, so small gaps are big differences. The line through the top-left dots is the frontier: nothing beats those on both cost and accuracy at once.
Jev sits alone on the far left: 67.8% accurate at $0.0004 a case and 0.4 s. The most accurate LLM workflow (sol) gets 74.1% at $0.08 and 23 s. So Jev gives up about six points of accuracy to be roughly 200x cheaper and 60x faster. Against an LLM at the same accuracy (terra, 67.9%), it's about 75x cheaper and 25x faster.
Two things before anyone quotes this. "Accurate" means agreeing with a reference answer built from two frontier LLMs (GPT-6 Astra and Fable 5.1), not with human labels. And TypeSafe's own team built the workflows. They say so themselves.
The part I trust most doesn't depend on Jev at all. Every model scored higher, cheaper and faster as a workflow than as one big prompt. Small typed questions beat one giant prompt. That's the lesson I'm taking into my own harness.
What people built with it
Jev playing a Subway Surfers-style runner. Clip: @_MaxBlade
JevBot running from a zombie in Minecraft. Clip: Hyperion (YouTube)
A simulated drone deciding to climb over a pipe. Clip: @RomanSlack1
Jev playing Doom, from TypeSafe's launch. Clip: @typesafeai
Wikiracing, Jev against three LLMs, from TypeSafe's launch. Clip: @typesafeai
All of these are games or simulators with a short list of moves. Nobody is driving a real car with this.
Where Jev goes in my blog scout

Every morning a scout reads new papers and lab news, Claude scores each one, and I accept or skip what it surfaces. Most of those calls are quick "is this worth my time" judgments, which is exactly the kind of call Jev is built for.

- Shadow first. Jev scores the same candidates as Claude and decides nothing. That's been running since Sep 20.
- Measure it against me, not against Claude. My accept and skip calls are the ground truth, including the candidates Claude threw out, which I'll start keeping.
- Then a cheap first pass. If it holds up, Jev screens everything and Claude only reads the shortlist.
I keep the final say either way. It's too early for numbers. I'll share them when there's enough data to mean something.
Why I care

AlphaGo paired a fast policy network, which suggests a few good moves, with a slow search that checks them. Fast guess, slow check. That's the shape I want.
The game and robot demos are the same idea in a tight loop: decide ten times a second, cheaply.
And the practical pain: last Saturday my Claude pool hit 94% of its window and my router started refusing judgment work. A lot of those calls were small yes/no decisions that never needed a frontier model.
Where Jev goes in my RSI harness

My harness improves itself in a loop. It collects failures, has Claude propose one fix, replays that fix against history, and opens a draft PR that I merge. Jev fits in four places:
- As the router itself, instead of a growing pile of regexes.
- Triage before Claude, so the expensive proposer only runs on clusters worth a round.
- A quick filter on Claude's proposals before the replay.
- A cheap judge for evals, checked against human labels.
The replay keeps the final say. Jev never merges anything. (Related: yesterday I wrote about why self-improving harnesses need a regularizer.)
"How is this different from TabFM?"

A teammate asked this, and it's a fair question. Both skip the training step, both give you probabilities from one forward pass, and both were pretrained on synthetic data.
The difference is where the task comes from. TabFM (Google, June 2026) reads a table. You hand it labelled rows as context and it predicts the label for new rows; fit() just stores them. Jev never sees labelled examples. It learns the task from how you word the question and the options.
So for my scout, TabFM could learn straight from my past accept/skip calls, if I turned each candidate into table columns. Jev can't use those labels at all.
Rule of thumb: tables with a labelled history, TabFM. Messy text and a question, Jev. Running both on the scout is probably my next experiment.
What would make me trust it
- It holds up against my own calls, not just against Claude's.
- That includes the candidates Claude rejected, not only the ones it picked.
- Until then it stays in shadow, in the scout and in the harness.
Credits
- Jevpilot, a 3D driving sim where Jev makes the calls: @jpschroeder
- Jev playing a Subway Surfers-style runner: @_MaxBlade
- JevBot running from a zombie in Minecraft: Hyperion (YouTube)
- A simulated drone deciding to climb over a pipe: @RomanSlack1
- Jev playing Doom, from TypeSafe's launch: @typesafeai
- Wikiracing, Jev against three LLMs, from TypeSafe's launch: @typesafeai
- Same 27 questions, Jev vs an LLM, from TypeSafe's launch: @typesafeai
- Screenshots: typesafe.ai, docs.typesafe.ai, evals.typesafe.ai
- Architecture measurements: archerhume, jev-from-scratch
- Diagrams: mine. Art: nano-banana.
- 01TypeSafe AI — Introducing System One Models & Jev
- 02TypeSafe docs — Quick start
- 03TypeSafe — Workflow evals
- 04archerhume — Jev's Architecture Unmasked
- 05Maverick-Ansh/jev-from-scratch
- 06TechCrunch — A new kind of AI model from a ChatGPT inventor is thrilling developers
- 07Wikipedia — Jev (AI model)
- 08Google Research — Introducing TabFM
- 09Jevpilot, a 3D driving sim where Jev makes the calls (@jpschroeder)
- 10Jev playing a Subway Surfers-style runner (@_MaxBlade)
- 11JevBot running from a zombie in Minecraft (Hyperion (YouTube))
- 12A simulated drone deciding to climb over a pipe (@RomanSlack1)