Writing
September 23, 2026 · 10 min read

Jev picks. It doesn't write.

A visual walk through TypeSafe's Jev: what it is, what's actually known about how it's built, what "zero hallucination" really means, and where I plan to use it: my blog scout and my self-improving harness.

jevsystem-one-modelscalibrationevaluationself-improvement

This is the deck for a talk I'm giving my team about Jev. I kept the words short on purpose. Scroll it like slides.

Nothing here is writing text

A white car in a 3D city sim ('Jevpilot', 'Jev engaged') waits at 0 km/h at a stop-sign intersection, then pulls away (10, then 18 km/h) and turns left along a blue planned-path arc, while the HUD shows Jev's current pick with its probability ('Trim left · 100%', 'Trim left · 99%', then 'Left · 100%'). Jevpilot, a 3D driving sim where Jev makes the calls. Clip: @jpschroeder

Every frame, the model picks one of a handful of moves. No sentences, nothing to parse. That's Jev.

What it is

TypeSafe AI's homepage: "The First (Public) System One Model; Jev Gives AI The Properties Of Code" typesafe.ai, Sep 23, 2026

Jev is TypeSafe AI's first "System One" model. You give it some state and a few typed questions. It gives back a pick for each question, with probabilities. It never writes a sentence.

Where the name comes from

Two systems of thought: a slow winding path on the warm side, a fast spark on the cool side

Kahneman's split from Thinking, Fast and Slow. System 1 is the fast gut call: is that a face, should I brake. System 2 is the slow, effortful kind: long division, writing an essay.

We built LLMs for System 2 work, then use them for System 1 calls all day long. Jev only does the gut call.

Writing vs picking

Animation: an LLM builds "The car should probably brake" one token at a time while Jev scores all four options in a single pass

An LLM builds its answer one token at a time, and every token waits for the one before it. Jev scores every option at once, in one pass.

Two terminals, 'Same 27 questions. Same order.' The TypeSafe panel has already returned all 27 typed answers with probabilities and confidences (cost $0.000081, completed in 0.114s). The LLM panel (gpt-5.6-terra) waits for its first token, streams plain answers, and finishes in 8.566s at $0.013880. The header then reads 'TYPESAFE 74.9× FASTER • 171.0× CHEAPER'. Same 27 questions, Jev vs an LLM, from TypeSafe's launch. Clip: @typesafeai

TypeSafe's own side-by-side: the same 27 questions to Jev and to an LLM. Jev is done in 0.114 s. The LLM takes 8.6 s.

What goes in, what comes out

Animation: state and three typed questions flow into Jev and three typed answers come out at the same moment

State goes in. Questions go in, each with options you define. Typed answers come out, all at the same time.

Three question types: Choice (up to 255 options), Score (2 to 10 ordered levels) and Noul (one probability). Every answer comes with a confidence.

Under the hood

Animation: shared state encoded once, a transformer block, three isolated question lanes whose options interact, probability readouts; colour marks what is confirmed, measured by outsiders, or unknown

Here's what is actually known about the network, sorted by how sure we can be.

  • TypeSafe says: it's transformer-based but not a language model. They call it "a new model architecture" with a parallel sampler. All answers come out in one pass, each question is answered on its own against the shared state, and it's trained only on synthetic data.
  • Outsiders measured it by poking at the API (not confirmed): a decoder LLM backbone with the text output swapped for a probability readout, a tokenizer closest to Qwen's, probably mixture-of-experts at roughly 10B active parameters. The state gets processed once and shared. Options inside one question can see each other; separate questions can't.
  • Nobody outside knows: the size, the base model, the training data, or even encoder vs decoder. There's no paper.

Outside measurements: archerhume, "Jev's Architecture Unmasked" and the 6.6M-parameter rebuild in jev-from-scratch.

How it's trained

TypeSafe's table comparing existing LLMs with System One + Jev: RLHF/RLVR vs RLCD, strings vs type-safe values, sequential vs parallel sampling, cost and speed From TypeSafe's launch post

LLM chat models get tuned with RLHF: the model is rewarded when a human rater likes the answer. Jev gets tuned with what TypeSafe calls RLCD, reinforcement learning for calibrated decisions. It's rewarded when its probabilities match what actually happened.

Animation: RLHF loop rewarded by a rater's thumbs-up vs RLCD loop rewarded when "brake 70%" matches the outcome

The funny part: Diogo Almeida, TypeSafe's CEO, spent about four years at OpenAI working on RLHF, InstructGPT and ChatGPT.

What "calibrated" buys you

Animation: a reliability curve moves from overconfident to hugging the diagonal

When a calibrated model says 70%, it's right about 70% of the time. That's what lets you put a threshold on it and trust the threshold.

One outside test measured a calibration error of 0.031 on 1,200 MMLU questions. Good, not perfect.

"Zero hallucination"

TypeSafe's chart of structured-output and tool-call error rates, Jev at 0%, followed by the note "Our number is not empirical. Schema matching is guaranteed." From TypeSafe's launch post. Read the last bullet.

TypeSafe puts Jev at 0% on these charts, and then tells you right underneath: "Our number is not empirical. Schema matching is guaranteed."

Animation: an LLM's answer runs off the rails while Jev can only land on one of the given options, once right and once wrong

So here's what "can't hallucinate" really means. It can't make up an answer that isn't on your list. It can still pick the wrong thing from the list. That second part is the one you have to measure yourself.

How you call it

TypeSafe docs quick start: a Python call with a Choice, a Score and a Noul question docs.typesafe.ai

from typesafe_sdk import Choice, Noul, TypeSafeClient

client = TypeSafeClient()          # reads TYPESAFE_API_KEY
r = client.system_one(
    state="car at 42 km/h, obstacle 12 m ahead, lane 2, raining",
    questions={
        "action": Choice(instructions="What should the car do now?",
                         criteria={"brake": "slow down now",
                                   "maintain": "keep speed",
                                   "accelerate": "speed up"}),
        "is_safe": Noul(instructions="The current speed is safe"),
    },
)
r.answers["action"].choice   # "brake", plus probabilities and a confidence
r.answers["is_safe"].noul    # e.g. 0.12

TypeSafe quotes 70 to 500 ms end to end, $0.042 per million input tokens, and free output. Running it over 176 of my blog candidates cost me about a cent.

What TypeSafe's evals actually show

TypeSafe's workflow evals: mean accuracy vs cost per case, with Jev far to the left on cost evals.typesafe.ai

TypeSafe took four automation jobs: triaging security incidents, watching agent traces, processing invoices and answering customer service tickets. They ran each job two ways. Circles are one big prompt that does the whole job. Diamonds are a workflow: the job is broken into many small typed questions, and plain code makes the final call from the answers.

Every dot is one model, averaged over the four jobs. Higher means more accurate. Further left means cheaper per case, and each gridline is 10x, so small gaps are big differences. The line through the top-left dots is the frontier: nothing beats those on both cost and accuracy at once.

Jev sits alone on the far left: 67.8% accurate at $0.0004 a case and 0.4 s. The most accurate LLM workflow (sol) gets 74.1% at $0.08 and 23 s. So Jev gives up about six points of accuracy to be roughly 200x cheaper and 60x faster. Against an LLM at the same accuracy (terra, 67.9%), it's about 75x cheaper and 25x faster.

Two things before anyone quotes this. "Accurate" means agreeing with a reference answer built from two frontier LLMs (GPT-6 Astra and Fable 5.1), not with human labels. And TypeSafe's own team built the workflows. They say so themselves.

The part I trust most doesn't depend on Jev at all. Every model scored higher, cheaper and faster as a workflow than as one big prompt. Small typed questions beat one giant prompt. That's the lesson I'm taking into my own harness.

What people built with it

A runner in a neon Subway Surfers-style game dodges barriers and a train across three lanes. The right-hand 'JEV is driving' panel shows per-lane probabilities (LEFT/MID/RIGHT) with the chosen move (e.g. 'MIDDLE · JUMP', 'LEFT · RUN', 'RIGHT · JUMP') and highlights the W/A/S/D key it presses. Jev playing a Subway Surfers-style runner. Clip: @_MaxBlade

Seen from a second player's camera in a birch forest: a bot named 'JevBot', a zombie spawned nearby (a 'Zombie Spawn Egg' tooltip is visible), and the zombie catching fire in daylight. The bot's terminal log (top right) prints 'hostile_appeared … hostile_threshold_crossed', then 'JEV action=FLEE confidence=0.93 danger=1.9 latency=173ms' and 'ACTION FLEE success (Fleeing from zombie toward (-10.7, 36.3))'. JevBot running from a zombie in Minecraft. Clip: Hyperion (YouTube)

A simulated quadrotor (chase camera) comes up to a long yellow pipe and climbs over it. On the right, the 'JEV SYSTEM ONE' panel shows the 64x48 onboard depth image, free space by sector, and 'TACTICAL JUDGMENT' probabilities, with 'climb' jumping to about 95% and 'CLIMBING OVER BARRIER' showing at the bottom. A simulated drone deciding to climb over a pipe. Clip: @RomanSlack1

A 'DOOM // OPS' console. Left: Doom gameplay with the shotgun firing at cacodemons and imps, then switching to the fist. Right: live Jev judgment bars for FIRING, GOAL, DODGE and MOVEMENT that shift as the fight goes on. The status line reads 'ATTACKING ENEMIES, FOCUSING ON CACODEMON A'. Jev playing Doom, from TypeSafe's launch. Clip: @typesafeai

The 'TypeSafe AI WIKIRACE' board for Baseball → Sun: the countdown hits 1 and the race starts. Jev 1.13.0 shows '1st place · Finished · 0.419 s · 3 hops' about a second later, while GPT-5.6 Terra, Claude Haiku 4.5 and Claude Sonnet 5 are still clicking through Wikipedia pages (Baseball, Sport, Home run, Energy expenditure) and the race clock runs past 0:05. Wikiracing, Jev against three LLMs, from TypeSafe's launch. Clip: @typesafeai

All of these are games or simulators with a short list of moves. Nobody is driving a real car with this.

Where Jev goes in my blog scout

The blog scout pipeline: papers and lab news, Claude assesses, Jev shadow-scores, I accept or skip

Every morning a scout reads new papers and lab news, Claude scores each one, and I accept or skip what it surfaces. Most of those calls are quick "is this worth my time" judgments, which is exactly the kind of call Jev is built for.

Animation: now, Jev shadow-scores next to Claude and decides nothing; next, Jev does a first pass, Claude reads only the shortlist, and my accept/skip calls check Jev

  1. Shadow first. Jev scores the same candidates as Claude and decides nothing. That's been running since Sep 20.
  2. Measure it against me, not against Claude. My accept and skip calls are the ground truth, including the candidates Claude threw out, which I'll start keeping.
  3. Then a cheap first pass. If it holds up, Jev screens everything and Claude only reads the shortlist.

I keep the final say either way. It's too early for numbers. I'll share them when there's enough data to mean something.

Why I care

A Go board: a fast teal intuition picks a move and a slow amber search tree grows from it

AlphaGo paired a fast policy network, which suggests a few good moves, with a slow search that checks them. Fast guess, slow check. That's the shape I want.

The game and robot demos are the same idea in a tight loop: decide ten times a second, cheaply.

And the practical pain: last Saturday my Claude pool hit 94% of its window and my router started refusing judgment work. A lot of those calls were small yes/no decisions that never needed a frontier model.

Where Jev goes in my RSI harness

Animation: the self-improvement loop, requests, router, failures, harvest, Claude proposes a fix, replay, draft PR, human merge, with four Jev slots lighting up

My harness improves itself in a loop. It collects failures, has Claude propose one fix, replays that fix against history, and opens a draft PR that I merge. Jev fits in four places:

  1. As the router itself, instead of a growing pile of regexes.
  2. Triage before Claude, so the expensive proposer only runs on clusters worth a round.
  3. A quick filter on Claude's proposals before the replay.
  4. A cheap judge for evals, checked against human labels.

The replay keeps the final say. Jev never merges anything. (Related: yesterday I wrote about why self-improving harnesses need a regularizer.)

"How is this different from TabFM?"

Animation: TabFM takes labelled table rows as its context and predicts a query row; Jev takes text plus a worded question and never sees the labelled rows

A teammate asked this, and it's a fair question. Both skip the training step, both give you probabilities from one forward pass, and both were pretrained on synthetic data.

The difference is where the task comes from. TabFM (Google, June 2026) reads a table. You hand it labelled rows as context and it predicts the label for new rows; fit() just stores them. Jev never sees labelled examples. It learns the task from how you word the question and the options.

So for my scout, TabFM could learn straight from my past accept/skip calls, if I turned each candidate into table columns. Jev can't use those labels at all.

Rule of thumb: tables with a labelled history, TabFM. Messy text and a question, Jev. Running both on the scout is probably my next experiment.

What would make me trust it

  • It holds up against my own calls, not just against Claude's.
  • That includes the candidates Claude rejected, not only the ones it picked.
  • Until then it stays in shadow, in the scout and in the harness.

Credits

References
  1. 01TypeSafe AI — Introducing System One Models & Jev
  2. 02TypeSafe docs — Quick start
  3. 03TypeSafe — Workflow evals
  4. 04archerhume — Jev's Architecture Unmasked
  5. 05Maverick-Ansh/jev-from-scratch
  6. 06TechCrunch — A new kind of AI model from a ChatGPT inventor is thrilling developers
  7. 07Wikipedia — Jev (AI model)
  8. 08Google Research — Introducing TabFM
  9. 09Jevpilot, a 3D driving sim where Jev makes the calls (@jpschroeder)
  10. 10Jev playing a Subway Surfers-style runner (@_MaxBlade)
  11. 11JevBot running from a zombie in Minecraft (Hyperion (YouTube))
  12. 12A simulated drone deciding to climb over a pipe (@RomanSlack1)