← Pooja Verma
BenchmarkedIssue · 25 Sep 2026
Product launch · about a 6-minute read

Jev: the AI model that can’t write a sentence

TypeSafe AI’s new model only makes typed decisions, fast and very cheaply. Early independent tests back the speed and cost, but its accuracy looks like a good small model’s, and it always answers, even when it shouldn’t.

Diagram: an LLM writes a paragraph about an invoice, while Jev returns scored options: fraud 0.07, clean 0.88, review 0.05
An LLM writes its answer out; Jev returns only the decision, with a probability for each option. Source: LangChain.

The short version

  • A new kind of model: Jev picks options, gives scores and estimates probabilities, but produces no text, code or chat. vendor
  • Very cheap: $0.042 per million input tokens, with output free; TypeSafe says replies take 70–500 ms. vendor
  • Speed and consistency hold up: LangChain measured 0.44 s and $0.00035 per judge call, with score variance 92–913× lower than the LLM judges. independent
  • Not smarter than LLMs: on 791 decisions, AY Automate found it roughly level with small models. independent
  • It always answers: without a “none of these” option, it flagged 0 of 30 out-of-scope messages, at 0.99 confidence. independent
  • Hard to get: new signups have been paused since 21 Sep (US Pacific). vendor

01What TypeSafe launched

TypeSafe AI, founded by former OpenAI researcher Diogo Almeida with Erik Gafni and Sasha Sheng, released Jev in early access on 15 Sep 2026. A paid press release announced a $40M seed round led by DCVC. press

TypeSafe calls Jev a “System One model”, after Daniel Kahneman’s fast, intuitive thinking. You send a state (text or JSON) plus any mix of three question types, and all of them are answered in parallel in one call: vendor

Choice, Score and Noul examples, each with a confidence value
The three question types. Every answer comes back typed, with probabilities and a confidence score. Source: LangChain.

Underneath is a new architecture, a “parallel sampler” that doesn’t generate token by token, and a training method TypeSafe calls RLCD (Reinforcement Learning for Calibrated Decisions). The name nods to economist William Stanley Jevons: cheaper intelligence means far more use. vendor

02TypeSafe’s own numbers

The headline claims don’t agree with each other. The launch post says “20–200x faster” and “40–400x cheaper”; the blog says 40–200× faster; the home page says 193.6× faster and 444.6× cheaper. TypeSafe’s own blog says the home-page figures are “on the higher end of real world gains”. vendor

Scatter chart of accuracy against cost per workflow: Jev at about 68% and the lowest cost, top models at about 74%
TypeSafe’s chart across its own four workflows: Jev agrees with the reference answer about 68% of the time, against roughly 73–74% for the top models, at a fraction of the cost. Source: TypeSafe.

The takeaway: Jev isn’t smarter than frontier models, but lands near mid-tier ones for far less. TypeSafe notes that its own team wrote the workflows, so “some bias could exist”, and that the reference answer averages GPT-6 Astra and Fable 5.1. vendor

03What independent testers found

TestResultKind
LangChain, agent-eval judge (5 runs, each scored 100×)Pass/fail agreement with a human: Jev 100%,
Terra 99.8%, Luna 96.4%, Claude Sonnet 4.6 80.0%
independent
LangChain, cost and speed0.44 s, $0.00035 per call;
$0.34 total vs $28.17 for Claude
independent
AY Automate, intent routing and injection detection (791 items)78.8–87.0%, about small-model level;
3.6× faster than GPT-5.6 Terra
independent
PriorBench, zero-shot on 400 items95.9% (keywords 77.2%, TF-IDF 66.0%)independent
HiringCafe, resume–job matchingSpearman 0.79 (cheap LLMs 0.72–0.77)community
@fazxes, safety classifier98.6% vs GPT-5.6 Luna 96.7%;
4.7× faster at the median
community
LangChain chart: Jev's score is a flat line across 100 repeats, while three LLM judges swing up and down
The same five cases, scored 100 times each: Jev’s line is flat, while the LLM judges wobble. Only five test cases in one domain, so promising rather than proof. Source: LangChain.
Post listing Spearman scores: Gemini 3.1 Flash-Lite 0.72, DeepSeek V4 Flash 0.73 and 0.77, Jev 0.79 at $0.02
HiringCafe’s test compared Jev only with cheap LLMs, at about a tenth of the cost of the next-cheapest option (the thread gives no cost unit). Source: @h_nilforoshan.

04What builders are doing with it

Unless marked otherwise, these are the builders’ own claims. community

Results table: 314 filled and 753 blank IRS form pages, 0 wrong
A tax-form classifier covering 261 IRS forms: 0 wrong across 1,067 pages, at about $0.001 per page. Source: kyotofin/tax-doc-classifier on GitHub.

Builders have converged on one pattern, summed up by @thegreatest_sv: “the big model plans, Jev picks, code does the rest.”

05Where it falls short

PriorBench note that Jev always answers, above a chart of accuracy by confidence threshold
PriorBench’s takeaway: accuracy is flat between the 0.50 and 0.95 confidence thresholds, then hits 100% at 0.99, covering 60.2% of traffic. Anonymous author, one day, one version. Source: PriorBench.

06Price and access

Jev costs $0.042 per million input tokens ($42 per billion), and output is free. Rate limits (250K tokens a second, 1,200 requests a minute) are “adjusting dynamically”. Servers are on the US West Coast; PriorBench saw a floor of about 430 ms from Europe through OpenRouter. vendor

TypeSafe blog: “We can’t prove it isn’t subsidized”
TypeSafe on its own pricing, which it expects to go down. Source: TypeSafe.

Access has been bumpy. A waitlist on 15 Sep gave way to open access on 20 Sep, then new signups were paused on 21 Sep (US Pacific) “due to demand”. Existing users keep working, and no reopening had been announced as of 25 Sep. vendor

07How it compares

OptionWhat the evidence saysKind
LLMs (GPT-5.6, Claude, Gemini Flash)Smarter and can write; slower and pricier for fixed-choice decisionsindependent
Jev first, Terra for unsure casesTerra-level accuracy at 26–28% of Terra’s cost (AY Automate)independent
Fine-tuned classifierNeeds training data, but may match Jev on one narrow taskcommunity
Laya (open weights, local)Claimed 11× faster decisions than Jev at Tetris, unverifiedcommunity
Span-01 (Respan)Claims “18% better than Jev” on its own benchmarkcommunity
Post by Nathan Flurry: “jev is just a *really* smart switch statement”
The skeptic’s summary, and the pattern he suggests: an LLM proposes options, Jev decides, code executes. Source: @NathanFlurry.

The interface is easy to copy: @tinkerapi says a $5 fine-tune of an open model can offer something similar (unverified). The open question is whether Jev’s quality and calibration hold up at this price. community

08Handle with care

  • “100× faster and cheaper” gets repeated widely, but TypeSafe’s own multiples range from 20× to 444.6×, and it calls the top ones “the higher end”. “Often tens to hundreds of times cheaper” is safer.
  • “Can’t hallucinate” is narrow: Jev always returns a valid option, but it can still pick the wrong one.
  • The Deel results (expense categorization 50% → 86%, 20–59× cheaper) were published by TypeSafe, not Deel, and the comparison LLM is unnamed.
  • LangChain co-hosted a livestream with TypeSafe, so its test isn’t fully arm’s-length.
  • Some secondary coverage gets basics wrong, including the founder’s name and crediting TypeSafe with the Minecraft demo, which was a community project.
  • A valuation figure attributed to Forbes in secondary coverage is unverified, so it isn’t reported here.
  • Signup status may have changed since 25 Sep.

09Sources

Benchmarked · the AI news, fact-checked · Facts as of 25 Sep 2026 · labels show who reported each number: vendor, independent, press, community