My verdict on how to run Jev locally is simple: you can't run the real Jev locally because TypeSafe hasn't released any weights, and the best local stand-in I've tested, Laya, is about seven times faster than Jev but clearly less accurate until you train it.
That's the whole verdict in one line.
Everything below is the evidence, round by round.
I ran Laya 0.3.4 on my own Mac on 21 September, against real Jev calls on the same test data.
Where a number came from somebody else, I'll tell you who published it.
Where I haven't tested something, like Cloudflare's new Clef models, I'll say that too.
The Verdict At A Glance
| Question | My answer |
|---|---|
| Can you run Jev locally? | No, Jev is hosted only and has no public weights. |
| Best local stand-in I've tested | Laya, free under Apache 2.0, from Convai Innovations. |
| Where Laya wins | Speed, cost per question, multilingual routing and calibration after tuning. |
| Where Jev wins | Accuracy out of the box and big lists of options. |
| Newest open contender | Cloudflare's Clef, released 1 October 2026, which I haven't tested yet. |
| Who should go local | Teams with fewer than 20 answer options who will train on their own data. |
| Who should stay on Jev | Anyone with big option lists or no time to train. |
Why "Run Jev Locally" Has No Direct Answer
Jev is TypeSafe AI's decision model, and it launched on 15 September 2026.
It doesn't write text.
It reads a situation, picks one answer from the list you supply, and returns a probability with it.
TypeSafe's site describes early access, a console sign-in and pricing per billion input tokens.
There's no mention anywhere on it of weights, on-premise deployment or an offline build.
Every route I've used to reach Jev, including OpenRouter, Vercel AI Gateway and OpenCode Zen, sends requests to a hosted model.
So any page promising a "Jev download" is really offering a different model that copies Jev's request shape.
That isn't automatically a bad thing.
It just means the real question is which local copy is worth your time.
Meet The Contenders
There are three things worth comparing here.
Jev is the hosted original, priced on OpenRouter and Vercel at $0.042 per million input tokens with output free.
Laya is a set of three small open models behind a Router, and its server copies Jev's POST /v1/systemone request format, according to its README.
Clef is Cloudflare's pair of open-weight decision models, Clef and the smaller Clef-flash, which Cloudflare says are fully Jev-API compatible.
| Spec | Jev | Laya | Clef |
|---|---|---|---|
| Who makes it | TypeSafe AI | Convai Innovations | Cloudflare |
| Runs locally | No | Yes | Yes, weights on Hugging Face |
| Licence | Proprietary, hosted | Apache 2.0 | Apache 2.0 |
| Model size | Not published | 322M to 421M parameters per model | 27B base for Clef, 9B for Clef-flash |
| Context window | 32k | 512 to 8,192 tokens depending on model | 64k |
| Max answer options | 255 | Limited by each model's room | Not tested by me |
| Tested by me | Yes | Yes | No |
Round 1: Speed
Laya wins this round comfortably.
| Setup | Time per question | Source |
|---|---|---|
| Laya with a GPU, one question | 33 ms | Laya team |
| Laya, ten questions in one call | 7 ms each | Laya team |
| Laya on a plain CPU | 193 to 464 ms | Laya team |
| Laya on my Mac | about 20 ms | My test |
| Jev, one question | 236 to 276 ms | Independent published tests |
| Jev from my desk through OpenRouter | about 350 ms per call | My test |
One warning stops this being a clean win.
Out of the box, Laya reloaded a model every time my test messages switched language, and that cost about 20 seconds per answer.
Setting Router(preload=True) fixed it, and the same messages dropped to about 19 ms each.
Round 1 verdict: Laya wins, as long as you turn on preload.
Round 2: Accuracy On Published Benchmarks
The Laya team published four shared tests against Jev's published numbers.
| Test | Laya | Jev |
|---|---|---|
| Typed decisions, 2,000 business choices | 0.766 | 0.727 |
| News sorting, four categories | 0.950 | 0.910 |
| Emotion, six feelings | 0.595 | 0.480 |
| Banking, 77 categories | 0.425 | 0.870 |
On paper, Laya wins four of five.
But there's a big asterisk I flagged in my Laya guide.
The 0.766 came from the Laya model trained on the practice version of that exact test.
The two untrained Laya models scored 0.362 and 0.342 on the same test.
Always picking the most popular answer scored 0.461, and random guessing scored 0.318.
The Laya team also measured its numbers on its own hardware, and they said themselves that they had no Jev access, so Jev's figures came from other people.
Round 2 verdict: Laya wins the scoreboard, but only the trained model earns that win.
Round 3: Accuracy On My Own Tasks
This is the round I care about most, because it's real work rather than a leaderboard.
| My test | Jev | Laya, untrained |
|---|---|---|
| 60 test emails into folders | 54 / 60 | 36 / 60 and 28 / 60 across two models |
| 24 contact-form messages: lead, vendor or noise | 24 / 24 | 16 / 24 |
| 10 article briefs matched to the right site | 10 / 10 | 8 / 10 |
Jev wins every row.
The one point I'd give Laya is honesty.
On the email test it never once claimed to be 85% sure, so it knew when it was guessing.
Round 3 verdict: Jev wins clearly on untrained, real-world sorting.
Round 4: Big Lists Of Options
Jev takes this one easily.
On the 77-category banking test, Jev scored 0.870 and Laya scored 0.425.
The reason is a design choice, not a bug.
Laya has a fixed amount of room for your list, so four options get roughly 60 words each while 77 options get three or four words each.
Jev handles up to 255 options.
The workaround is to split one big question into two smaller ones, but that's extra work Jev doesn't need.
Round 4 verdict: Jev wins whenever you have more than about 20 answers.
Round 5: Can You Trust The Sureness Number?
The probability is the whole point of these tools, because it tells your agent when to act and when to ask.
On calibration, where lower is better, Jev measured 0.246 and tuned Laya measured 0.081.
That means when tuned Laya says 90%, it's right about 90% of the time.
Untuned, Laya's English model sat at 0.466, which makes the number close to decoration.
The multilingual model ships with no tuning at all, so you'd need to tune it yourself.
Jev has its own weak spot here too.
On the emotion test, it gave the right answer a 0% chance 16% of the time.
Round 5 verdict: Tuned Laya wins, untuned Laya loses.
Round 6: Languages
Laya's Router is a genuine advantage.
It checks the language of each message before running anything, then sends it to the English model, the multilingual model or the paperwork model.
The Laya team reported 45 of 51 languages working well with the Router, against only 23 without it.
On my Mac I sent the same billing complaint in Hindi, Thai and Khmer.
The English model alone couldn't read it, but the Router sent it to the multilingual model and got "billing" every time.
I didn't run the same multilingual test on Jev, so I can't score Jev fairly here.
Round 6 verdict: Laya impresses, but this one isn't a fair head-to-head.
Round 7: Cost And Setup
Laya costs nothing per question once it's installed.
Jev costs $0.042 per million input tokens through OpenRouter or Vercel, with output free.
Vercel's free promo pricing ended on 25 September 2026.
OpenCode Zen still lists a free jev-1.13-free model for a limited time, and jevplayground.com lets you try the real model with no signup.
On setup, Jev wins for speed to first answer, because a hosted call needs no install.
Laya needs Python 3.10 or newer, a pip install laya, and ideally an afternoon on Kaggle's free GPUs to train it.
Round 7 verdict: Laya wins on running cost, Jev wins on time to first answer.
The Final Scoreboard
| Round | Winner |
|---|---|
| Speed | Laya |
| Published accuracy | Laya, trained model only |
| My real-world accuracy | Jev |
| Big option lists | Jev |
| Sureness number | Laya, after tuning |
| Languages | Laya, not a fair head-to-head |
| Cost per question | Laya |
| Time to first answer | Jev |
What About Clef?
Clef deserves a mention because it's the most serious new local option.
Cloudflare's blog says it released Clef on 1 October 2026 with open weights on Hugging Face under Apache 2.0.
Clef is post-trained from Qwen 3.8-27B with a vision encoder, and Clef-flash is a 9B sibling.
Cloudflare reported median latencies of 209.3 ms for Clef and 38.8 ms for Clef-flash, and claimed wins over Jev on several benchmarks.
Those are Cloudflare's own numbers on its own infrastructure, not independent tests.
I haven't run Clef myself, so it doesn't get a score from me today.
The obvious trade-off is size, because a 27B model needs far more hardware than Laya's sub-500M models.
Who Should Run A Local Jev Alternative
Pick Laya if your questions have fewer than 20 answers, you care about speed or keeping data on your own machine, and you're willing to spend an afternoon training it.
Stay on hosted Jev if you have big option lists, need accuracy on day one, or don't want to maintain anything.
Watch Clef if you have serious GPU hardware or want vision and a longer context, but test it on your own data before you trust it.
Use both if you want the best of each, because Laya's server and Jev share a request format, so switching is one URL change.
If you're comparing models for coding rather than decisions, my best AI models for coding ranking is the better read.
And if you like this round-by-round format, my DeepSeek Harness review and Claude Opus 5 vs GPT-5.6 comparison use the same approach.
๐ฅ Want to see how I put a decision layer in front of my agents? Inside the AI Profit Boardroom, I share the Laya and Jev decision-layer walkthroughs as I build them, sitting in front of Claude Code, Hermes and OpenClaw. Plus weekly coaching calls + 3,400+ members building real automations. โ Get access here
Related Reading
๐บ Video notes + links to the tools ๐
๐ฅ Learn how I make these videos ๐
๐ Get a FREE AI Course + Community + 1,000 AI Agents ๐
๐ฅ Not sure which decision model fits your setup? Bring your question to one of the four weekly coaching calls inside the AI Profit Boardroom, where 3,400+ members are testing these tools on real work. โ Join the AI Profit Boardroom
Also On Our Network
- ๐ How to set up a local Jev stand-in, step by step
- ๐ What running Jev locally means for your business costs
- ๐ Plugging a local Jev into your Agent OS
- ๐ The full Laya AI test against Jev
FAQ: Local Jev Questions Answered
Is Jev open source?
No, Jev is a proprietary hosted model from TypeSafe AI and its weights are not public.
That's why you can't run Jev locally.
Is Laya better than Jev?
Laya is faster and free per question, and the trained version beat Jev on several published tests.
Untrained, it lost to Jev on every one of my own sorting tests, so it depends on whether you'll train it.
Is Clef better than Jev?
Cloudflare claims Clef beats Jev on several benchmarks, but those are Cloudflare's own numbers.
I haven't tested Clef yet, so I can't give you an independent verdict.
Can you run Jev locally with Ollama?
No, because there's no Jev model file to load into Ollama or any other local runner.
You'd be running a different model, such as Laya or Clef, that copies Jev's request format.
What's the cheapest way to use real Jev?
OpenCode Zen lists jev-1.13-free as free for a limited time, and jevplayground.com lets you try it with no signup.
Otherwise it's $0.042 per million input tokens through OpenRouter or Vercel, with output free.
About Julian
I'm Julian Goldie, an AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom, which has 3,400+ members.
I help business owners scale with AI agents, automation, and SEO.
- I have 400,000+ YouTube subscribers who watch my AI tool tests every week.
- I built a 7-figure agency from the ground up.
- I run daily AI training inside the Boardroom.
- I wrote two Amazon best-sellers on SEO and agency growth.
โ Get my best AI training inside the AI Profit Boardroom
Final verdict: you can't run the real Jev locally, so if you want to run Jev locally in practice, train Laya on your own data for speed and privacy, and keep hosted Jev for big lists and day-one accuracy.











