Skip to main content
PromptQuorumBuilt for humans. Structured for AI.

Julia-1 vs Jev: Local vs Cloud Decision Models (2026)

Julia-1 is a 144.3M-parameter, Apache 2.0 decision model built on mmBERT-small that runs locally on CPU for classification and routing tasks. Jev, the cloud decision model TypeSafe AI launched days earlier to viral press coverage, is a closed API with no published parameter count. Both score a fixed list of answer choices instead of generating text β€” the difference is where each one runs.

Quick Answer

Julia-1 and Jev are both "decision models" built to score a fixed list of answer choices rather than generate free text, but they sit at opposite ends of the deployment spectrum. Julia-1 is a 144.3M-parameter, Apache 2.0 model from Supersonic Labs that runs on CPU with no API key or per-token cost. Jev is TypeSafe AI's closed cloud API, priced per input token, with no published parameter count or open weights.

  • β–ΈJulia-1: 144.3M parameters, Apache 2.0, runs locally on CPU, ~550 MB on disk
  • β–ΈJev: cloud-only API from TypeSafe AI, $0.042 per million input tokens, architecture undisclosed
  • β–ΈBoth score fixed answer choices instead of generating text β€” the same emerging "decision model" category
Model ComparisonsIntermediate

Key Takeaways

  • βœ“Julia-1 (Supersonic Labs) is a 144.3M-parameter, Apache 2.0 decision model built on mmBERT-small β€” it runs on CPU, needs no API key, and weighs ~550 MB on disk
  • βœ“Jev (TypeSafe AI) is a closed, cloud-only decision model that launched to viral press coverage on September 15, 2026 β€” pricing is $0.042 per million input tokens, architecture undisclosed
  • βœ“Both score a fixed list of 2–20 candidate answers instead of generating text β€” the same "decision model" category, opposite deployment models
  • βœ“No public benchmark has tested Julia-1 and Jev head-to-head on the same task

What Is Julia-1?

Julia-1 is a 144.3M-parameter decision model built by Supersonic Labs on top of mmBERT-small, a multilingual encoder from Johns Hopkins' Center for Language and Speech Processing (CLSP). Given a context, a question, and 2 to 20 candidate answers, Julia-1 scores every option and returns the best one β€” it does not generate free text.

The model keeps mmBERT-small's attention layout: every third layer attends globally across the full input, and the other layers attend only within a Β±64-token local window. A 256,000-row vocabulary embedding accounts for 98M of the 144M total parameters, and a single request only reads the embedding rows for the tokens it actually uses β€” one reason the model runs on CPU without a GPU.

Julia-1 ships under the Apache 2.0 license, weighs 550.5 MiB on disk (FP32 weights), and accepts up to 8,192 combined tokens across context, question, and candidate answers. It targets classification, routing decisions, ordered scoring, and yes/no judgments β€” the kind of fixed-choice decision an AI agent makes before calling a tool or escalating a ticket, not open-ended generation.

Its base model, mmBERT-small, is separately notable: JHU CLSP trained it on more than 1,800 languages (1,833 in the final training phase) under the MIT license. That figure describes mmBERT-small's pretraining coverage β€” Julia-1's own accuracy has only been measured across 52 locales in the MASSIVE benchmark, so treat the 1,800-language number as a base-model fact, not a claim about Julia-1's multilingual accuracy.

How Accurate Is Julia-1?

Julia-1 scores 94/100 on AG News (4-label topic classification) and 86/100 on the DAIR Emotion benchmark (6 labels), both vendor-reported by Supersonic Labs. On the 52-locale MASSIVE intent benchmark it answers 110,573 of 154,648 items correctly (71.5%), and on a mixed set of typed decisions it scores 1,463 out of 2,000 (73.15%).

The weaker result is a 100-example pilot on Banking77, a 72-label banking-intent benchmark: Julia-1 scored 64/100 against an 87% reference score. Supersonic Labs' own model card is explicit that "100-example pilots are encouraging signals, not guarantees," and flags long, ambiguous label sets like Banking77's 72 categories as Julia-1's current weak point.

None of these numbers have been independently reproduced outside Supersonic Labs' own model card as of publication β€” treat them as vendor-reported until a third party verifies them.

BenchmarkLabelsScoreSource
AG News494/100Vendor-reported
DAIR Emotion686/100Vendor-reported
MASSIVE (52 locales)intent set71.5% (110,573/154,648)Vendor-reported
Banking77 (pilot)7264/100 vs. 87% referenceVendor pilot, 100 examples

Should You Run Julia-1 Locally or Call Jev's Cloud API?

Jev, launched by TypeSafe AI on September 15, 2026, is the cloud counterpart to Julia-1's local approach β€” and it launched to far more attention. The launch video drew roughly 40 million views on X, TypeSafe AI raised a $40 million seed round led by DCVC around the announcement, and investors have since approached the company at valuations above $10 billion, per Bloomberg's coverage.

Jev is built by a former OpenAI researcher and is positioned as a "System 1" model: it returns typed, calibrated decisions β€” not text β€” for routing, scoring, triage, and moderation, the same task category as Julia-1. Unlike Julia-1, Jev is cloud-API-only through partners such as DigitalOcean's Model Catalog. TypeSafe AI has not published Jev's parameter count or model architecture; outside observers only suspect, unconfirmed, that it is transformer-based on an open-weight foundation.

Pricing is per-token rather than free-after-download: Jev charges $0.042 per million input tokens with output tokens currently free, and TypeSafe AI reports 70–500 millisecond end-to-end latency under standard cloud API conditions. No public benchmark has tested Jev and Julia-1 head-to-head β€” the two have not been evaluated on the same task set.

For CPU-only local setups outside the decision-model category, see best Ollama models for CPU-only machines.

Frequently Asked Questions

What is a "decision model," and how is it different from an LLM?β–Ύ
A decision model takes a context, a question, and a fixed list of candidate answers, then scores or picks one β€” it never generates open-ended text. A general-purpose LLM can do this too, but it also has to generate, format, and parse a text response for every call, which costs more tokens and adds latency. Julia-1 and Jev are both decision models built specifically for this narrower job.
Can Julia-1 replace an LLM call for ticket routing?β–Ύ
For a fixed set of routing categories, yes in principle β€” that is the use case Supersonic Labs built it for. Its own benchmarks show 94% accuracy on 4-label topic classification (AG News) but only 64/100 on the 72-label Banking77 pilot, so accuracy depends heavily on how many categories you route between. Test it against your own category list before replacing a production LLM call.
Is Julia-1 free to use?β–Ύ
Yes β€” it ships under the Apache 2.0 license, so you can download and run it without a per-request fee. You still pay for your own compute, but Julia-1 needs no GPU and runs on CPU.
Does Jev run locally like Julia-1?β–Ύ
No. Jev is available only as a cloud API β€” through partners such as DigitalOcean's Model Catalog β€” and TypeSafe AI has not released its weights or architecture.
How much does Jev cost to use?β–Ύ
TypeSafe AI prices Jev at $0.042 per million input tokens, with output tokens currently free, as of its September 2026 launch.
What is mmBERT-small, and is it the same as Julia-1?β–Ύ
mmBERT-small is the 140M-parameter multilingual encoder from Johns Hopkins CLSP that Julia-1 is built on, trained on more than 1,800 languages under the MIT license. Julia-1 adds a decision-scoring head on top of it and is a separate, Apache 2.0-licensed release from Supersonic Labs, evaluated on 52 locales rather than the full 1,800+.
How accurate is Julia-1 on real-world tasks?β–Ύ
It varies by task: 94/100 on AG News, 86/100 on DAIR Emotion, 71.5% across 52 locales on MASSIVE, and 64/100 on a 100-example Banking77 pilot against an 87% reference score. All of these figures are vendor-reported by Supersonic Labs and have not been independently reproduced.
Has anyone directly compared Julia-1 and Jev on the same benchmark?β–Ύ
Not publicly, as of this writing. The two come from different vendors and target the same "decision model" category, but have not been tested head-to-head on a shared evaluation set.