Jev is an AI model that does not talk. You send it a situation and a list of typed questions. It sends back choices, scores, and yes/no probabilities that your code can act on.
For people who already know what an API is and have never used a System One model · Official docs: docs.typesafe.ai
Contents
- What Jev is
- What Jev is not
- The mental model
- The three question types
- Anatomy of one request
- Your first call
- How to write state
- Confidence, thresholds, and review
- Patterns you will actually use
- When to use Jev (and when not to)
- Models, price, and limits
- Common beginner mistakes
- Glossary
- Where to go next
1. What Jev is
Jev is TypeSafe AI’s first public System One model, announced on 15 September 2026. You give it two things:
- State — the situation. A support email, a ticket, a log line, a JSON snapshot of a page, a resume, a retrieved passage.
- Questions — named decisions you defined in advance. Each question is one of three types: Choice, Score, or Noul.
Jev answers every question in one parallel pass and returns typed values plus probabilities. There is no chat transcript. There is no paragraph to parse. Your program reads a field and branches.
“Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” — TypeSafe AI
The name is a nod to economist William Stanley Jevons. His paradox says that when something gets cheaper, people use more of it. TypeSafe’s bet is that cheap, fast decisions will show up in places you would never spend a chat model. “System One” is a Kahneman metaphor: fast judgment, not slow deliberation. It is a product name, not a claim that Jev thinks like a human.
TypeSafe was founded by Diogo Almeida, who previously worked on the RLHF methods behind InstructGPT / ChatGPT. The company positions Jev against chat models on decision work, not writing work.
2. What Jev is not
| Jev is | Jev is not |
|---|---|
| A decision model for software | A ChatGPT competitor for chatting |
| Choice / Score / Noul over your state | A writing, coding, or image assistant |
| Probabilities + confidence for thresholds | A guarantee that the answer is correct |
| Complementary to LLMs in a workflow | “JSON mode on a smaller LLM” |
| Text in, typed values out | A vision, audio, or video model |
Jev cannot draft the refund email. It can tell you, with a probability, whether the customer asked for a refund. An LLM can write the email after your code decides to send one.
Hallucination, sort of. Jev cannot invent a fourth department if you only listed three. The output is schema-bound. It can still be wrong about the text. That is what confidence and human-review gates are for.
3. The mental model
Most “AI in the product” stacks collapse three jobs into one chat model: understand, decide, write. Jev only takes the middle job.LLMJevYour codeWrites textSummarisesExplains to a humanPicks an optionScores a rubricSays yes / no, with pThresholdsPermissions, side effectsMath, dates, retriesFigure 1. Keep generation, judgment, and action in separate boxes. Jev only owns judgment.Stateticket / log / JSONQuestionsnamed + typedJevone parallel passAnswers+ p + confidenceFigure 2. One POST. State is ingested once. Every question is scored against that same state.
Three consequences fall out of this:
- Output tokens are free because there is almost nothing to generate. You pay for input.
- Questions are independent. Adding a 14th question does not rot the first 13. TypeSafe’s pitch is that this is the opposite of stuffing more instructions into a chat context.
- Policy lives in your code. Jev does not refund anyone. A threshold you chose does.
4. The three question types
Jev has exactly three primitives. Picking the wrong one is the most common early mistake.
| Type | You ask | You define | You get back | Use it for |
|---|---|---|---|---|
| Choice | Which of these options? | Up to 255 named options, each with a description | choice, probabilities, confidence | Routing, labeling, picking a tool, picking a next click |
| Score | Where on this scale? | An ordered rubric of 2–10 levels | score (weighted mean), legend, probabilities, confidence | Priority, severity, quality, intensity |
| Noul | Is this statement true? | A yes/no claim, optional true/false criteria | noul — a single probability from 0 to 1 | Flags, filters, guard checks, “does this match?” |
You can mix all three in one request. Every question sees the same state and is answered on its own.
Choice
Use Choice when the answer is one item from a closed list and the items are not ordered.
{
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery status, delays, lost packages",
"billing": "Charges, invoices, payment problems",
"other": "Anything that is not returns, shipping, or billing"
}
}
}
The response names the winning option and gives you the full distribution:
{
"type": "choice",
"choice": "returns",
"confidence": 1.0,
"probabilities": {
"returns": 1.0,
"shipping": 0.0,
"billing": 0.0,
"other": 0.0
}
}
confidence is not a second opinion. It is a summary of how peaked the distribution is. All mass on one option → high confidence. Mass split across two departments → low confidence, even if one of them “won.”
Always include an escape hatch such as other or none if the list might not cover the world. Otherwise Jev is forced to pick a wrong bucket.
Score
Use Score when the answer sits on an ordered spectrum you can describe in steps. Levels are numbered from 0 in array order. The returned score is the probability-weighted mean of those numbers, so it can land between levels.
{
"severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
Example answer:
{
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": { "0": 0.0, "1": 0.57, "2": 0.43 }
}
Read that as: mostly level 1, some evidence for level 2, not sure enough to treat it as a hard 1 or 2. A score of 1.0 with confidence 1.0 is a different beast from a score of 1.0 with probability split 50/50 on levels 0 and 2. Always look at probabilities.
Write levels as situations, not adjectives. “Broken feature, workaround exists” is usable. “Moderately severe” is not. The model does not see the numbers or compare levels to each other; it matches the state to each description.
Noul
Noul is TypeSafe’s word for a binary probability. It answers “is this true?” with a number from 0 (no) to 1 (yes). There is no separate confidence field. With only two outcomes, the single number is the whole distribution.
{
"refund_requested": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?",
"criteria": {
"true": "Asks to get money back, reverse a charge, or be refunded",
"false": "Complains or asks for options but does not request a refund"
}
}
}
A value of 0.94 means “very likely yes,” not “94% refunded.” A value of 0.5 means the model cannot tell. Phrase the question so that high = the event you care about. Avoid double negatives.
Noul is not a scale. “How angry is this customer?” is a Score. “Is the customer angry?” is a Noul. “Is the customer angry and asking for a refund?” is two Nouls, combined in code.What kind of answer do you need?Yes or no→ NoulOne of a listno order → ChoiceA position on a scaleordered levels → ScoreFigure 3. If you cannot say which box you are in, split the question until you can.
5. Anatomy of one request
Everything goes to one endpoint:
POST https://api.typesafe.ai/v1/systemone
Authenticate with a bearer token from the TypeSafe console. Header: Authorization: Bearer $TYPESAFE_API_KEY.
Minimum body:
{
"model": "jev-latest",
"state": "Please call me tomorrow morning to discuss the setup.",
"questions": {
"callback_requested": {
"type": "noul",
"instructions": "Does the sender explicitly ask for a phone call?"
}
}
}
Minimum useful response:
{
"model": "jev-1.13.0",
"answers": {
"callback_requested": {
"type": "noul",
"noul": 0.97
}
},
"usage": {
"input_tokens": 280,
"output_tokens": 22
}
}
jev-latestis an alias. As of late September 2026 it points atjev-1.13.0. The response tells you which version actually answered. Pin a version once you have tuned thresholds.statecan be a string, a JSON object, or an array of text values. No images, audio, or video. Convert those to text first.- Question keys (
callback_requested) are yours. They come back underanswerswith the same names. Treat them like function names: stable, specific, boring.
6. Your first call
1. Create a key at console.typesafe.ai/keys.
2. Export it locally. Do not put it in a repo.
export TYPESAFE_API_KEY="tsk_..."
3. Run this from a terminal.
curl --fail-with-body --silent --show-error \
https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
--data-binary @- <<'JSON'
{
"model": "jev-latest",
"state": "My card was charged twice for order A-104. Please refund the extra $48.",
"questions": {
"refund_requested": {
"type": "noul",
"instructions": "Does the customer request a refund?"
},
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery, delays, lost packages",
"billing": "Charges, invoices, payment problems",
"other": "Anything else"
}
},
"urgency": {
"type": "score",
"instructions": "How urgent is this message?",
"criteria": [
"No time pressure; a normal queue is fine",
"Wants a prompt reply but no hard deadline",
"Says they are blocked, losing money, or need this today"
]
}
}
}
JSON
Python equivalent with the official SDK:
pip install typesafe-sdk
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # reads TYPESAFE_API_KEY
response = client.system_one(
state="My card was charged twice for order A-104. Please refund the extra $48.",
questions={
"refund_requested": Noul(
instructions="Does the customer request a refund?"
),
"department": Choice(
instructions="Which team should handle this?",
criteria={
"returns": "Exchanges, wrong or damaged items",
"shipping": "Delivery, delays, lost packages",
"billing": "Charges, invoices, payment problems",
"other": "Anything else",
},
),
"urgency": Score(
instructions="How urgent is this message?",
criteria=[
"No time pressure; a normal queue is fine",
"Wants a prompt reply but no hard deadline",
"Says they are blocked, losing money, or need this today",
],
),
},
)
print(response.answers["refund_requested"].noul)
print(response.answers["department"].choice)
print(response.answers["urgency"].score)
You can also try questions in the TypeSafe Playground before writing code.
7. How to write state
State is not a prompt. Do not write “You are a helpful support agent…” Jev is not role-playing. Hand it a record you would hand a sharp colleague.
- Filter first in code. Strip signatures, unsubscribe blocks, raw HTML, 40 previous tickets the current question does not need.
- Name the parts if you send JSON.
{"latest_message": "...", "plan": "pro", "open_invoices": 2}beats a blob. - Point questions at named parts when the state is large: “In
latest_message, does the sender ask for a call?” - Do not ask Jev to do arithmetic or date math. Extract the pieces (month, day, year as Choice questions if you must), then compute in code. Official docs are explicit that Jev 1.13 reads dates as text, not as ordered quantities.
- Keep secrets out unless the question needs them. API keys, full card numbers, and other people’s data do not belong in state by default.
A usable support state looks like this:
{
"customer": {"plan": "pro", "tenure_months": 14, "open_tickets": 1},
"latest_message": "Charged twice for A-104. Refund the extra $48 today.",
"order": {"id": "A-104", "total": 48, "status": "paid"}
}
That is a record. “Please carefully analyse the following customer communication and…” is a prompt. Jev wants the record.
8. Confidence, thresholds, and review
Jev’s guarantee is on shape, not on truth. Your product contract is a threshold.Lowauto-no / other pathMiddle bandhuman reviewHighauto-actPut the two cutoffs in config, not in a promptFigure 4. Almost every production Jev call is a three-way gate, not a boolean.
For a Noul:
YES = 0.85
NO = 0.15
p = answers["refund_requested"].noul
if p >= YES:
issue_refund() # or enqueue it
elif p <= NO:
route_to_bot()
else:
send_to_review()
For Choice / Score, gate on confidence as well as the winning label. A choice of billing with confidence 0.28 is not a routing decision; it is a queue for a person.
Where to put the line depends on the cost of being wrong:
- Raise the yes-threshold when a false yes is expensive (refunds, destructive tools, paging someone at 3am).
- Lower it when a missed yes is expensive (safety flags, fraud, abuse).
- Start conservative. Log every decision with state hash, model version, raw probabilities, and the action you took. Move the line with data, not vibes.
Vendor calibration claim, in one sentence: a reported 0.8 should happen about 80% of the time as a group rate, not as a promise on one ticket. Measure that on your own labels before you trust it with money.
9. Patterns you will actually use
Ask everything at once
Questions run in parallel. Output is free. If you might need return_reason only when department is returns, still ask it. Ignore the unused answers in code. This is cheaper and faster than a second HTTP call.
Confidence-gated routing
Classify, then only auto-route the high-confidence slice. Everything else goes to a human or to a slower LLM. This is the default production shape.
Composite scores
Do not ask Jev for “priority.” Ask severity, frustration, and report quality as three Scores. Normalize each to 0–1 (score / (len(criteria) - 1)) and mix them with weights you own:
priority = 0.6 * severity_norm + 0.3 * frustration_norm + 0.1 * quality_norm
Weights are product policy. Keep them in config.
Cascade
A cheap model (or a heuristic) proposes. Jev verifies field-by-field with Nouls. Only the failures go to a large reasoning model or a human. This is how people are already using Jev next to GPT-class extractors: Jev is the cheap, typed check, not the writer.
Agent harness
The agent loop stays in code. At each fork — which tool, which element, allow / ask / deny — Jev picks from a menu you built from the current page or tool list. An LLM, if present, only writes arguments that must be prose. Browser agents such as community projects in this style exist because Jev cannot invent a selector that was not offered.
Real-time loops
Vendor-reported latency is 70–500 ms. That is why people put Jev inside request paths and game ticks rather than overnight batch jobs. Still treat it as a network call: timeout, retry on 429, fail closed on destructive actions.
10. When to use Jev (and when not to)
| Good first jobs | Leave it alone |
|---|---|
| Route a support email to a team | Write the reply |
| Filter RAG passages that actually bear on the claim | Generate the answer from those passages |
| Score ticket urgency or bug severity | Compute a date window or a total |
| Decide whether a tool call is destructive | Execute the tool |
| Pick the next UI action from a listed menu | Invent CSS selectors or code |
| Check whether a citation supports a sentence | Author the sentence |
| Flag “customer asked for a human” | Hold a conversation with the customer |
A useful test: if the output must be a string a person reads, you want an LLM. If the output must be a value an if statement reads, you want Jev. If the output must be a number the universe already defined (a sum, a timestamp, a hash), you want code.
11. Models, price, and limits
Figures below are from TypeSafe’s models page around 20 September 2026. Rate limits were described as dynamic while demand is high. Check the live docs before you bank a design on them.
| Item | Published figure |
|---|---|
| Current version | jev-1.13.0 (alias jev-latest) |
| Input price | $0.042 per million tokens ($42 per billion) |
| Output price | Free |
| Context | 64k tokens per request; 32k for state plus the longest question |
| Input modalities | Text only (string, object, or array of text) |
| Published rate limits | 250,000 tokens/sec · 1,200 requests/min |
| Vendor latency | 70–500 ms typical |
| Choice options | Up to 255 |
| Score levels | 2 to 10 |
Speed and cost comparisons you will see in launch posts (on the order of “~200× faster / ~400× cheaper than a frontier LLM”) are vendor workflow numbers on decision tasks, not “Jev writes a novel 200× cheaper.” Treat them as a reason to benchmark your own loop, not as a universal constant.
A 429 means you are over a limit. Official SDKs retry with backoff. If you call HTTP yourself, honour Retry-After.
12. Common beginner mistakes
- Using Jev as a chatbot. There is no reply to show the user. Pair it with an LLM if you need sentences.
- One question per HTTP call. Batch. Parallel questions are the product.
- Choice for a yes/no, or Noul for a scale. Re-read Figure 3.
- Vague criteria. “Billing stuff” loses to “Charges, invoices, payment problems — not delivery status.”
- No
otherbucket. Forced-choice on an incomplete list is how you get confident nonsense. - Acting on the label and ignoring the distribution. 0.61 returns / 0.35 billing is a dual-route or a review, not “returns.”
- Thresholds buried in prompt text. Put 0.85 in config. Version it. A/B it.
- Asking Jev to count, add, or order dates. Extract, then compute.
- Compound Nouls. “Angry and wants a refund” is two questions.
- State as a novel. Filter in code. Name fields. Ask atomic questions.
- Trusting confidence as correctness. Confidence is peakedness of the distribution. Measure accuracy on labeled tickets.
- Pinning nothing. When
jev-latestmoves, your 0.8 may not mean what it meant last month.
13. Glossary
| Term | Meaning here |
|---|---|
| System One model | TypeSafe’s name for models that emit typed decisions instead of prose. |
| State | The input record questions are judged against. |
| Primitive / question type | Choice, Score, or Noul. |
| Noul | Probability that a yes/no claim is true (0–1). |
| Choice | Pick one named option from a list you defined. |
| Score | Probability-weighted position on your rubric. |
| Confidence | How peaked a Choice or Score distribution is (0–1). |
| Criteria | Option descriptions (Choice), level descriptions (Score), or true/false notes (Noul). |
| RLCD | Reinforcement Learning for Calibrated Decisions — TypeSafe’s training slogan. Not a public paper you can reproduce from. |
| Fan-out | Asking many speculative questions in one call, then branching in code. |
| Jev Engineering | Community name for the split: LLM writes, Jev decides, code acts. |
14. Where to go next
- Run the curl in section 6 against a real message from your product. Then change one sentence and watch the numbers move.
- Read the official primitives in order: Noul, Choice, Score.
- Work through learnjev.com/start if you want the two-hour tutorial path.
- Put a three-way gate in front of one real action (route, refund flag, tool allow). Log model version, probabilities, and what you did.
- Only then wire an LLM next to it for the sentences humans read.
This guide is a beginner map, not a substitute for TypeSafe’s docs. APIs, aliases, prices, and rate limits move. Confirm against docs.typesafe.ai before you ship.

Leave a Reply