An assessment of TypeSafe's Jev model, a survey of every place the platform classifies or normalises today, and a shortlist of where Jev should replace an LLM call or back up a rule that gives up. Rules that work stay as they are.
Date 26 September 2026Status Investigation only, no code changedScope 9 services and shared packages, ~60 touchpointsModel assessedjev-1.13.0
Summary
The headline
Jev is a strong fit for PolicyCheck where we already rely on an LLM, and as a backstop where a rule gives up. It should not replace deterministic rules that work. The main gain is not cost: today our classification confidence is a hard-coded number, and Jev would give us a real, calibrated one that switches on pipeline branches already in the code.
0.85
Confidence is fixed, not measured
Document classification always reports 0.85 and never asks for confirmation. Every downstream gate that depends on confidence is currently inert.
0.0%
The heuristics lean on the LLM
On 503 labelled seed documents, the keyword layers alone score 0% on cover notes, both endorsement types and schedules of authority.
4
LLM paths built but never wired
Endorsement fallback, chat routing, a shadow router and a second-opinion arbiter were designed and left unused, most likely on cost and latency grounds. Jev removes that blocker.
$0.08
Cost of a full offline benchmark
Classifying all 503 seed documents costs about eight cents. A pilot carries negligible model spend. The real gating items are data residency and compliance.
Recommendation
Run a one-week offline benchmark on the seed corpus, then two weeks in shadow mode on document classification and endorsement matching. Cut over behind flags only once calibration is proven and TypeSafe has answered five compliance questions. No tenant data leaves the platform until then. Keep Amazon Bedrock as the permanent fallback.
Part 1
What Jev is
Jev is the first "System One" model from TypeSafe AI, released 15 September 2026. It is a decision model, not a language model told to output JSON. You give it text plus a fixed set of questions whose answers you define in advance. It returns every answer in a single parallel pass, each with a calibrated probability distribution.
It cannot write a sentence, and that is the point. It cannot invent a label that is not in your list, and its confidence numbers are trained to mean something: answers given 80% should be right about 80% of the time.
Three question types
Choice
"Which one of these?"
Returns
Chosen option, probability per option, confidence
Limit
Up to 255 options; add "other" when the list may not cover every input
Score
"Which level?"
Returns
Level (can be fractional, e.g. 1.4), probabilities, confidence
Limit
2 to 10 ordered, described levels
Noul
"Is this true?"
Returns
Probability of yes, 0 to 1
Limit
The probability is the signal; no separate confidence
Many questions of any type can share one input in one request. Adding a question adds input tokens but almost no latency, and each answer is independent of the others.
What a request looks like
// one request, three independent answersconst r = await client.systemOne({
state: { filename: "END--BAA-2025-Territory.pdf", text: first16kChars },
questions: {
docType: choice("What kind of Lloyd's market document is this?", {
BAA: "A binding authority agreement between managing agent and coverholder",
BINDER_ENDORSEMENT: "An endorsement that amends a binding authority agreement",
POLICY_ENDORSEMENT: "An endorsement that amends an individual policy",
OTHER: "None of the above",
}),
isReinsurance: noul("The document concerns reinsurance"),
scanQuality: score("How readable is the extracted text?",
["Clean digital text", "Minor OCR noise", "Heavily garbled"]),
},
});
r.answers.docType.choice; // "BINDER_ENDORSEMENT"
r.answers.docType.confidence; // 0.91
r.answers.isReinsurance.noul; // 0.03
TypeScript and Python SDKs exist, plus a Pydantic AI adapter. Input is text only; scanned content must be transcribed first.
Published numbers
Vendor-supplied. We have not reproduced these independently.
Metric
Jev
Frontier LLM (vendor comparison)
End-to-end latency
70 – 500 ms
3 – 329 s
Input price
$0.042 / M tokens
$0.20 – $10 / M tokens
Output price
Free
~5× input
Malformed structured output
0% by construction
0.58% – 45.5%
Context window
64k tokens (32k for input + longest question)
n/a
Rate limits
250k tokens/s, 1,200 req/min
n/a
What it cannot do
Generate anythingNo summaries, free-text extraction or explanations. The pattern is: find candidates with regex or an LLM, then let Jev pick.
Read between the linesNegations are taken literally. A question whose "yes" means a bad outcome underperforms.
Count or calculateArithmetic and counting are unreliable and worsen with list size.
Compare datesPick date values with Choice; do the ordering in code.
Ignore noiseAccuracy drops as irrelevant text is added. Filter first.
Act as a security boundaryUser-controlled input can steer it.
Match English elsewhereOther languages are handled, but not equally well.
Handle huge label setsMore than 255 options needs a two-stage scheme.
Compliance is the gating item. TypeSafe says Jev is not trained on customer data, and zero-retention is available on enterprise accounts, but not by default. Servers were described as West Coast only, with no UK or EU region, data residency statement or SOC 2 / ISO reference found. Jev is not on Amazon Bedrock, which is the only provider PolicyCheck uses today. Until these are resolved, the pilot is limited to seed data and shadow mode.
Part 2
How PolicyCheck classifies today
There is exactly one LLM classifier on the per-document hot path: document type plus product type, via detectMetadata in llm-service. Everything else is deterministic (regex, keyword tables, alias maps, database lookups) or an LLM call that only runs when a cheaper step fails. All LLM traffic runs through llm-service and LiteLLM to Amazon Bedrock, with nova-micro as the default classification model.
Where classification sits in ingestion
Slice sync and exact-checksum duplicate check
Text extraction (Docling, MinerU, pdf-parse, OCR) alongside PII and quality pre-checks
Classification: LLM document and product type, then a stack of deterministic overrides main opportunity
Template verification: LMA template classifier and section tagging
Near-duplicate check, language detection
Metadata: insurer, product and coverholder resolution opportunity
Extraction: LLM tiers, clause segmentation, term tagging
Hydration: class, currency, territory, occupancy normalisation; identity and date fallbacks
Endorsement overlay: operation classification, clause match opportunity
Asynchronous enrichment: number tagging, document agents, signature tags
Six findings from the survey
Document classification confidence is not a measurement.It always returns 0.85 with no confirmation required. The early-classification threshold (0.7), the endorsement association gate (≥ 0.9) and the low-confidence tag therefore never change behaviour.
One LLM call answers ten questions as free text, then five keyword layers second-guess it.Our own baseline on 503 labelled documents shows that without the LLM's answer, cover notes, both endorsement types and schedules of authority score 0%. The heuristics only work when the LLM has already done the job.
Four LLM classification paths were designed and never connected.Product-type mapping, the endorsement fallback, chat agent routing and the chat shadow router. Each was most likely left off because an LLM call per document or per message was too slow or expensive. Jev's economics are the missing precondition.
Insurer resolution creates duplicate records.When a fuzzy match fails, a new insurer is auto-created as pending review, so near-duplicates accumulate. A "same entity?" check over the top candidates is the cheapest fix.
Vocabularies are duplicated.Class of business exists in three places and document type in five. We should settle one canonical list per concept before adding Jev, or it becomes a sixth.
Inbound email classification is probably not reaching the LLM.The call targets the wrong route, omits a required field and reads a response field that does not exist. The filename fallback is doing all the work. Not verified at runtime, but three independent mismatches sit in the same call path.
Part 3
Ranked opportunities
A deterministic rule gives the same answer every time, costs nothing and can be explained in an audit. On a Lloyd's-market platform that is worth keeping, so Jev is not a replacement for rules that work. We applied four rules, in order, to every touchpoint:
Replace
Replace an existing LLM callThe decision is already model-made. Jev makes it calibrated, sub-second and unable to return an off-list label.
Fill slot
Fill a slot designed for an LLM but never wiredFour such paths exist in the code today.
Leftovers
Rule first, Jev only on what the rule gives up onA 0.3 passthrough, a default of "other" or "unknown", missing data, an auto-created record. The rule still runs first and still wins when it fires.
Keep
Otherwise keep the ruleA heuristic that works is not replaced. Where we have no failure data, the item is marked Measure until we do.
Audit safeguard. For authority and compliance outcomes that are rule-based today, Jev may raise a suggestion for a reviewer but never sets a pass or fail itself.
R Replace LLM callS Fill unwired slotT Rule first, Jev on leftoversM Measure firstK Keep deterministic
Replace the LLM's free-text "document kind" question with one Jev request over the filename, first 16k characters, upload context and frontmatter:
Document type: Choice over the ~50 document types plus OTHER, each described from our Lloyd's domain notes
Product type: Choice over 20 products plus OTHER
Supporting document only: Noul, which lets us skip a 15–47 s full parse
Endorsement side: binder or policy, asked only when the existing weighted scoring is too close to call
The filename, frontmatter, reinsurance and body-text rules all stay and still override when they fire. What goes are the two keyword chains that exist only to translate the LLM's free text into our categories.
A real probability replaces the fixed 0.85, so the three inert gates start working and low-confidence documents can be flagged for confirmation. Early classification becomes cheap enough to leave on. About $0.00017 per document.
Which BAA section is amended: Choice over the heading list, replacing the LLM matching call. Its 0.5 filter becomes a real threshold.
Per-clause change details: moves from the LLM diff call to Jev.
Category and impact: asked only when the keyword table finds no match, instead of today's "administrative, low impact, 30% confidence". This is the fallback that was designed and never wired.
The operation regex (add, remove, amend…) stays as it is.
J6
Inbound email attachment triage Replace
Replaces E1 · keeps E2
Replace the misrouted LLM call with a direct Jev request over filename, first 8k characters and subject: document type (11 options), issuance / change / supporting, and whether it is in scope.
Fixes a path that is probably broken today. The filename table stays as an input and as the fallback.
J3
Insurer resolution misses Leftovers
I25 · PL1
The fuzzy search and the automatic link at a score of 85 or above are unchanged. Only below 85, before a new insurer would be auto-created, Jev checks whether any of the top 5 candidates is the same legal entity. A new record is created only when "none" wins with high confidence.
Directly reduces the growth of duplicate pending-review insurer records.
Tier 2
Other LLM calls to replace
J8Clause typeThe LLM still finds clause boundaries; Jev labels each clause from 12 types.
J13Products from client needsProbabilities give a ranked top 5 in one call, instead of asking an LLM to stick to a list.
J14Compliance yes/no verdictsAlready LLM-made. Jev gives the verdict; the LLM only writes the rationale when a check fails or warns.
I29Identity candidatesWhen the regex cascade finds more than one UMR or coverholder, Jev picks instead of the LLM fallback.
Rule first, Jev only on the leftovers
J7Coverage statusOnly when the regex returns "unknown".
J10Class of businessOnly on the low-confidence passthroughs the alias table cannot place.
J11Occupancy and tradeOnly for values currently kept as free text.
J12Bordereau subtypeOnly when the heuristic returns "unclassified".
J15Authority rule topicOnly when the map falls through to "other".
J16Authority check missing dataA reviewer suggestion with a stored probability, never the verdict itself.
Measure first
J5 Chat routing. The keyword router may be working fine. Run Jev through the existing shadow harness and log agreement; replace only if disagreements show the router is wrong.
J4 Monetary figure subtypes. Needs an error rate for the keyword window first.
J9 LMA clause match cut-offs. Needs evidence the similarity thresholds misjudge.
Term tagging, party names, finding severity.
Keep deterministic
Filename, frontmatter, reinsurance and body-text overrides.
Enum aliasing, currency codes, value cleaning and rule-code routing. A lookup is faster and exact.
Language detection, scan quality and PII detection.
Chat guardrails. The vendor states Jev is not a security boundary.
Where it does not belong
Anything that writes text: extracted values, summaries, rationales, chat answers.
Anything numeric: premium sums, date ordering, counting clauses.
Safety, tenant isolation, authorisation, and final verdicts on rule-based authority checks.
Non-English documents, until accuracy on them is measured.
Part 2 · detail
Full touchpoint inventory
Every place the platform classifies or normalises, with its current approach and our recommendation under the four rules. "Hot" runs per document, page, email or chat turn. File references were checked against the codebase on 26 September 2026.
Ref
Service
Task
Today
Rec.
Location
Part 4
Pilot plan
Before any tenant document is sent
Question for TypeSafe
Why it matters
1
UK or EU inference region, or contractual data residency
Lloyd's market data; current statement is West Coast servers
2
Zero data retention on our account, in writing
"Not trained on" is not the same as "not retained"
3
SOC 2 / ISO 27001 report, DPA, sub-processor list
Tenant isolation and auditability are our top two platform priorities
4
Availability via Bedrock or a gateway we can route through LiteLLM
We run one gateway today; a second provider path is a platform change
5
Model pinning and deprecation policy for jev-1.13.0
Thresholds are tuned per version; silent upgrades break calibration
Until questions 1–3 are answered, the pilot uses only the Lloyd's seed corpus. No tenant data.
Phases
Phase 01 week
Offline benchmark, no production change
Add a Jev runner to the existing baseline harness (503 labelled seed documents) for J1.
Report strict and family-level accuracy per document class, plus a calibration curve (stated confidence against observed accuracy).
Pass bar: match or beat the current system on every focus class, and beat the heuristics-only result by a wide margin on the four classes at 0%.
Model cost of the run: about $0.08.
Phase 12 weeks
Shadow mode
Add Jev as a provider in llm-service, selectable per operation from the existing admin runtime settings.
Run J1 and J2 in shadow: log Jev's answer, confidence and latency alongside production.
Run J5 as a measurement only, beside the keyword chat router. No switch without disagreement data.
Measure agreement, disagreement categories, and p50 / p95 latency from eu-west-2.
Phase 2after sign-off
Cut over J1, J2 and J6 behind flags
Confidence-gated, with thresholds tuned from the Phase 0 calibration curve:
≥ 0.85Auto-accept0.60 – 0.85Accept, ask for confirmation< 0.60Fall back to current path
Retire the two free-text translation chains once agreement is stable. The filename, frontmatter, reinsurance and body-text rules stay.
Phase 3as capacity allows
Remaining opportunities
J3 and the rest of Tier 2. "Measure first" items only once we have failure data.
Part 4 · continued
Risks and economics
Vendor maturityThe model was eleven days old at the time of writing and all performance figures are vendor-run. Bedrock stays as a permanent fallback; Jev is never the only route to a classification.
Calibration is per versionPin jev-1.13.0, record the model version on every response in the audit trail, and re-run the benchmark on each upgrade.
Literal readingOption descriptions must be written carefully. Our Lloyd's domain definitions should be reused verbatim.
Noisy inputSend the first 16k characters and filename, not the whole document.
A sixth vocabularyConsolidate duplicated document-type and class lists first, or new probabilities land on labels other services alias differently.
Indicative model cost
Assumes ~4k tokens per document, 1.5k per chat turn and 2k per endorsement, at prices published 26 September 2026.
Workload
Per unit
10,000 / month
100,000 / month
J1 Document classification
$0.00017
$1.68
$16.80
J5 Chat routing (shadow)
$0.00006
$0.63
$6.30
J2 Endorsements
$0.00008
$0.84
$8.40
The current nova-micro classification is already inexpensive ($0.15 per million input tokens, $0.60 output), so direct model savings are modest. The value is in latency: sub-second answers let early classification stay on and skip full parses for supporting documents. And it is in real confidence, which activates the confidence-gated pipeline branches that exist in code but never fire.