Codeaza Technologies Prepared for PolicyCheck  ·  Commercial in confidence
Jev in PolicyCheck
Technical investigation · Classification & normalisation

Where a calibrated classifier fits in PolicyCheck

An assessment of TypeSafe's Jev model, a survey of every place the platform classifies or normalises today, and a shortlist of where Jev should replace an LLM call or back up a rule that gives up. Rules that work stay as they are.

Date 26 September 2026 Status Investigation only, no code changed Scope 9 services and shared packages, ~60 touchpoints Model assessed jev-1.13.0
Summary

The headline

Jev is a strong fit for PolicyCheck where we already rely on an LLM, and as a backstop where a rule gives up. It should not replace deterministic rules that work. The main gain is not cost: today our classification confidence is a hard-coded number, and Jev would give us a real, calibrated one that switches on pipeline branches already in the code.

0.85

Confidence is fixed, not measured

Document classification always reports 0.85 and never asks for confirmation. Every downstream gate that depends on confidence is currently inert.

0.0%

The heuristics lean on the LLM

On 503 labelled seed documents, the keyword layers alone score 0% on cover notes, both endorsement types and schedules of authority.

4

LLM paths built but never wired

Endorsement fallback, chat routing, a shadow router and a second-opinion arbiter were designed and left unused, most likely on cost and latency grounds. Jev removes that blocker.

$0.08

Cost of a full offline benchmark

Classifying all 503 seed documents costs about eight cents. A pilot carries negligible model spend. The real gating items are data residency and compliance.

Recommendation

Run a one-week offline benchmark on the seed corpus, then two weeks in shadow mode on document classification and endorsement matching. Cut over behind flags only once calibration is proven and TypeSafe has answered five compliance questions. No tenant data leaves the platform until then. Keep Amazon Bedrock as the permanent fallback.

Part 1

What Jev is

Jev is the first "System One" model from TypeSafe AI, released 15 September 2026. It is a decision model, not a language model told to output JSON. You give it text plus a fixed set of questions whose answers you define in advance. It returns every answer in a single parallel pass, each with a calibrated probability distribution.

It cannot write a sentence, and that is the point. It cannot invent a label that is not in your list, and its confidence numbers are trained to mean something: answers given 80% should be right about 80% of the time.

Three question types

Choice

"Which one of these?"

Returns
Chosen option, probability per option, confidence
Limit
Up to 255 options; add "other" when the list may not cover every input

Score

"Which level?"

Returns
Level (can be fractional, e.g. 1.4), probabilities, confidence
Limit
2 to 10 ordered, described levels

Noul

"Is this true?"

Returns
Probability of yes, 0 to 1
Limit
The probability is the signal; no separate confidence

Many questions of any type can share one input in one request. Adding a question adds input tokens but almost no latency, and each answer is independent of the others.

What a request looks like

// one request, three independent answers
const r = await client.systemOne({
  state: { filename: "END--BAA-2025-Territory.pdf", text: first16kChars },
  questions: {
    docType: choice("What kind of Lloyd's market document is this?", {
      BAA: "A binding authority agreement between managing agent and coverholder",
      BINDER_ENDORSEMENT: "An endorsement that amends a binding authority agreement",
      POLICY_ENDORSEMENT: "An endorsement that amends an individual policy",
      OTHER: "None of the above",
    }),
    isReinsurance: noul("The document concerns reinsurance"),
    scanQuality: score("How readable is the extracted text?",
      ["Clean digital text", "Minor OCR noise", "Heavily garbled"]),
  },
});

r.answers.docType.choice;      // "BINDER_ENDORSEMENT"
r.answers.docType.confidence;  // 0.91
r.answers.isReinsurance.noul;  // 0.03

TypeScript and Python SDKs exist, plus a Pydantic AI adapter. Input is text only; scanned content must be transcribed first.

Published numbers

Vendor-supplied. We have not reproduced these independently.

MetricJevFrontier LLM (vendor comparison)
End-to-end latency70 – 500 ms3 – 329 s
Input price$0.042 / M tokens$0.20 – $10 / M tokens
Output priceFree~5× input
Malformed structured output0% by construction0.58% – 45.5%
Context window64k tokens (32k for input + longest question)n/a
Rate limits250k tokens/s, 1,200 req/minn/a

What it cannot do

  • Generate anythingNo summaries, free-text extraction or explanations. The pattern is: find candidates with regex or an LLM, then let Jev pick.
  • Read between the linesNegations are taken literally. A question whose "yes" means a bad outcome underperforms.
  • Count or calculateArithmetic and counting are unreliable and worsen with list size.
  • Compare datesPick date values with Choice; do the ordering in code.
  • Ignore noiseAccuracy drops as irrelevant text is added. Filter first.
  • Act as a security boundaryUser-controlled input can steer it.
  • Match English elsewhereOther languages are handled, but not equally well.
  • Handle huge label setsMore than 255 options needs a two-stage scheme.
Compliance is the gating item. TypeSafe says Jev is not trained on customer data, and zero-retention is available on enterprise accounts, but not by default. Servers were described as West Coast only, with no UK or EU region, data residency statement or SOC 2 / ISO reference found. Jev is not on Amazon Bedrock, which is the only provider PolicyCheck uses today. Until these are resolved, the pilot is limited to seed data and shadow mode.
Part 2

How PolicyCheck classifies today

There is exactly one LLM classifier on the per-document hot path: document type plus product type, via detectMetadata in llm-service. Everything else is deterministic (regex, keyword tables, alias maps, database lookups) or an LLM call that only runs when a cheaper step fails. All LLM traffic runs through llm-service and LiteLLM to Amazon Bedrock, with nova-micro as the default classification model.

Where classification sits in ingestion

  1. Slice sync and exact-checksum duplicate check
  2. Text extraction (Docling, MinerU, pdf-parse, OCR) alongside PII and quality pre-checks
  3. Classification: LLM document and product type, then a stack of deterministic overrides main opportunity
  4. Layout normalisation, canonical markdown, parse-completeness gate
  5. Template verification: LMA template classifier and section tagging
  6. Near-duplicate check, language detection
  7. Metadata: insurer, product and coverholder resolution opportunity
  8. Extraction: LLM tiers, clause segmentation, term tagging
  9. Hydration: class, currency, territory, occupancy normalisation; identity and date fallbacks
  10. Endorsement overlay: operation classification, clause match opportunity
  11. Asynchronous enrichment: number tagging, document agents, signature tags

Six findings from the survey

  • Document classification confidence is not a measurement.It always returns 0.85 with no confirmation required. The early-classification threshold (0.7), the endorsement association gate (≥ 0.9) and the low-confidence tag therefore never change behaviour.
  • One LLM call answers ten questions as free text, then five keyword layers second-guess it.Our own baseline on 503 labelled documents shows that without the LLM's answer, cover notes, both endorsement types and schedules of authority score 0%. The heuristics only work when the LLM has already done the job.
  • Four LLM classification paths were designed and never connected.Product-type mapping, the endorsement fallback, chat agent routing and the chat shadow router. Each was most likely left off because an LLM call per document or per message was too slow or expensive. Jev's economics are the missing precondition.
  • Insurer resolution creates duplicate records.When a fuzzy match fails, a new insurer is auto-created as pending review, so near-duplicates accumulate. A "same entity?" check over the top candidates is the cheapest fix.
  • Vocabularies are duplicated.Class of business exists in three places and document type in five. We should settle one canonical list per concept before adding Jev, or it becomes a sixth.
  • Inbound email classification is probably not reaching the LLM.The call targets the wrong route, omits a required field and reads a response field that does not exist. The filename fallback is doing all the work. Not verified at runtime, but three independent mismatches sit in the same call path.
Part 3

Ranked opportunities

A deterministic rule gives the same answer every time, costs nothing and can be explained in an audit. On a Lloyd's-market platform that is worth keeping, so Jev is not a replacement for rules that work. We applied four rules, in order, to every touchpoint:

  1. Replace
    Replace an existing LLM callThe decision is already model-made. Jev makes it calibrated, sub-second and unable to return an off-list label.
  2. Fill slot
    Fill a slot designed for an LLM but never wiredFour such paths exist in the code today.
  3. Leftovers
    Rule first, Jev only on what the rule gives up onA 0.3 passthrough, a default of "other" or "unknown", missing data, an auto-created record. The rule still runs first and still wins when it fires.
  4. Keep
    Otherwise keep the ruleA heuristic that works is not replaced. Where we have no failure data, the item is marked Measure until we do.
Audit safeguard. For authority and compliance outcomes that are rule-based today, Jev may raise a suggestion for a reviewer but never sets a pass or fail itself.
R Replace LLM call S Fill unwired slot T Rule first, Jev on leftovers M Measure first K Keep deterministic

Tier 1: do these first

J1

Document classification Replace

Replaces L1 · L5 · retires L2 · L4 · keeps L3 · I2 · I3

Replace the LLM's free-text "document kind" question with one Jev request over the filename, first 16k characters, upload context and frontmatter:

  • Document type: Choice over the ~50 document types plus OTHER, each described from our Lloyd's domain notes
  • Product type: Choice over 20 products plus OTHER
  • Supporting document only: Noul, which lets us skip a 15–47 s full parse
  • Endorsement side: binder or policy, asked only when the existing weighted scoring is too close to call

The filename, frontmatter, reinsurance and body-text rules all stay and still override when they fire. What goes are the two keyword chains that exist only to translate the LLM's free text into our categories.

A real probability replaces the fixed 0.85, so the three inert gates start working and low-confidence documents can be flagged for confirmation. Early classification becomes cheap enough to leave on. About $0.00017 per document.

J2

Endorsement matching Replace Fill slot

Replaces I22 · I23 · fills the I20 fallback · keeps I21
  • Which BAA section is amended: Choice over the heading list, replacing the LLM matching call. Its 0.5 filter becomes a real threshold.
  • Per-clause change details: moves from the LLM diff call to Jev.
  • Category and impact: asked only when the keyword table finds no match, instead of today's "administrative, low impact, 30% confidence". This is the fallback that was designed and never wired.

The operation regex (add, remove, amend…) stays as it is.

J6

Inbound email attachment triage Replace

Replaces E1 · keeps E2

Replace the misrouted LLM call with a direct Jev request over filename, first 8k characters and subject: document type (11 options), issuance / change / supporting, and whether it is in scope.

Fixes a path that is probably broken today. The filename table stays as an input and as the fallback.

J3

Insurer resolution misses Leftovers

I25 · PL1

The fuzzy search and the automatic link at a score of 85 or above are unchanged. Only below 85, before a new insurer would be auto-created, Jev checks whether any of the top 5 candidates is the same legal entity. A new record is created only when "none" wins with high confidence.

Directly reduces the growth of duplicate pending-review insurer records.

Tier 2

Other LLM calls to replace

J8Clause typeThe LLM still finds clause boundaries; Jev labels each clause from 12 types.
J13Products from client needsProbabilities give a ranked top 5 in one call, instead of asking an LLM to stick to a list.
J14Compliance yes/no verdictsAlready LLM-made. Jev gives the verdict; the LLM only writes the rationale when a check fails or warns.
I29Identity candidatesWhen the regex cascade finds more than one UMR or coverholder, Jev picks instead of the LLM fallback.

Rule first, Jev only on the leftovers

J7Coverage statusOnly when the regex returns "unknown".
J10Class of businessOnly on the low-confidence passthroughs the alias table cannot place.
J11Occupancy and tradeOnly for values currently kept as free text.
J12Bordereau subtypeOnly when the heuristic returns "unclassified".
J15Authority rule topicOnly when the map falls through to "other".
J16Authority check missing dataA reviewer suggestion with a stored probability, never the verdict itself.

Measure first

  • J5 Chat routing. The keyword router may be working fine. Run Jev through the existing shadow harness and log agreement; replace only if disagreements show the router is wrong.
  • J4 Monetary figure subtypes. Needs an error rate for the keyword window first.
  • J9 LMA clause match cut-offs. Needs evidence the similarity thresholds misjudge.
  • Term tagging, party names, finding severity.

Keep deterministic

  • Filename, frontmatter, reinsurance and body-text overrides.
  • Endorsement operation, transaction type, participant role, authority reference topology.
  • Enum aliasing, currency codes, value cleaning and rule-code routing. A lookup is faster and exact.
  • Language detection, scan quality and PII detection.
  • Chat guardrails. The vendor states Jev is not a security boundary.

Where it does not belong

  • Anything that writes text: extracted values, summaries, rationales, chat answers.
  • Anything numeric: premium sums, date ordering, counting clauses.
  • Safety, tenant isolation, authorisation, and final verdicts on rule-based authority checks.
  • Non-English documents, until accuracy on them is measured.
Part 2 · detail

Full touchpoint inventory

Every place the platform classifies or normalises, with its current approach and our recommendation under the four rules. "Hot" runs per document, page, email or chat turn. File references were checked against the codebase on 26 September 2026.

RefServiceTaskTodayRec.Location
Part 4

Pilot plan

Before any tenant document is sent

Question for TypeSafeWhy it matters
1UK or EU inference region, or contractual data residencyLloyd's market data; current statement is West Coast servers
2Zero data retention on our account, in writing"Not trained on" is not the same as "not retained"
3SOC 2 / ISO 27001 report, DPA, sub-processor listTenant isolation and auditability are our top two platform priorities
4Availability via Bedrock or a gateway we can route through LiteLLMWe run one gateway today; a second provider path is a platform change
5Model pinning and deprecation policy for jev-1.13.0Thresholds are tuned per version; silent upgrades break calibration

Until questions 1–3 are answered, the pilot uses only the Lloyd's seed corpus. No tenant data.

Phases

  1. Phase 01 week

    Offline benchmark, no production change

    • Add a Jev runner to the existing baseline harness (503 labelled seed documents) for J1.
    • Report strict and family-level accuracy per document class, plus a calibration curve (stated confidence against observed accuracy).
    • Pass bar: match or beat the current system on every focus class, and beat the heuristics-only result by a wide margin on the four classes at 0%.
    • Model cost of the run: about $0.08.
  2. Phase 12 weeks

    Shadow mode

    • Add Jev as a provider in llm-service, selectable per operation from the existing admin runtime settings.
    • Run J1 and J2 in shadow: log Jev's answer, confidence and latency alongside production.
    • Run J5 as a measurement only, beside the keyword chat router. No switch without disagreement data.
    • Measure agreement, disagreement categories, and p50 / p95 latency from eu-west-2.
  3. Phase 2after sign-off

    Cut over J1, J2 and J6 behind flags

    Confidence-gated, with thresholds tuned from the Phase 0 calibration curve:

    ≥ 0.85Auto-accept 0.60 – 0.85Accept, ask for confirmation < 0.60Fall back to current path
    • Retire the two free-text translation chains once agreement is stable. The filename, frontmatter, reinsurance and body-text rules stay.
  4. Phase 3as capacity allows

    Remaining opportunities

    • J3 and the rest of Tier 2. "Measure first" items only once we have failure data.
Part 4 · continued

Risks and economics

Vendor maturityThe model was eleven days old at the time of writing and all performance figures are vendor-run. Bedrock stays as a permanent fallback; Jev is never the only route to a classification.
Calibration is per versionPin jev-1.13.0, record the model version on every response in the audit trail, and re-run the benchmark on each upgrade.
Literal readingOption descriptions must be written carefully. Our Lloyd's domain definitions should be reused verbatim.
Noisy inputSend the first 16k characters and filename, not the whole document.
A sixth vocabularyConsolidate duplicated document-type and class lists first, or new probabilities land on labels other services alias differently.

Indicative model cost

Assumes ~4k tokens per document, 1.5k per chat turn and 2k per endorsement, at prices published 26 September 2026.

WorkloadPer unit10,000 / month100,000 / month
J1 Document classification$0.00017$1.68$16.80
J5 Chat routing (shadow)$0.00006$0.63$6.30
J2 Endorsements$0.00008$0.84$8.40

The current nova-micro classification is already inexpensive ($0.15 per million input tokens, $0.60 output), so direct model savings are modest. The value is in latency: sub-second answers let early classification stay on and skip full parses for supporting documents. And it is in real confidence, which activates the confidence-gated pipeline branches that exist in code but never fire.

Sources