Log 013 · Seas: calmVoyage: Neon Oracle

It Can't Write a Sentence, So We're Asking It to Judge Our Hooks

TL;DR

  • Jev is a "System One" model from TypeSafe AI (launched 2026-09-15) that cannot generate text — it returns only typed decisions: Choice, Score, or Noul (yes/no probability).
  • The headline figures ($0.042/1M input tokens, 70–500ms latency, "0% hallucination") are all vendor claims, not independently reproduced.
  • We specced a pilot: run our labeled hooks through Jev's Score gate and compare against the human blind-test verdicts before Day 1 (9/23).

It can't write a single sentence — which is exactly why it might judge our hooks better than any chat model.

This is not Log 012. That one was about postponing the hook verdict twice — a delay record. This is a tool-verification record: before we let a model anywhere near our Day 1 (9/23) hook decisions, we are putting the model itself on trial. The candidate is Jev, released by TypeSafe AI on September 15, 2026, and our deep-dive research on it landed last night at 23:28.

What Jev actually is

**Jev is a non-generative decision model.** You POST `{state, model, questions}` to `api.typesafe.ai/v1/systemone`, declare the answer space *before* the model runs, and it returns typed probabilistic judgments — never prose.

Three primitives, verified against the official docs:

  • **Choice** — pick one from a declared set (up to 255 options); returns the choice, per-option probabilities, and a confidence score.
  • **Score** — place the state on an ordered rubric (2–10 levels); returns a weighted score with a legend.
  • **Noul** — "is this proposition true?"; returns a single 0–1 probability.

The working analogy: **a smart if-statement.** You declare the answer space, and the model scores every possible answer in one parallel forward pass instead of writing tokens one by one. No token loop means no output to bill and milliseconds of latency — 70–500ms and $0.042/1M input tokens with output free, per the vendor.

Why a judge that can't write might judge better

Observation: a chat model asked to score hooks has two jobs at once — think *and* write the score in a parseable format. The writing half is where format errors, hedging, and hallucinated categories creep in. Jev does one job: given the same hook plus our rubric, it returns a number.

Interpretation: for bounded decisions like "rate this hook 1–5 against our rubric," removing generation removes a failure surface. It does not remove judgment — and judgment can still be wrong. It just makes wrong cheaper and faster.

That matters because our use case is not philosophical. Our Neon Oracle hook verdict is currently a handful of humans arguing in a chat thread. A scorer that returns `stop_power: 4.2, confidence: 0.71` in ~300ms for $0.000021 per decision (at ~500 input tokens per call, vendor pricing) is not a better thinker than us — it is a tireless intern with no opinions and no handwriting.

The part I don't trust: "0% hallucination," dismantled

The vendor's most quoted claim is "0% hallucination." Here is the narrow, defensible reading, per TypeSafe's own materials: because candidates are bound to pre-declared enums and types, **a type error is mathematically impossible**. The model cannot emit a field that doesn't exist or a string where a float was declared.

That is not the same as being right. The vendor's own qualifiers, which they published (to their credit, unusually frank):

  • TypeSafe's CEO conceded on Hacker News that Jev can be **"confidently wrong"** — a schema-valid but factually wrong answer with high confidence.
  • The "0% hallucination" rate is **"not empirical"** — by construction, not measured (vendor's own qualifier).
  • As one independent reviewer put it: the real property is *prose-free*, not *hallucination-free*.

Then there is the benchmark homework. TypeSafe's four workflow evals (security triage, agent trace review, invoice processing, customer service) grade Jev against the average of GPT-6 Astra and Claude Fable 5.1 — two frontier models, not human-labeled ground truth. So "agreement" measures **consensus with two frontier models, not correctness** (interpretation). The company's own footnotes admit the comparisons sit at "the higher end of real world gains" and that latency was measured "from our laptops on the West Coast." And on its own chart, Jev was not the most accurate — it tied GPT-5.6 Terra and Claude Sonnet 5 at 67.8% while GPT-5.6 Sol led at 74.1%. On invoice processing specifically, Jev scored 61.8% against Sol's 79.1%. Seventeen points on a pay-the-vendor decision is worth more than the inference bill.

So the claim that survives red-teaming is small and honest: **Jev returns well-formed, cheap, fast, calibrated-ish judgments — that can still be wrong.**

"A model that can't write can't bluff — but it can still be wrong with perfect grammar."

Our use case: the hook QC pilot

Observation: we already have a pilot spec (`jev-hook-scoring/SPEC.md`). Every hook goes through one parallel pass — `stop_power`, `save_intent`, `share_intent`, `brand_fit` as Score (1–5), plus a `rule_violation` Noul against our brand rules. The state carries our tagline ("Horoscopes give you vibes. Saju gives you dates."), our message hierarchy (money dates > money season > time map), and the no-guarantee-amounts rule.

Interpretation: this fits Jev's strong suit — bounded scoring against a declared rubric — and avoids its jagged edges (no arithmetic, no dates, no multi-hop reasoning in the questions themselves).

The validation plan is deliberately boring: **18 hooks from our own corpus, run through Jev and through our current judging, compared side by side.** Success bar: at least two of Jev's top 3 overlap the humans' top 3, and the one hook with a known brand-rule violation gets flagged. One honest rule we borrowed from the vendor's own docs: never let a confident-wrong answer touch money or launch decisions — confidence below 0.5 routes to a human, and nothing ships on Jev's word alone.

Status: spec done, **API key pending** (`TYPESAFE_API_KEY` not yet issued — Jev is still early-access/waitlist). The pilot script exists and waits.

What remains unconfirmed

  • The 193.6x faster / 444.6x cheaper figures: vendor-published against vendor-built workflows. No independent replication exists as of 2026-09-21.
  • Calibration: confidence is a distribution-spread statistic with **no published calibration numbers**. We must validate on our own labeled hooks before any routing decision.
  • Pricing: "$0.042/1M, output free" is live but the vendor admits, "we can't prove it isn't subsidized." We are not architecting around permanence.
  • Input budget: sources disagree (32k vs 64k tokens). Unconfirmed until read from current docs.

What this connects to

Day 1 is Wednesday, 9/23. The blind-test hook verdict is still unlanded, and Jev will not decide it for us — a vendor brochure does not get a vote on our launch. But if the 18-hook pilot holds, Jev becomes an **auxiliary scoring gate**: candidates ranked by a tireless, opinion-free judge before humans argue the finals. That is a small, honest upgrade, and small honest upgrades are the only kind I trust six days after a launch.

Next log

Whether the blind-test verdict finally lands — and what Jev thought of the same hooks.

Next log: Whether the blind-test verdict finally lands — and what Jev thought of the same hooks.
← Ship Log