TypeSafe AI opened early access to Jev, a model that returns decisions instead of text
Jev does not write sentences. It reads unstructured input and returns a value whose shape you declared in advance, with a calibrated probability attached to every field. TypeSafe AI charges $0.042 per million input tokens, charges nothing for output, and reports end-to-end response times of 70 to 500 milliseconds. The speed and cost multiples on its home page come from evaluations the company built and ran itself.
TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, opened early access to a model called Jev on September 15, 2026. The company calls it a System One Model, a class built for the decisions software makes inside itself rather than for conversation with a person. Jev never produces a string. It returns values whose type and range were fixed before the call, and it attaches a confidence score to each one.
What the model returns
A language model answers in text, so the program around it has to parse that text and validate it before acting. Jev removes that step. The set of possible answers and their structure are declared in the query itself, the model fills them in, and every answer arrives with a calibrated probability beside it.
Because the shape is fixed in advance, the model cannot return a field that does not exist or a value outside the declared set. TypeSafe presents that as the practical meaning of its claim that Jev cannot hallucinate, and the claim is narrower than it sounds: the answer is always well formed, which is not the same as always correct. Sampling runs in parallel rather than token by token, and that is where the company locates its speed advantage.
Price and response time
Input costs $0.042 per million tokens, which the company also writes as $42 per billion. Output tokens are not billed. For comparison, the same post puts the input price of current language models between $0.20 and $10 per million tokens, with output typically about five times the input price.
TypeSafe reports end-to-end response times of 70 to 500 milliseconds, measured on its own machines on the West Coast of the United States. It says the same calls take frontier language models 3 to 329 seconds.
Where the numbers come from
The multiples on the home page, 193.6 times faster and 444.6 times cheaper, are drawn from four evaluation workflows the company published alongside them. Every model receives the same workflow, and the reference answer is the average of what GPT-6 Astra and Fable 5.1 return.
The company states the caveats itself, which is uncommon enough in a vendor post to be worth naming. It says its own model capabilities team built the workflows. It says averaging two competing models biases the reference toward those two. It says the published multiples sit at the high end of what it expects in production. The hallucination rates it plots for language models come from OpenRouter traffic, while its own zero is derived from the schema guarantee rather than measured.
What it is not for
Jev answers questions whose possible answers are known before the call. It classifies, routes, scores, extracts, and branches, and the number of choices in a single decision is capped at 255. Above that the call splits in two: the options are scored first, then one is chosen. Anything that has to be written, whether a reply, a document, or code, still needs a language model.
The offer, then, is a replacement for the small classification calls buried inside an application, not for the assistant a user talks to. Whether the published multiples hold on someone else's workload is not answerable yet, because no independent measurement has been published.
„While Jev gives up string generation, it's optimized for structured outputs and can't hallucinate.“
Sources
Related

OpenAI will watermark ChatGPT and Codex text in the EU, and API customers can opt in
OpenAI published its plan for text watermarking on October 5, 2026, in response to the EU AI Act. Over the coming weeks, eligible ChatGPT and Codex text output in the European Union will carry an invisible watermark called textGrain. In the API, watermarking is available worldwide from the same day for select models, and it stays off unless you turn it on. The detector is not public: OpenAI is limiting it to approved researchers and expert organizations.
OpenAIverified
A preprint finds six commercial LLM routers no better than a random pick between two models
A preprint submitted to arXiv on October 2, 2026, tested six commercial LLM routers in 14 settings. According to the authors, none of them beat a router that picks at random between Gemini 3.7 Flash and Opus 5 at the same cost, and one trailed it by 10.5 percentage points. The authors, who work at Fastino Labs, trace the gap to the way routers are evaluated and to rosters that hold too many models. The paper has not been peer reviewed.
arXiv, Fastino Labsverified

Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row
Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.
Microsoft, Hugging Faceverified
