TypeSafe's public comparison adapter, typesafe-ai/system-one-adapter-python, asks competing models to supply probabilities, not merely pick an answer. That difference is becoming a practical purchasing question. Unsloth has added a local implementation of the decision interface, while a published independent experiment shows that tightening the required accuracy reverses which of two model-routing setups sends fewer requests to an expensive model.
The new choices change how to read a speed chart. A buyer needs to define both the information an answer must contain and the standard a completed task must meet. Matching the interface and matching the outcome are related experiments, but they are not interchangeable.
TypeSafe's September 15 launch announcement claims large speed and cost advantages for Jev's structured decisions. It also explicitly acknowledges that its LLM comparison wrapper requests probability-bearing answers, which tend to take longer and cost more than discrete decisions. That is a disclosed design choice, not evidence of a hidden handicap.
What the application actually consumes
Our September 22 report examined the probability interface. The adapter's schema generator supports two jobs: return one allowed label, or return a number for every option. These modes predate the latest release. A distribution can help a workflow distinguish uncertainty from a decisive choice; an application that only consumes the winning label may have less to buy.
A fresh example makes the distinction concrete. Pi's September 28 Jev router asks whether programming work is standard or complex, chooses a planning model, and switches to its implementation model after the first successful edit. It reads whether the complex option reaches 0.5. With only two options, that can amount to choosing a label; the example does not prove that an entire probability distribution earns its keep. Nor do its scripted tests establish savings on completed work.
TypeSafe's workflow evaluation asks a richer question: keep the workflow fixed, combine model judgments with code, and compare against reference answers generated by two large models at high reasoning settings. Evaluated models use provider defaults. That is a defined comparison target. It is not the only useful target for a team choosing a classifier or router.
One percentage point can change the cheaper route
An independent September 18 intent-classification study asks what happens when a first model passes uncertain cases to a stronger one. Its 200-item CLINC150 sample compares Jev and a small LLM as the first stage, with GPT-5.6 Terra as the destination for escalated requests. The author publishes the predictions and analysis, including corrections to initially overstated conclusions.
We recalculated the threshold results from the published rows. If the cascade may finish one percentage point below Terra's 91.5% accuracy on that sample, the minimum escalation rates are 22% for Jev and 48.5% for the small LLM. Require exact parity with Terra, and those rates become 100% and 73%. The apparent routing advantage changes direction when the acceptance criterion changes.
The mechanism is a threshold jump. Many Jev answers share the highest confidence value, so the routing rule cannot separate them by confidence. Achieving parity on this sample requires sending every case onward. This is not a general victory for the small LLM: the rates are retrospective minima chosen on the same finite sample, the preregistered finding was ambiguous, and thresholds tested on held-out halves missed the intended accuracy target for both setups.
These are request-routing fractions, not a measured total bill or a reproduction of TypeSafe's workflow charts. Their value is more practical: “almost the same accuracy” and “the same accuracy” can buy very different systems. Set that requirement before shopping for the cheapest first call.
The probability interface now has a local path
The choice is also broader than Jev versus a language model writing probability-shaped JSON. Unsloth's September 25 integration implements /v1/systemone locally using Laya, an encoder-based decision model. Its September 28 follow-up tightens request validation and supplies the request identifier expected by TypeSafe clients.
That makes a local alternative inspectable, rather than merely imaginable. It also makes naming important: the catalog accepts aliases such as jev-latest and resolves them to the configured Laya checkpoint. It is serving Laya, not Jev's weights. The feature must be enabled; available checkpoints have short, differing context limits, and the route rejects overflow. Shared request shape does not establish equivalent decisions.
For a team with short classification inputs and a reason to operate locally, that is a concrete candidate to test. We inspected the merged implementation, not an installed release or a live model run. Its reported zero output tokens describe a non-text-generation interface; they do not make local inference free. Hardware, latency, errors and escalation still belong in the comparison.
Count the work behind the answer
The TypeSafe adapter remains useful here as an accounting instrument. Its evaluation loop separates final-response tokens from cumulative returned-attempt totals. It can send malformed output back with a correction request. Corrective and transient retries are separate settings; both default to zero. Retries are a configured part of an experiment, not an unavoidable tax on rivals.
The September 22 v0.2.1 release, still the repository’s latest revision when checked on September 29, stops particular provider refusal and incomplete-response signals from consuming corrective retries. It also preserves missing usage as unknown, rather than zero. The regression cases show that a later response cannot repair an earlier missing count. Returned-response totals still cannot price a failed request that supplies no usable accounting.
None of this pins the vendor's historical charts to the reviewed adapter revision, or quantifies how much its output requirements contributed to the advertised advantage. It gives a buyer a better experiment: choose the necessary answer shape, fix a held-out error target, and compare hosted Jev, an LLM and a suitable local decision model with every attempt and escalation included. The useful winner is the implementation that meets the application's requirement at the cost the whole workflow actually incurs.