In September TypeSafe released the first public model in its System One class of models.1 The model, JEV (version 1.13), takes its class name from Kahneman’s System 1, which stands for fast and intuitive thinking. Rather than generating text, the model takes a text and a list of questions with fixed answers as input and then outputs an answer to each question along with a probability for every answer option. TypeSafe claims the model is much cheaper and faster to run than LLMs, and that it produces better-calibrated probabilities through a post-training process it calls “reinforcement learning for calibrated decisions.”
If you annotate or scale political text, you are probably wondering whether JEV is really cheaper, faster and better calibrated than other options, and whether it gives up any accuracy to get there. To find out, Matt DiGiuseppe and I re-ran seven applications from the political science literature in a new working paper, now on arXiv (DiGiuseppe and Denney 2026). For each application, we scored JEV against the same human benchmark as the original paper, along with a current low-cost commercial model (GPT-6 Luna) and an open-weight model you can run locally (Qwen3.8-27B).2
Overall, JEV is fairly close to the other models in accuracy. It is generally a bit worse than Luna on the reading tasks but better than both models on the V-Dem codes. It is not cheaper than Luna at batch prices, but it was much faster for us. With each question asked only once, it also produced better-calibrated probabilities than Luna’s token probabilities. Compared with Qwen, it depends on the task.
Seven Text as Data Replications
We examine seven applications of annotation and scaling. In the first five, the model has access to the text itself. The last two, which we call recall tasks, give the model just a name, and it must answer solely on the basis of what it learned during training.
tweet relevance for content moderation (Gilardi et al. 2023);
sentiment in tweets about two Supreme Court decisions (Ornstein et al. 2025);
attack types in 37,709 Global Terrorism Database incidents (Brandt et al. 2026);
pairwise scaling of open-ended survey answers on economic knowledge (DiGiuseppe and Flynn 2026);
sentence-level positions in UK party manifestos and European Parliament speeches (Le Mens and Gallego 2025);
left–right positions of European parties from their names alone (Di Leo et al. 2025);
V-Dem expert codes for 53 indicators in 171 countries, again from the country name alone (Weidmann et al. 2026).
To be as consistent with the original studies as possible, we used their prompt wording where JEV allowed it but left out the instructions about the format of the answer. In cases where a study had used a numeric scale, we rephrased the scale in terms of JEV’s worded answer levels. In cases where a study had allowed for a “not applicable” answer, we added a separate yes/no question instead. All of these changes are described in the appendices of our paper. I go through the questions in order, starting with accuracy, then cost and speed, and finally the probabilities.
Accuracy: close, but rarely ahead of Luna
Figure 1 shows the level of agreement that each model achieves with the human benchmark (each study uses its own measure for this). No single model outperforms all others, with JEV underperforming Luna on tweets, manifestos and EP speeches, matching it on survey pairs and party positions, and doing slightly better at identifying the type of attack. It also outperforms both models (and the published GPT-4o codings) on V-Dem. Compared with Qwen, JEV does better on attack types and both recall tasks, worse on tweet sentiment and the 2023 tweets, and about the same elsewhere.

One caveat applies to the recall tasks. Both the V-Dem data (version 14) and the expert party placements were publicly available before all three models were trained. To the extent that they were part of the training data, this is a form of data contamination, and some of the agreement may come from the models recalling the benchmark. This matters most for the comparison with GPT-4o’s V-Dem codes, since GPT-4o was trained before V-Dem version 14 was released.
Cost and speed: faster, not cheaper
At OpenAI’s Batch prices (which apply when you are willing to wait up to a day for results), Luna is between 34 and 55 percent cheaper than JEV on five of the seven tasks and within 3 percent of JEV on the other two (see Table 1). If you need your results immediately and pay the more expensive Standard prices, JEV is cheaper on six of the seven tasks.

TypeSafe’s pricing is less helpful than it first seems. TypeSafe does not charge for output, which matters for a reasoning LLM that writes many output tokens, but in a classification task the labels are so short that input makes up most of the cost. Luna wrote around four tokens for each answer, while JEV billed 1.2 to 3.3 times as many input tokens per decision, because it counts the question and all of the answer options as input. Qwen was more expensive still through a hosted API, at 2.35 to 5.41 times the cost per decision of JEV, but a 27B model is small enough that you might be able to run it on a good computer or on your university’s supercomputer (like Leiden University’s ALICE).
In our runs, JEV was much faster than the other two models (0.09 to 0.27 seconds per decision, against 0.75 to 1.04 for Luna and 0.39 to 0.87 for Qwen), although for most projects you can run a batch job and come back the next day for results.
Probabilities: easy to get, and better calibrated than Luna’s
TypeSafe’s third claim, and in my view the most interesting, is about JEV’s probabilities. Researchers typically use such probabilities in one of two ways. Some send the answers the model is unsure about to a human coder, or search a corpus for a rare category. What matters then is that the probabilities rank the right answers above the wrong ones, which AUROC measures. Others build a continuous scale from the probabilities or read them as rates. That requires the probabilities to be calibrated, meaning that the model is right 80% of the time when it gives an answer 80% probability, which the expected calibration error (ECE) measures. Chat-tuned LLMs are known to be overconfident, and many APIs give token probabilities for only a few of the top tokens, or none at all (OpenAI gives at most the top five for Luna, and none when it reasons).
Take a look at Figure 2 here. In the party-comparison task, 74% of Luna’s probabilities are smaller than 0.01 or larger than 0.99. For JEV the figure is 18%, and for Qwen 16%. If you ask a model which of two parties is more right-wing and the two are close to each other on the political spectrum, a reasonable answer is a probability somewhere in the middle of the range. JEV and Qwen give that kind of answer far more often than Luna does. Since Qwen does it too, near-certainty is not inherent to LLMs in general. Or at least not to Qwen.

Extreme probabilities are a warning sign, but they are not proof of poor calibration. To check calibration, we compare the probabilities with the human labels, and Figure 3 shows the results. It covers every task where we have human labels and probabilities from all three models, which includes the tasks we replicated and four tasks we added. On all eight reading tasks where we ask each question only once and read Luna’s probability from its answer token, Luna has a higher calibration error than JEV. When we ask each pair in both orders and average, the gap narrows, and on the survey pairs it disappears. On the task where Luna writes out its probability (attack type), Luna is better calibrated than JEV.
Against Qwen, each model is better calibrated on its own kind of task, and by quite a lot. Qwen is better calibrated on all four pairwise tasks. JEV is better calibrated on three of the four labelling tasks, on attack type and on both recall tasks. On V-Dem, for example, JEV’s calibration error is 0.06, while Qwen’s is 0.15.

The ranking results point the same way. Compared with JEV, Luna’s AUROC is lower on seven of the eight tasks where each question was asked once, and equal on the other. Compared with Qwen, it is lower on all eight. JEV and Qwen are not consistently different from each other.
JEV! Should you use it?
JEV does not save you money if you can wait for batch results, and it is not more accurate than Luna on most reading tasks. On the other hand, it is very fast, and unlike Luna it gives you a clean probability for each answer, rather than making you infer it from a handful of tokens that often have probabilities close to 0 or 1. For most social science work, the probabilities are the more useful of the two. The speed difference is big, but it does not matter unless you need the answers immediately, as in market research, rather than a day later from a batch job.
Qwen is at least as accurate as JEV on most reading tasks and better calibrated on the pairwise ones. If you are able to run an open-weight model yourself, for example on a university cluster, you may not need JEV. JEV is ahead of Qwen on the recall tasks and on attack type, though, in both accuracy and calibration.
TypeSafe has not released any of JEV’s internals (its architecture, training data or reward), so there is little to judge it on beyond how it performs on tasks like these. As with any LLM, that means validating it against human coding for the task you are interested in.
***
The full paper, with all seven replications, the prompts and the appendices, is on arXiv: Denney, Steven, and Matthew DiGiuseppe. 2026. “JEV versus LLMs: Accuracy, Cost and Calibration on Seven Political Science Replications.” Preprint, arXiv, October 5. https://doi.org/10.48550/arXiv.2610.06625
On 6 October, after we had finished our runs, OpenAI released a similar endpoint in public beta, the Decisions API. It runs on GPT-6 Luna, one of the models we compare JEV to, and returns probabilities for all answer options. It also only charges for input, like JEV does. We have not tested it.
Both GPT-6 Luna and Qwen3.8-27B ran with reasoning turned off. Qwen ran through OpenRouter on a single host (Parasail) at 8-bit precision, not on our own hardware.


