Less than a week after Jev was released, the open-source community had already started reproducing the idea of a model that does not write out an answer and instead returns options with probabilities. The shared goal is simple: skip token-by-token generation and output choices and probabilities directly. The implementations, though, have already split into several very different tracks.
Reading scores from existing models
SemIf takes the most straightforward route and does not retrain a model at all. It uses an existing Qwen model, stops the process when the model is about to answer A, B, or C, reads the score for each option, and converts those scores into probabilities.
Simple Jev also relies on an existing large model, but tries to cut repeated computation further. The same source material is read once, then reused across multiple questions before separate answer judgments are made. Its author said clearly that the project reproduces Jev's usage pattern, not full equivalence in accuracy, speed, or probability calibration.
Dedicated decision models
Laya drops the text-generating large model entirely and uses an encoder model that looks more like a traditional classifier. The upside is a smaller and faster system. The weakness is also clear. In one task with 77 answer choices, Laya reached 42.5% accuracy, while Jev posted 87.0%.
Von follows the same dedicated decision-model direction. It has about 400 million parameters, does not generate text, and completes Choice, Noul, and Score in a single forward pass. The result looks closer to a new classifier that can read natural language than to a standard language model.
Verdict is smaller still, at about 150 million parameters. Its author focused on two issues: cases where the model is wrong but still reports very high confidence, and cases where changing the order of answer options changes the result as well. The project puts more weight on making both probabilities and judgments stable.
Keeping a large model but adding decision structure
Kev keeps the large-model backbone. It uses Qwen as the base and adds a dedicated decision structure so the same input is read once while multiple questions each get their own probability calculation. A third party that had previously reverse-engineered Jev had also suggested that Jev might use a similar design.
In Kev's own tests on new questions not seen during training, the 8B version reached about 78% accuracy, compared with about 86% for Jev. The larger gap showed up in confidence: Kev still had some wrong answers that were assigned confidence above 90%.
Nimble puts its emphasis on training. It uses Qwen3.5-9B and prepares a set of post-training data where changing only one fact flips the correct answer. On 324 test samples that were not part of training, Nimble selected the reference answer in 90.1% of cases, versus 93.2% for Jev. Its author also warned that the model still cannot guarantee that a reported 90% confidence really corresponds to a 90% chance of being correct.
Using diffusion to fill answers directly
OpenJev goes the furthest by switching to DiffusionGemma. It leaves several answer positions blank first, then fills them in one shot in a way that resembles how a diffusion model completes an image. On an RTX PRO 6000, it reported a median latency of about 94 ms for a single request.
Four paths in a matter of days
After only a few days, OpenJEV-related work has already split into four routes: directly reading answer scores from existing models, training dedicated decision models, turning large models into decision models, and using diffusion models to fill in answers directly.
So far, the results suggest that Jev's overall direction may be easy to imitate. The real separation point is still accuracy, generalization, and probability calibration. A model may be willing to report 90% confidence. The harder question is how trustworthy that 90% really is.

