SGLang has introduced a native decision API that allows off-the-shelf large language models such as Qwen3.8-27B, as well as multimodal models, to handle selection, judgment, and scoring tasks without changing model weights or running another round of fine-tuning. Instead of generating free-form text first and then parsing the response, developers can provide the current state and a set of candidate answers, and the model returns the probability for each option directly. The method uses the model’s built-in next-token probabilities rather than altering the model itself.
SGLang also demonstrated the approach with Qwen3.8-27B in Pokémon FireRed. In the demo, the model used the live game state to decide whether to attack, switch Pokémon, or heal, with each decision taking less than 100 ms. It cleared the Elite Four and the Champion in one run. Because Qwen3.8-27B supports image input, the same setup can also be used for multimodal decision-making. In addition, SGLang added a /v1/systemone endpoint compatible with the TypeSafe SDK used by Jev, allowing applications already built on that SDK to keep working by pointing to an SGLang instance. The feature is now in the nightly build and is planned for release in v0.5.21.
SGLang has added a native decision API that lets existing large language models such as Qwen3.8-27B, along with multimodal models, perform selection, judgment, and scoring tasks without changing weights or retraining.
Under the setup, developers provide the current situation and several candidate answers. The model then returns the probability of each option directly, instead of generating a block of text first and having that output parsed afterward.
Reading option probabilities directly
According to the description, the approach uses the model’s built-in next-token probabilities. If the candidates are A, B, and C, a standard chat model would usually continue generating text. SGLang reads the model’s probability assigned to A, B, and C and checks which answer the model prefers. The model itself is unchanged; only the way the output is read is different.
Pokémon FireRed demo runs with sub-100 ms decisions
SGLang used Qwen3.8-27B in a Pokémon FireRed demo. Based on the live game state, the model decided whether to attack, switch Pokémon, or heal. Each decision took less than 100 ms, and the run cleared the Elite Four and the Champion in one pass.
Because Qwen3.8-27B supports image input, the same method can also handle multimodal decision-making.
New endpoint added for Jev SDK compatibility
SGLang also introduced a /v1/systemone endpoint that is compatible with the TypeSafe SDK used by Jev. Applications already using that SDK can keep calling the service after changing the service address to their own SGLang instance.
The new feature has entered the nightly version and is planned for formal release with v0.5.21.
This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan. Disclaimer:
The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.
Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.