StartLux open-sources decision model that tops Jev on public benchmarks

StartLux open-sources decision model that tops Jev on public benchmarks

N
News Editor
2026-10-03 04:03:10
StartLux, a Shanghai AI company, has released StartLux-Decision, a fully open-source decision model family with 0.8B, 2B, 4B, 9B, and 27B variants plus quantized files for local deployment. According to the team’s self-test using public evaluation tools and a Sept. 28, 2026 snapshot of the Decision Index leaderboard, the 27B model scored 63.88 on Decision Index 0.2.1, ahead of Jev 1.13 at 57.91, and outperformed Jev on 31 of 38 benchmarks. In a separate seven-test suite from the Shanghai AI Lab Intern-Decision team, the model posted an average accuracy of 91.82%. StartLux also shared a head-to-head chess comparison in which both models saw the same board state, used no extra search, and chose each move directly. The company said StartLux-Decision-27B won 35 of 36 games and performed better on average loss, best-move hit rate, and blunder count. The release is being framed not just as a benchmark result, but as part of StartLux’s broader local AI stack, where a dedicated decision model handles high-frequency, low-latency choices inside agent systems while larger models take on more complex reasoning.

On Sept. 30, StartLux rolled out StartLux-Decision, a family of decision models that the company says is fully open source. In the team’s own test, using public evaluation tools against a Sept. 28, 2026 snapshot of the public Decision Index leaderboard, the 27B model scored 63.88 on Decision Index 0.2.1. Jev 1.13 scored 57.91. Across 38 benchmarks, StartLux-Decision finished ahead of Jev on 31.

StartLux open-sources decision model that tops Jev on public benchmarks 2

StartLux also put out a hands-on comparison using chess. Same board position. No extra search. Each model had to pick every move directly. StartLux said StartLux-Decision-27B produced checkmate and won 35 of 36 games against Jev 1.13. The company also said its model posted better average loss, best-move hit rate, and blunder-count numbers.

The release comes from Shanghai AI company StartLux, also called YuanDian XingHui. The article frames it as the firm’s second eye-catching result in less than five months since it was founded. Before this, the local model StartLux-27B beat the 284B DeepSeek-V4-Flash in testing by the China Academy of Information and Communications Technology under the Ministry of Industry and Information Technology, and came in roughly on par with DeepSeek-V4-Pro while using about 1/60 of the parameters.

Five model sizes and local deployment options

StartLux shipped the decision model in five sizes: 0.8B, 2B, 4B, 9B, and 27B. It also released quantized versions meant for local deployment.

  • GitHub: https://github.com/StartLuxLabs/Startlux-Decision
  • Hugging Face collection: https://huggingface.co/collections/startlux-models/startlux-decision-6abba92b301b573fa154d493

In the independent Decision Index 0.2.1 evaluation, the 27B model got 63.88. It ranked above Jev 1.13 on 31 of 38 benchmarks; the article calls Jev 1.13 Jev’s latest version. In a separate seven-test setup, StartLux-Decision-27B posted average accuracy of 91.82%.

That second test set came from the Shanghai AI Lab Intern-Decision team and follows a different method from Decision Index. And the article is explicit here: these numbers were self-reported by the StartLux team, based on public evaluation tools and a public leaderboard snapshot dated Sept. 28, 2026.

Why a dedicated decision model matters for agents

The report is built around one question: why would an agent system need a model focused specifically on judgment?

StartLux open-sources decision model that tops Jev on public benchmarks 3

In StartLux’s examples, the decision model is not mainly there to generate open-ended text. It is there to pick straight from predefined options, make binary calls, or assign grades and levels. That, in the article’s telling, is the dividing line between a decision model and a general generative model. A general model would usually write natural language first, then leave a program to parse it afterward.

One demo used a customer-support ticket with the message: "The order has been charged twice, but the system still shows unpaid." The system used StartLux-Decision to make three judgments in one request: send the case to the correct team, set the urgency, and assign a severity level.

Another demo was about shopping. The agent’s task was to buy "the cheapest 8-pack of AA batteries with free shipping and send it to the home address." The system judged model type, quantity, price, and delivery terms. At each turn, StartLux-Decision looked at the current page state, picked the next action, and decided whether the task was done. After every action, the page state was refreshed and the cycle repeated until the order had been placed and the task requirements had been confirmed.

A third demo moved to workplace collaboration. The target was to invite a specified member to the Design team and give that person the Editor role. The model had to choose the correct team, the correct member, and the correct permission level, then decide whether the invitation flow had been completed.

Those demos share one thing. The answer space is fairly clear. Browser controls, support departments, and tool lists are already defined. The article sets that against the open-ended generation tasks where large language models tend to shine, like drafting emails or building plans. But many agent steps are much smaller: a yes-or-no decision, one pick out of three, a five-level rating, or a tool selection.

The report places all of this in a longer shift. In traditional software, engineers hard-coded a lot of business logic. Rules decided which workflow applied after an amount crossed a threshold, which state triggered which action, and which permission branch a user entered. Then large language models absorbed much of that explicit logic, with one model handling understanding, classification, judgment, planning, and generation. In the agent stage, the article says, that labor is splitting again: deterministic flows go to code, search goes to embedding systems, complex planning goes to reasoning models, and high-frequency choices start moving toward specialized decision models.

Latency and compute allocation

StartLux also shared speed figures. Under a single H200 GPU, BF16, local HTTP, and short-request conditions, StartLux-Decision-4B needed an average of 26 milliseconds to answer three questions at once. The 0.8B and 2B versions averaged 12.2 milliseconds and 15.5 milliseconds.

StartLux open-sources decision model that tops Jev on public benchmarks 4

The team said it used CUDA Graph, fast linear attention operators, and a method that merged multiple questions into one forward pass, cutting repeated computation and invocation overhead.

The article’s argument is simple enough: once an agent task includes dozens of decision nodes, the compute spent on each step starts shaping total response time and operating cost. That is the backdrop it gives for the recent rise of decision models. A lot of these jobs look tiny at first glance: choosing a tool, deciding if a task is finished, checking which workflow a piece of information should enter, or figuring out the next click on a page. Small stuff. But when agents run for long stretches and at scale, those small judgments can turn into the most frequently called layer in the whole system.

How StartLux fits decision models into local AI

The article leans on a September interview with StartLux co-founder and CTO Guo Quanwei to explain what the company means by “local AI.” Guo was talking about more than a small offline model running on a PC. His definition covered models, quantization, inference systems, hardware adaptation, and upper-layer pieces such as Agent Harness, tools, search, and continuous learning, all aimed at deployment on users’ local devices.

That creates a much bigger systems problem. A local agent that runs for long periods has to deal with VRAM and memory limits, precision choices, task scheduling across 4B, 9B, and 27B models, long-context management, recovery after interrupted steps, memory read-and-write strategy, and the cost tradeoffs tied to higher-level judgment.

StartLux-Decision is presented as a part built for one specific layer in that stack. In StartLux’s internal breakdown, StartLux-27B is the general-capability layer. The decision model takes the high-frequency choices. Agent and Memory sit above them. Quantization, inference systems, and hardware adaptation sit below as support.

The logic starts with local compute limits. Cloud systems can stretch model size across GPU clusters. Local systems cannot. They have to keep running over time inside fixed limits on VRAM, memory, bandwidth, and power. So one question becomes central: how to use each unit of compute well.

StartLux open-sources decision model that tops Jev on public benchmarks 5

The article connects that point to earlier post-training experiments on StartLux’s 27B model. Guo described model parameters as a high-dimensional space where scale sets capacity and training redistributes abilities within that capacity. In those experiments, post-training aimed at agent capabilities pushed GPQA from 83.33 to 90.40, DeepSearchQA from 47.04 to 59.39, and also improved Tau3-Banking. At the same time, GAIA dropped from 57.57 to 45.70, and IFEval dropped too.

The takeaway, as the article sees it, is that fixed parameter budgets force tradeoffs. Training reallocates resources across abilities. Decision models take that idea one step further by offloading frequently called, clearly defined tasks to a lighter model built for the job.

Quantization data and engineering tradeoffs

That is also why StartLux released five model sizes and three GGUF quantization formats: BF16, Q8_0, and Q4_K_M.

For the 4B model, the Q8_0 file is about 4.48 GB. In the team’s published requantization retest on 231 public JevBench questions, that version hit 100% decision consistency with the original weights and answered 204 questions correctly, matching the original. The smaller Q4_K_M version, about 2.71 GB, posted 98.3% decision consistency and answered 201 questions correctly.

The article treats this size spread as a very practical engineering choice for a local agent system. A long-running personal AI stack may use several model types at the same time: a small decision model for low-latency classification, a larger model for harder reasoning, specialized modules for retrieval, Memory for long-term state, and Agent Harness for coordination across applications. In that view, models are starting to look like different processing units inside a broader compute system.

Another point the report highlights is probabilistic output. StartLux-Decision can return a probability distribution across candidate options, letting the system build routing and branching strategies. If the model is confident enough on a familiar task, the system can move ahead directly. If several candidates come out with similar probabilities, the system can ask for more information, hand the task to a stronger model, or trigger human confirmation.

But the article adds a clear warning. Those probabilities still need calibration against the business context. For high-risk actions such as payments, deletions, or permission changes, existing authorization and confirmation mechanisms are still required.

StartLux open-sources decision model that tops Jev on public benchmarks 6

A three-day first cycle

Another figure that caught attention: three days. According to the team, the stretch from defining the direction of the decision model to finishing the first round of development and validation was about three days.

The report credits that pace to StartLux’s long-built Auto Research system under its RSI framework. AI is already involved in data construction, training, evaluation, and failure analysis, and the article says the company’s last two model projects both showed that pattern.

In an earlier interview, Guo explained the number behind the claim that “70% of experiment execution is handled with AI participation.” If you narrow the scope to routine engineering work like config generation, training launch, evaluation runs, and result collation, automation is already above 95%, he said. The 70% figure refers to AI participation in the more central research and decision layer.

The report separates the easy part from the hard part. Automating training launch is relatively mature. The harder work starts after training ends: figuring out why a model failed, what data should be added, whether the issue sits in the algorithm or in the order of reasoning, which variable should be adjusted in the next experiment, and whether a line of work still deserves more compute. StartLux calls that higher-order process Auto Research.

The article gives a finance-related example. When the system found that a model had calculated directly instead of returning to the underlying source data, failed samples were fed into the research system. Multiple AI models then worked together to analyze causes and suggest hypotheses, such as weak data coverage, flaws in the reasoning chain, or missing reinforcement for the relevant behavior. The system then evaluated the value of each proposal, automatically built data, launched training, ran evaluations, checked for capability regressions, and fed the new results into the next round.

Inside that process, AI is taking on work the article describes as research decision-making. StartLux-Decision is presented as one implementation of that pipeline for a new model category. Once the team chose to build a decision model, it had to define interface standards and build datasets and evaluation systems around action selection, constraint understanding, and structured output. Abnormal cases from training were then pushed back into the iteration loop, with measured results deciding which changes stayed.

StartLux open-sources decision model that tops Jev on public benchmarks 7

StartLux’s RSI ladder

The article says that in a market where foundation models change fast and the lead time of any single weight release is getting shorter, the portability of the R&D system matters more. Guo said that in StartLux’s work, most accumulated data and training experience can be reused when switching within the same family of base models, while the portability of the pipeline, evaluation, and failure-analysis system is even higher.

From that, the company defines its long-term core assets as the ability to spot effective data, deal with failed samples, maintain an experiment evaluation framework, and automatically generate the next round of experiments.

StartLux then lays out five levels of Recursive Self-Improvement, or RSI:

  1. Level 0 is Self-correction, where the model can retry after detecting an error.
  2. Level 1 adds Memory and Experience accumulation.
  3. Level 2 is Automated Training, where AI can generate data, write code, launch training, and run evaluation.
  4. Level 3 is Auto Research, where AI analyzes failures from the previous round, proposes new hypotheses, and helps decide the next research direction.
  5. Level 4 is full RSI in Guo’s definition: once AI improves, its ability to improve AI also improves, forming a recursive feedback loop.

StartLux currently places itself at Level 3.

In that framework, StartLux-Decision is described in two ways at the same time. It is a product sped up by the Auto Research pipeline, and it is also a component that the same pipeline may be able to call later on. Automated research itself contains plenty of discrete decisions: which failures deserve deeper analysis, which tool to call next, how to rank several experiment plans, and when to stop chasing a line of work.

For tasks with a clearly defined candidate set, bringing a decision model into that loop is presented as a sensible next step. Still, the article says the real benefit has to be tested in future research workflows.

That fits StartLux’s broader view of Auto Research as a heterogeneous multi-model system. Models with different parameter counts, training paradigms, and functions — including prediction, discrimination, and algorithmic roles — work together in the process, with a verifier and algorithms deciding which experiments remain.

StartLux open-sources decision model that tops Jev on public benchmarks 8

The industry’s focus is shifting toward system integration

The last section places StartLux-Decision on a broader timeline. Starting with Jev’s release in mid-September, a little over two weeks brought OpenAI’s Decisions API launch at its developer event, Cloudflare’s release of the open-source decision model Clef, and continued momentum for previously visible projects such as Laya and Personal Agent. The article says the sector is reaching the same conclusion unusually fast: inside agent systems, the intelligence called most often is frequently not broad generation but small, high-frequency, tightly bounded judgment. Those tasks justify a dedicated layer that is cheap and responds in milliseconds, instead of calling a full large model every time.

The report also argues that StartLux got onto this path earlier. Citing Guo’s earlier interview with Machine Heart, it says the company’s idea of “local” was never “a relatively small model that can be downloaded to a computer.” It was “a complete personal AI system.” The bottom layer is made up of models, quantization, and inference systems. The middle layer is Agent Harness and Memory built for local environments. The top layer is the user-facing product. In that structure, StartLux-27B and StartLux-Decision map to the general-capability layer and the high-frequency decision layer, while work on infrastructure, quantization, and inference continues in parallel.

That is why the article does not present StartLux-Decision as a one-off release chasing a trend. It presents it as a planned piece inside a bigger system map.

The same reading is applied to the three-day development cycle. The report says the meaning is not just speed. It is a sign that StartLux’s Auto Research system has become portable across categories under its RSI approach. If data construction, evaluation, and failure-analysis systems can be reused from one model type to another, then the time from project kickoff to first validation can shrink in a systematic way.

The article ends with both a promise and a limit. A system does not prove itself on an architecture chart. It proves itself by completing real tasks over long periods with stability. Guo is quoted in the report saying that long-task stability, failure recovery, context management, and overall user experience are harder than simply getting a model to run. Based on the progress of filings, StartLux’s first public local-intelligence experience version may arrive within the year. The article says that will be the first full test of the company’s design.

This report originated from the WeChat public account Machine Heart (ID: almosthuman2014), with the byline listed as “Focused on AI.”

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.