Naive AI open-sources its first model and puts AI-led AI research at the center

Naive AI open-sources its first model and puts AI-led AI research at the center

N
News Editor
2026-09-28 09:14:40
Naive AI, a startup founded by Dai Jifeng’s team, has released and open-sourced its first model, Naive-N0.5-Flash, positioning it as both a product built with AI assistance and a model trained to take on AI research work itself. The company’s technical report lays out a development process in which researchers define goals, constraints, and evaluation criteria, while AI systems handle large parts of implementation, experimentation, analysis, and validation. The release comes as major labs have started publishing hard numbers on how much AI is already contributing to model research. Anthropic recently said 26% of its internal AI research work is now “Claude-led,” up from less than 1% in February. OpenAI said it had reached its “automated research intern” goal, with about 3.1 agent workdays running in parallel for every human workday in its research team. Against that backdrop, Naive AI is presenting a more explicit version of the same direction: building a research system where AI is involved from day one in creating the next generation of models. Naive-N0.5-Flash has 309B total parameters, 15.5B active parameters, native 1M-token context, MIT-licensed weights and inference code, and API pricing disclosed in yuan. The report also details two case studies, NaiveRT and AutoWM, showing how the model was used in runtime optimization and world-model research, alongside benchmark results in paper reproduction, machine learning engineering, post-training, coding, and long-horizon tasks.

Naive AI has released its first model, Naive-N0.5-Flash, and framed the launch around a clear thesis: use AI to build the next generation of AI. According to the company’s technical report, the model was developed with AI participation and was also trained to take on AI research tasks itself.

Naive AI open-sources its first model and puts AI-led AI research at the center 2

That idea has moved quickly from theory to measurable practice across top labs over the past few weeks. Ten days ago, Anthropic published internal data showing that 26% of its internal AI research work is now “Claude-led,” meaning researchers provide a high-level instruction and Claude completes most of the task end to end under human supervision. In February, that figure was still below 1%.

A few days earlier, OpenAI said it had achieved the “automated research intern” goal it set last year. On a standard workday basis, the company said its research team now runs about 3.1 agent workdays in parallel for every one human workday invested.

On Sept. 21, OpenAI also proposed that the United States lead the creation of global technical standards for frontier AI and listed recursive self-improvement, or RSI, as a priority topic. In China, Zhipu said on Sept. 17 that an Infra Agent powered by GLM-5.3 had built the inference service for GLM-5.3-Flash from scratch on a domestic 100,000-GPU cluster.

Naive AI is taking that direction a step further. The startup, founded by Dai Jifeng’s team in February, was built around a research structure that does not keep humans at the center of every step. Coding, running experiments, monitoring workflows, and analyzing results are assigned to AI models. Human researchers focus on three things: setting goals, defining constraints and evaluation criteria, and making final decisions.

First model built on open-weight foundations

Naive AI is focused on post-training open models and building agents rather than pretraining a foundation model from scratch. Its approach starts with open-weight models and then moves through architectural modification, incremental pretraining, and post-training.

The report places that choice in a broader shift. Open ecosystems are becoming the starting point for a new wave of labs. Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, released its first model, Inkling, in July. Inkling was trained from scratch, but its architecture followed the mixture-of-experts design used in DeepSeek-V3.

Naive AI goes further by continuing training and modification directly on top of an open base model. In the company’s framing, the competitive question for new labs is no longer only whether they can train a base model from zero, but whether they can reshape an existing one faster and better for their own needs.

That is where Naive AI is placing its bet. Naive-N0.5-Flash has 309B total parameters, 15.5B active parameters, native support for a 1M-token context window, and a focus on coding and AI research.

Naive AI open-sources its first model and puts AI-led AI research at the center 3

The model has a dual role. It was built with AI involvement, and it was trained to perform AI research work. The technical report says AI participated in exploring attention architectures and in optimizing training, inference, and deployment systems. Researchers set goals, constraints, and evaluation standards. AI then implements candidate solutions, runs experiments, and adjusts based on results. Researchers still retain control over key decisions such as architecture choices.

Two case studies: NaiveRT and AutoWM

The company included two supporting case studies to show what that workflow looks like in practice.

In work on the NaiveRT inference runtime, researchers and AI ran 151 experiments over six days and reduced the time for one full speculative decoding cycle on the same system from 12.3 milliseconds to 3.4 milliseconds.

In another task, Naive-N0.5-Flash took on world-model research. After 400 cumulative hours and 15 major experiments, with ongoing adjustments to data, training plans, and inference strategy, the process produced AutoWM, which the report says has entered the top tier of WorldArena.

Naive AI intern and Tsinghua PhD student Shiqian Su also shared an example on X showing the use of Naive-N0.5-Flash to train a Jev model. The resulting multimodal Jev-like model outperformed Jev and the open-source Laya on the games Tetris and Snake.

In that setup, environment and data preparation, model and training-method selection, test construction, and the later stages of training, evaluation, debugging, and iteration were all handled by the agent. The report says the agent also adjusted its plan based on evidence gathered during the process.

Those examples are meant to show a broader attempt: putting models into the next round of research so they can help solve real problems in later iterations. The starting point, though, is still model capability itself. Can the model read papers, implement methods, and identify promising directions after a chain of experiments?

Benchmark results focused on AI research tasks

Naive AI’s first technical report presents results across paper reproduction, machine learning engineering, post-training, coding, and long-horizon tasks.

Naive AI open-sources its first model and puts AI-led AI research at the center 4

On PaperBench, which measures paper reproduction ability, Naive-N0.5-Flash scored 63.2, ahead of Claude Opus 4.7, GPT-5.5, and MiniMax M3 in the report. For a model aimed at AI research, that score offers a direct view into how much of a research method described in a paper it can turn into an actual implementation.

On MLE-bench-30, a benchmark derived from Kaggle competition tasks, the model scored 73.7%. Those tasks involve training models and running experiments around data.

On PostTrainBench, which requires language-model post-training under a fixed compute budget, it scored 37.5.

The report treats those benchmarks as covering different parts of the research process. Paper reproduction tests understanding and implementation of prior work. Machine learning engineering measures the ability to solve concrete tasks. Post-training tasks ask the model to search for improvements under resource constraints. Together, they are presented as evidence for whether a model can participate in research.

Another comparison in the report stands out. On Sol-ExecBench for GPU operator optimization, NanoChat AutoResearch for training improvement under a fixed budget, and NanoGPT SpeedRun for training acceleration, the comparison target is Recursive Superintelligence, a company focused on automated AI research. Richard Socher is its CEO, and Tian Yudong is listed as a co-founder.

That company said in June that its system reduced NanoGPT Speedrun from 79.7 seconds to 77.5 seconds, after a community had spent more than two years optimizing it and set 83 human record refreshes along the way. Naive-N0.5-Flash pushed that result to 73.8 seconds.

Coding ability underpins much of this work. On NL2Repo, which asks models to generate a code repository from natural-language requirements, Naive-N0.5-Flash scored 71.9, above DeepSeek-V4.1-Flash. On ALE-CLI, a benchmark for long-horizon professional tasks, it scored 32.4, close to GPT-6 Astra and Opus 5.5.

The report also points to a practical feature of research work: the longer a task runs, the more historical information it tends to accumulate. A training error may need to be traced through earlier code changes. Choosing the next experiment may require revisiting prior results and configurations. The model’s native 1M-token context is meant to preserve that history across code, tool outputs, and experiment logs.

Naive AI open-sources its first model and puts AI-led AI research at the center 5

Under the disclosed release plan, Naive-N0.5-Flash weights and inference code are licensed under MIT, and an API will be provided. Pricing for input, output, and cache reads is set at 0.6 yuan, 2.6 yuan, and 0.07 yuan per million tokens, respectively.

How the model was built

The first technical challenge in developing Naive-N0.5-Flash was efficiency at a 1M-token context length. The team started from the open-weight MiMo-V2.5 base model. Most layers in that model use sliding-window attention, or SWA, to focus on local information, while a smaller number of global-attention layers preserve long-range information. At a million-token context, those global-attention layers account for a large share of decoding cost.

Naive AI targeted those layers and replaced global attention with DeepSeek Sparse Attention, or DSA. DSA first uses a lightweight indexer to score historical positions, then selects 2,048 of them for the backbone network to attend to. SWA handles a local 128-token window. The network mainly uses a pattern of five SWA layers for every one DSA layer, forming a hybrid sparse-attention architecture.

The team also used GQA with four KV groups and designed a lightweight indexer with 16 query heads to further adjust indexing and attention cost. The report notes that the DSA indexer still scans the full history and that those layers still keep the full KV cache. The main gain comes from reducing backbone attention computation and memory access.

Reducing the indexer’s query heads from 64 to 16 significantly lowered long-context sparse-attention latency. The report says this design was explored jointly by researchers and AI. Researchers set the target and evaluation protocol, requiring better long-context decoding efficiency without sacrificing model quality. AI implemented candidate architectures, ran ablation studies, and summarized the results. Among the options that met the target, researchers chose the one with the simpler engineering implementation.

After the attention structure changed, the model had to adapt to a new way of reading information. Under the 1M-token setup, Naive-N0.5-Flash went through 3.25T tokens of multi-stage training.

  • The first 50B tokens were used to warm up the indexer. The rest of the model was frozen, the layers scheduled to switch to DSA continued using global attention, and the indexer was trained with KL-divergence loss using those attention distributions as supervision.
  • The model then switched to sparse attention and continued pretraining for 3T tokens with language-modeling loss, adapting to the new information-selection and aggregation mechanism while strengthening coding and AI research ability.
  • Finally, it underwent supervised fine-tuning on another 200B tokens, with the learning rate gradually reduced.

AI also handled a validation task during the transition. Using global-attention results as a reference, it analyzed the recall of the top-2048 positions selected by the indexer. The report says AI found and fixed a numerical-stability issue in top-k selection and then independently verified the fix. Researchers used that work to determine the switching point and precision threshold and to unify validation standards across training and deployment.

Training system and memory management

Another layer of work happened in the training system. Million-token sequences have to be distributed across multiple GPUs, and the way information moves between cards directly affects training efficiency.

Naive AI open-sources its first model and puts AI-led AI research at the center 6

When analyzing the hybrid architecture, AI identified a key difference in access patterns. DSA needs the full history. SWA only needs a local left-side window. Based on that, it proposed and validated a hybrid sequence-parallel scheme: DSA layers use Ulysses sequence parallelism and all-to-all redistribution to access the full sequence, while SWA layers use overlapping shards and exchange only the required overlap with neighboring left-side shards.

That reduces communication complexity for the SWA part from O(L), which grows with sequence length, to O(w), which depends on window size. The same idea was extended to inference, where long-context prefilling, row-wise sharded computation, and index retrieval each use a matching context-parallel strategy.

Memory management was also refined. For intermediate activations created by sparse indexing and attention alignment, AI analyzed recomputation paths and reduced offloading granularity to the output of individual operators, allowing different intermediate results to use different offloading strategies. Top-k indices were kept explicitly so the selected positions could be reused during recomputation. Checks on long sequences also surfaced positional-encoding precision issues and the risk of sequence-index overflow.

According to the disclosed figures, the training system can complete 1T tokens of training in about four days under the 1M-token setup using 512 GPUs. The infrastructure supporting research operations also includes sandboxing, compute, and permission-management platforms. The company says those systems support nearly 10 million sandbox runs per week, with peak concurrency reaching 100,000.

Those details give concrete meaning to the phrase “AI participates in research.” In the report, that means understanding computational dependencies, finding numerical or performance issues, and using experiments to test whether changes work. Researchers set constraints and make key decisions. AI handles a large share of analysis, implementation, and validation.

NaiveRT: 151 experiments in six days

NaiveRT is an inference runtime aimed at a practical issue in long-horizon reinforcement learning. A single trajectory may need to generate a large number of tokens, and a small number of especially slow trajectories can hold up an entire training batch. Faster decoding for each sequence can reduce waiting time for those samples.

To address that, researchers and AI ran 151 optimization experiments in six days, and 63 of the resulting changes were adopted. AI analyzed whole-network performance, implemented candidate solutions, and checked results through numerical validation and multi-GPU testing. The work covered operator fusion, data movement, and GPU execution scheduling.

Some directions failed. One MoE operator-fusion path was tried for seven consecutive rounds. Every version produced correct numerical results, but end-to-end performance fell each time. The analysis showed that most of the original operator-boundary overhead had already been hidden by execution overlap, while fusion introduced extra synchronization. Researchers stopped that line of work.

Naive AI open-sources its first model and puts AI-led AI research at the center 7

In the final result, NaiveRT reduced one full speculative decoding cycle, including draft generation, verification, sampling, and commit, from 12.3 milliseconds in SGLang to 3.4 milliseconds on the same system, a 72.4% latency drop.

On eight frontier GPUs, the system also recorded a peak single-stream decoding speed of 2,122 tokens per second. The report says that figure came from the best one-second window among 41 HTML/SVG generation requests, with thinking mode disabled and prefilling time excluded.

The report breaks down the 151 experiments as well: 63 changes were adopted, 71 failed validation or were rolled back, and 17 were exploratory. Each downward step in the progress curve corresponds to one or more changes entering the actual execution path.

AutoWM: 400 hours and 15 major experiments

AutoWM extends the same approach into world-model research. A researcher gave Naive-N0.5-Flash a research goal, a compute budget, and an evaluation protocol. From there, the model kept pushing forward on design, training, and evaluation, deciding what to try next based on experimental results.

The work began with reproducing FlowWAM. Early adjustments to classifier-free guidance and timesteps produced limited gains, so the model shifted toward data and training design. It rewrote video descriptions, expanded the training set from 2,500 videos to 22,500, filtered out lower-scoring samples, and adjusted frame sampling and rollout structure. Results improved further after switching to a stronger base model and pairing it with higher-quality data.

As the experiments progressed, the model found that some WorldArena metrics could be computed without real reference videos and could therefore guide output selection. It began searching outputs generated at different timesteps, modeled frame selection as a knapsack problem solved with dynamic programming, and then explored Best-of-N video selection and post-processing strategies. Larger search scales did not always help. In one setup, N=32 performed worse than N=16.

After 400 cumulative hours and 15 major experiments, the team measured a score of 77.43 on video-quality evaluation under the public WorldArena-1 Track 1 protocol, above the then-public best score of 73.64 recorded in the report.

The report says that result reflects the combined contribution of training improvements, inference-time filtering, and post-processing, while also documenting how the model adjusted its research path in response to experimental feedback.

Naive AI open-sources its first model and puts AI-led AI research at the center 8

Placed inside the RSI debate

Naive AI presents these efforts as part of an exploration of recursive self-improvement. The idea is to let models participate in the research and development of later models and systems, then gradually expand the range of work they can take on.

The company says that when it entered the large-model race in 2026, it chose a different route: AI building AI. Starting from open base models, it is strengthening the Naive series around research tasks while also letting AI participate deeply in architecture changes, training, and system optimization. The newly released technical report and the two case studies are meant to provide a first set of concrete, testable outputs from that strategy.

The next question is what those outputs can do for the next round of research. Faster inference can shorten generation wait times in experiments. Better ability to read papers, write code, and analyze results could help researchers find effective directions sooner. If those gains keep accumulating, the team may be able to run more productive exploration under the same time and resource budget.

That is the RSI loop Naive AI says it wants to build: let AI participate in the development of later models and systems, then use the improved capability to push the next round of improvement.

The industry, though, is still debating what RSI actually means. In its Sept. 21 proposal, OpenAI said fully autonomous RSI has not yet happened and should not be pursued before it can be done safely. Anthropic has also stressed that Claude is not yet fully autonomous in any category of research work included in its statistics.

Dream-RSI, released in mid-September by Google and Google DeepMind among others, takes a different route. Model weights stay fixed, while agents learn better search strategies from past search records. In other words, different groups are answering different questions about what exactly is being improved: the weights, the system, or the method of doing research.

Naive AI’s current practice still keeps researchers in charge of direction while AI handles a large amount of execution. What sets it apart in the report is the stated goal: to gradually merge the product of research with the executor of research. The Naive series is both the model produced by this system and the researcher trained to work inside it. The two case studies show the potential for AI to take on concrete research tasks. Whether that can keep translating into progress for the next generation of models will depend on later iterations.

This article was sourced from the WeChat public account “机器之心” (ID: almosthuman2014), by RSI.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.