Stanford professor Percy Liang opens Marin 535B model training to the public

Stanford professor Percy Liang opens Marin 535B model training to the public

N
News Editor
2026-08-24 13:07:10
Stanford professor and Simile AI founder Percy Liang has kicked off training for Marin 535B-A23B and is making the process public in real time. The model is described as having 535 billion total parameters and 23 billion active parameters, trained on 18.75 trillion tokens across 11 GB200 NVL72 systems, or about 792 GB200 GPUs. Liang said the run is expected to last roughly three months and consume about 2.7e24 FLOPs, followed by a post-training phase. Before starting the main run, the team completed a four-stage scaling ladder ranging from 1.6B-A61M with 48 billion tokens to 27.7B-A1.2B with 926 billion tokens. Liang said those smaller runs were used to surface and debug issues early and to forecast the behavior of the larger “hero run.” The project has drawn attention because it exposes details that are usually hidden in frontier model development, including data mix, processing methods, training configuration, code, experiment design, live loss curves, and model-state tracking. The team has also published links to a data overview page, a Weights & Biases dashboard, and a GitHub issue page where the work can be followed in detail.

Stanford professor and Simile AI founder Percy Liang said Marin Open Lab has started training Marin 535B-A23B and is running the entire process in public. Over the past two days, the announcement gained traction on X after user Max For AI wrote, 「Too crazy — now you can literally watch a 535B model being trained live.」

Stanford professor Percy Liang opens Marin 535B model training to the public 2

According to Liang, Marin 535B-A23B has 535 billion total parameters and 23 billion active parameters. The team prepared 18.75 trillion tokens for training and deployed 11 GB200 NVL72 systems, equivalent to about 792 GB200 GPUs.

The run is expected to continue for about three months, with total training compute estimated at roughly 2.7e24 FLOPs. After the main training run is finished, the team plans to continue with post-training.

A four-stage scaling ladder came first

Liang said the team trained a four-stage Scaling Ladder before launching the main job. Those runs ranged from 1.6B-A61M with 48B tokens to 27.7B-A1.2B with 926B tokens.

Stanford professor Percy Liang opens Marin 535B model training to the public 3

He said the smaller-scale experiments served two purposes: finding and debugging potential problems in advance, and using training behavior at different scales to forecast the team’s “hero run.”

Liang also said, 「This is the largest training run we have done so far, so we really do expect some unexpected situations during the process.」

The main model is already running, with live metrics open

The main run is now live, and users can directly monitor the training curve in real time. Liang said the team will continue disclosing training data details, experiment logs, and engineering issues as the run progresses.

Stanford professor Percy Liang opens Marin 535B model training to the public 4

The public materials include:

  • data-level details such as training-data mixture ratios, processing methods, and how domain-specific data enters the model;
  • engineering details including training configuration, code, and experiment design;
  • process metrics such as live loss changes, model state, and predictions for different stages.

The team also published specific links for following the run through a data overview page, a Weights & Biases dashboard, and a GitHub page with detailed information.

Public links

Data composition page:
https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling

Stanford professor Percy Liang opens Marin 535B model training to the public 5

Live training on Weights & Biases:
https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng

Details on GitHub:
https://github.com/marin-community/marin/issues/8435

Opening up a process that is usually treated as a black box

The announcement drew attention in part because the training process behind models such as GPT-4, Claude, and Gemini is usually opaque to outsiders. People generally only see end-model capabilities, benchmark scores, and limited paper disclosures. Data ratios, failed runs, effective hyperparameters, loss behavior, and whether scaling-law predictions held up are rarely visible.

Stanford professor Percy Liang opens Marin 535B model training to the public 6

Marin is trying to change that by treating model training more like an open process that can be observed, discussed, and worked through in public.

Community discussion has moved into the training details

Reactions on X have continued since the announcement. Some users described the effort as the real spirit of open source, saying the team is not hiding problems even if the run drifts off course or encounters failures.

Others said the striking part is not just the size of the model, but the fact that people can now watch a large-model run evolve step by step through WandB. One post compared it to opening the door to a dark room inside a top AI lab and letting everyone see what is happening inside.

Stanford professor Percy Liang opens Marin 535B model training to the public 7

Some discussion has already become highly technical. James Thewlis, who has been following the run in real time, asked why norms showed a spike around step 500. He later answered his own question in the comments: 「This spike is entirely caused by router_bias. It is not a parameter learned through gradient training, but part of the token balancing heuristic in an MoE model, so it follows dynamics that differ from ordinary model parameters.」

The original report said there have been many similar exchanges, turning the project into not only a viewing channel but also a place for learning and discussion.

Liang also teaches a course on building language models from scratch

Beyond the live training run, Liang also offers a theory course, CS336: Language Models From Scratch, aimed at helping students fully understand how a large language model is built.

Stanford professor Percy Liang opens Marin 535B model training to the public 8

Course link:
https://www.youtube.com/playlist?list=PLoROMvodv4rMqXOcazWaTUHhq-yembLCV

The report said the next three months will not only be about the model’s final capability. Observers will also be watching for the failures, crashes, and unexpected events that may appear during training.

Reference links cited in the report include related X posts from Max For AI, Percy Liang, and CopyRebeldia. The article was sourced from the WeChat public account Jiqizhixin, with authorship credited as 「关注AI的」.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
70

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.