Stanford Lab Kickstarts Massive AI Model Training
Stanford AI Lab recently announced that Percy Liang started training for the Marin 535B-A23B model in an open process. The pretraining stage, which covers 80% of the run, is expected to last around three months and will be done on 11 GB200 NVL72 clusters. This is part of a larger effort to validate AI model pretraining processes and scaling laws through high-performance computing.
The researchers first validated their approach using a four-rung scaling ladder from 1.6B-A61M to 27.7B-A1.2B models on smaller token counts to debug systems and forecast performance. This validation process is crucial in ensuring the success of the larger project. The training data consists of 18.75T tokens, which will be used for roughly three months at a rate of 2.7e24 FLOPs before post-training begins.
The use of high-performance computing and open processes highlights the lab's commitment to transparency and collaboration in AI research.