ByteDance Cranks Up AI Video Generation Efforts with World Models
ByteDance, the parent company of TikTok, is making significant strides in developing artificial intelligence (AI) systems that can generate video in real-time with spatial awareness. This effort has been prioritized for 2026, with a focus on world models that encompass real-time interactive video and 3D-aware generation.
The technical building blocks for this project have been laid out through a series of releases from ByteDance's Seed research team. In February 2026, Seedance 2.0 was launched, which focused on physics-accurate multimodal generation. This was followed by Seedance 2.5 in July 31, which pushed the boundaries further with 30-second single-pass audio-video clips and enhanced spatial controls.
More recently, an arXiv paper published on August 24 introduced autoregressive adversarial post-training (AAPT), a model that can generate video at 24 frames per second with low latency. This model is capable of producing minute-long coherent outputs.
In March 2026, ByteDance released Helios, an open-weight model that achieved approximately 19.5 fps for minute-long videos running on a single GPU. Additionally, the systems accept pose and camera inputs, enabling applications like virtual human generation where users can direct the output rather than just prompt it.
ByteDance has elevated world models as a top-tier AI focus, with the goal of matching Google's Genie 3 benchmark for interactive world simulation. The company is putting significant resources behind this effort, with training data budgets reaching eight figures in RMB. The MultiMedia Lab is also commercializing 4D and 6DoF spatial video pipelines for volumetric content.