Perplexity AI Unveils Custom Serving Infrastructure for Faster, Cheaper AI Search
Perplexity AI has been working on custom serving infrastructure to make its AI search engine more efficient. The company, which handles around 400 million queries a month, has been building this infrastructure since at least February 2025.
The core of Perplexity's retrieval improvements is the pplx-embed family of models. These compact embedding models come in two sizes: a 0.6 billion parameter variant and a larger 4 billion parameter version. Both were released in February 2026, with quantization techniques reducing storage requirements by up to 32x.
The more ambitious piece of infrastructure is the Runtime-Optimized Serving Engine (ROSE). Introduced around February 2025, ROSE is designed to serve a wide range of model architectures, including Mixture-of-Experts (MoE) format. Perplexity addressed MoE serving with custom RDMA networking optimizations that achieve up to 97.1% bandwidth efficiency.
In June 2026, Perplexity unveiled a hybrid local-cloud inference orchestrator. The system automatically routes tasks between on-device models running on local hardware and more powerful cloud-based models, depending on the complexity of the query.