NVIDIA Unveils Secure AI Inference on Blackwell GPUs
NVIDIA has made significant strides in its confidential computing technology, enabling secure AI inference on its Blackwell GPUs. The company's latest advancements allow for the execution of large language model workloads in memory-encrypted environments, retaining 96-98% of performance seen in non-confidential configurations.
This solution caters to high-value, regulated industries such as healthcare and financial services, where data security during inference is critical. By incorporating confidential virtual machines (CVMs), encrypted NVLink, and confidential Blackwell GPUs, NVIDIA aims to make secure deployments viable without significant performance trade-offs.
NVIDIA evaluated the performance impact of enabling confidential computing using its TensorRT LLM inference framework. The company's tests demonstrated that confidentiality can coexist with near-peak performance, introducing a latency overhead of just 1.2-4.3% at concurrency levels of 1-16.