IBM and RPI Propose New Error Protection Microarchitecture for HBM AI Inference
Researchers from Rensselaer Polytechnic Institute and IBM's T.J. Watson Research Center have proposed a new microarchitecture called REACH to strengthen error protection for High-Bandwidth Memory used in AI inference workloads.
The team, led by Rui Xie, notes that the cost of High-Bandwidth Memory motivates stronger controller protection capable of supporting a wider range of device error rates.
They suggest using established inner codes to correct common errors and identify unresolved chunks, reserving a long outer code for known-erasure repair.