Optimizing Recommendation Syst... Note

Optimizing Recommendation Systems with JDK’s Vector API

Netflix's Ranker service, responsible for personalized recommendations, had a CPU hotspot in its video serendipity scoring feature. Initially, the scoring utilized a costly, nested loop structure for calculating similarities. The first optimization involved batching calculations to leverage matrix multiplication, improving the algorithm. However, an initial implementation using matrix multiplication introduced overhead, like GC pressure and poor cache behavior. Further optimization involved switching to flat buffers, improving data layout and cache locality. Then, a thread-local buffer reuse strategy eliminated per-request allocations. BLAS libraries were explored but didn't provide significant gains due to JNI overhead. Finally, the JDK Vector API was integrated, enabling SIMD, with a fallback to optimized scalar code for compatibility. This streamlined approach resulted in a 7% drop in CPU utilization and a 12% drop in average latency. Ultimately, the CPU consumption of the scoring feature was reduced from 7.5% down to roughly 1%, and the CPU per request significantly decreased by 10%.
CdXz5zHNQW_TmUGwTILLl.png