Building a Production-Grade GEMM Library with oneDNN's BRGeMM Micro-Kernels
Part 2 of the “Accelerating Deep Learning on Modern CPUs” series Recap: Where We Left Off In Part 1, we walked through the evolution of matrix multiplication on CPUs — from a naive triple loop ...