The update defaults the Hadamard mul_mat operation to a CPU routine, enhancing performance for CPU-based inference. It also expands support across various platforms, including new CUDA 13.3 DLLs for Windows x64 and KleidiAI enablement for macOS Apple Silicon.
llama.cpp's continued optimization for CPU and diverse hardware broadens access to local LLM inference, reducing reliance on high-end GPUs.