The latest llama.cpp release includes a fix for DeepseekV4 to clear cache only for sequences, optimizing performance. It also expands support for various hardware backends, including Apple Silicon, Intel, Vulkan, ROCm 7.2, OpenVINO, SYCL, CUDA 12/13, and HIP across macOS, iOS, Linux, Android, and Windows platforms.
Ongoing llama.cpp optimizations and broad hardware support reduce friction for local LLM inference, increasing demand for diverse compute hardware.