The update introduces full support for f16 SET_ROWS operations, equivalent to f32, across Vulkan and CPU backends, enhancing performance and efficiency for inference workloads.
llama.cpp's expanded f16 support improves inference efficiency on a wider range of hardware, reducing compute requirements for local LLM deployments.