The latest llama.cpp release routes large matrix multiplications to medium tiles on Adreno GPUs and fixes `llama-cli` breaking issues for longer prompts with `q4_0` quantized networks due to insufficient shared memory. The update also includes various platform builds for macOS, iOS, Linux, Android, and Windows, supporting different hardware backends like Vulkan, CUDA, ROCm, and OpenVINO.