The latest llama.cpp server release enhances its edit tool and introduces a tools_io abstraction, streamlining local LLM inference workflows. It also adds support for CUDA 13.3 DLLs on Windows x64, alongside existing support for CUDA 12.4, Vulkan, ROCm 7.2, OpenVINO, SYCL, HIP, and OpenCL Adreno across various platforms.