PyTorch increased shard counts for its longest-running trunk tests, including CUDA 13.0 (NVIDIA L4 and g4dn.12xlarge GPUs) and Windows CPU pools, to reduce wall-clock time per shard to under an hour. This enhancement aims to improve time-to-signal and node utilization by better distributing test workloads across available compute resources.
PyTorch's infrastructure optimization improves GPU and CPU utilization for its CI/CD, making existing compute more efficient for development.