forked from Karylab-cklius/vllm
Add per-token NaN/Inf detection to all RMSNorm CUDA kernels by piggybacking on the existing variance reduction. NaN/Inf propagates through sum-of-squares naturally, so a single isnan(variance) || isinf(variance) check on thread 0 after the CUB reduction detects it -- zero additional kernel launches, memory reads, or register pressure. Each block writes to its own int8 flag slot (no atomics). Controlled by VLLM_NAN_DETECT=1. When enabled: - Bypasses Oink/batch-invariant paths (with warning) - Reports per-token, per-layer NaN/Inf with layer names - Distinguishes real-token errors from padding-token warnings - CUDA-graph compatible (fixed flag buffer address) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: Tyler Michael Smith <tlrmchlsmth@gmail.com>