PyTorch Just Made AI Training Safer by Fixing a Silent Bug
A small code change could prevent mysterious crashes in AI apps you use daily.
Packing a 16-bit atomic means fastAtomicAdd needs a bound on how far it may look for a pairing partner. Callers holding a raw pointer supply that bound themselves, and getting it wrong is silent — too large, and the pairing can write past the allocation. That's why roughly 20 scatter sites that subscript a packed accessor still use the plain atomic rather than hand-deriving a bound.
This PR adds an accessor overload that derives the offset and bound itself: offset is the sum over d of index[d] * stride[d], and span is 1 plus the sum over d of (size[d] - 1) * stride[d]. Span is the distance to the last addressable element, not the element count — the two differ for a non-contiguous accessor, whose gaps are legal pairing targets.
It converts the 5 sites that subscript a packed accessor (ReplicationPadding, UpSampleLinear1d, FractionalMaxPool2d/3d). They lose their index arithmetic and no longer have a bound to get wrong. It also adds coverage for the accessor overload's derived bound, including the non-contiguous case where element count and last addressable offset differ — the case a
- PyTorch fixed a bug that could silently corrupt data during AI training, leading to crashes or wrong results.
- The update automatically calculates safe memory limits, so developers don't have to do error-prone math.
- This makes AI models more reliable and could reduce frustrating bugs in apps you use.
Why It Matters
This fix makes AI software more stable, reducing crashes and errors in apps you rely on.