参考
A Formal Analysis of the NVIDIA PTX Memory Consistency Model by Lustig et al., 2019
Dissecting the Turing GPU Architecture through Microbenchmarking (Jia et al., 2019)
PTX Documentation: Link
SASS Instruction List: Link
https://mlc.ai/modern-gpu-programming-for-mlsys/chapter_background/index.html
FlashAttention-4: Faster Attention with Asynchrony and Low-precision