9d4ae950 | arXiv:2205.14135v2 [cs.LG] 23 Jun 2022 FLASHATTENTION: Fast and Memory-Efficient | ['q02', 'q04', 'q11', 'q12']
3ed43e6e | Algorithm 3 FLASHATTENTION-3 backward pass with warp specialization Require: Mat | ['q02', 'q03', 'q04', 'q06', 'q07', 'q11', 'q12']
ae6d8565 | it can begin executing a different warp. In practice, this can be more performan | ['q03', 'q04', 'q09', 'q12']
ba7760bd | Waiter spins in SMEM Waiter sleeps until all threads arrive Direct copy to SMEM  | ['q03', 'q04', 'q07']
4aef3fe9 | Manuscript submitted to ACM Dissecting the NVIDIA Hopper Architecture through Mi | ['q03', 'q04', 'q07', 'q12']
31f4cc99 | 486 continues on next page 20. Compute Capabilities CUDA C++ Programming Guide,  | ['q03', 'q04', 'q07', 'q12']
6a6bb659 | www.nvidia.com Kernel Profiling Guide v2021.2.1 33 Chapter 7. ROOFLINE CHARTS Ro | ['q04', 'q07']
57c99885 | Pipelines of simple map operations can be optimized by tradi- tional loop fusion | ['q04', 'q05', 'q11', 'q12']
96ae5bec | The key to working around these problems is exploiting the GPU's L2 cache, which | ['q04', 'q07', 'q11']
cb57ee01 | Undoubtedly, both larger datasets and datasets for new domains will serve as imp | ['q04', 'q06']
3c0e6e5e | are Tesla-qualified. Our T4 and P4 experiments ran on HPE Proliant DL360 Gen9 se | ['q04', 'q12']
8c69fe39 | We present Triton, a language and compiler centered around the concept of tile,  | ['q05', 'q11', 'q12']
bebb11b3 | arXiv:2006.06762v5 [cs.LG] 15 Oct 2023 ARTIFACT EVALUATED usenix ARTIFACT EVALUA | ['q05', 'q11', 'q12']
c761f479 | ABSTRACT FP8 is a natural progression for accelerating deep learning training in | ['q06']
31b6314f | Larger models usually require more compute and memory resources to train. These  | ['q06']
c6acfa61 | Occupancy is the ratio of the number of active warps per multiprocessor to the m | ['q07']
db840e88 | i !! Chapter 1. Overview Guide The NVIDIA® Magnum IO GPUDirect® Storage Overview | ['q08']
5365100b | In this paper, we observe that existing LLM serving sys- tems [31, 60] fall shor | ['q08', 'q11']
07f3c182 | Prior work With the introduction of multi-core and GPU devices, multiple paralle | ['q10', 'q11']
ca78c2d6 | 1.11.8 Reducing the memory consumption AlphaFold model has high memory consumpti | ['q10']
6575255c | Code availability Source code for the AlphaFold model, trained weights and an in | ['q10']
dddfbc83 | To address this challenge, we developed MMseqs2, a parallelized, open-source sof | ['q10']
22ab01a2 | Published work, together with our experiments, consistently show that bare-metal | ['q12']
