This training provides a practical introduction to optimizing PyTorch-based AI models on GPUs, with a strong focus on performance profiling and custom Triton kernels. The aim is to equip attendees with the skills needed to identify performance bottlenecks and accelerate GPU-based computations. By the end of the training, attendees will be able to:
- manage CPU–GPU memory transfers and reason...
This training provides a practical introduction to optimizing PyTorch-based AI models on GPUs, with a strong focus on performance profiling and custom Triton kernels. The aim is to equip attendees with the skills needed to identify performance bottlenecks and accelerate GPU-based computations. By the end of the training, attendees will be able to:
- manage CPU–GPU memory transfers and reason...