Agentic GPU Programming for MLSys#
Machine learning systems depend on fast GPU kernels for training and serving. Attention, matrix multiplication, and fused operators account for substantial work in these systems. Improving their implementations can reduce the time and resources needed to run a model.
Making these kernels fast requires in-depth knowledge of algorithms, GPU hardware, and programming models. Coding agents can increasingly carry out this work: finding relevant implementations, writing kernels, interpreting diagnostics, and choosing what to try next. Using these capabilities effectively requires a workflow that coordinates the work and an environment that gives agents access to the knowledge and feedback they need.
This book introduces the main elements of agentic GPU programming. You will learn how a compiler-driven harness equips an agent to express optimization ideas, diagnose errors, measure performance, and retain the results worth keeping. Together, these capabilities form an effective optimization flow.
The book develops a compiler foundation, domain-specific program analysis, a knowledge base, and GPU benchmarking and profiling. It then shows how agent workflows compose with this compiler harness to guide an optimization run.
The material grows out of real-world experience building and using agentic GPU programming systems. It is a companion book to Modern GPU Programming for MLSys, and we plan to integrate it into the Machine Learning Systems course series at Carnegie Mellon University.
This book is open source. Contributions, corrections, and examples are welcome through the GitHub repository.
How This Book Is Organized#
Part I, Elements of Agentic GPU Programming. This part introduces the compiler harness and the agent workflows that compose with it. It opens with an overview, then develops domain-specific program analysis, a knowledge base, and benchmarking and profiling through examples and interactive diagrams. It closes with a short chapter on different kinds of agent workflows and how they compose with the compiler harness.
Part II, TIRx Harness Overview. This part introduces TIRx, one concrete instance of a compiler harness and the environment used for the examples in the rest of the book. It walks through the TIRx compiler foundation and the KCoral benchmark server with its execution interface.
Part III, Agentic GPU Programming in Action. This part is a hands-on tutorial that puts the preceding chapters to work and carries agentic GPU programming out end to end. It looks closely at the feedback the harness returns over the course of a run and at the interaction patterns between the agent and each element of the harness, and it closes with a few advanced tips. Grouped GEMM introduces the workflow; recorded Kimi Delta Attention runs supply the diagnostic and review examples.
Part I, Elements of Agentic GPU Programming
Part II, TIRx Harness Overview
Part III, Agentic GPU Programming in Action