Scroll horizontally to explore the full diagram.
measure Collect
Supports the components above
TIRx Harness: the components from Part I
TIRx provides the compiler foundation beneath the knowledge base, program analysis, and GPU evaluation. Select a component to inspect its role.
Agent workflow: kernel development
The agent uses the harness to express kernels, find implementation knowledge, diagnose errors, and obtain GPU measurements.
TIRx foundation: program representation
TIRx provides the program representation used for compilation and analysis. We write kernels through TIRx-lite, a low-level Python interface.
Knowledge base: implementations and hardware references
TIRx Harness’s knowledge base includes a kernel zoo with ports from FlashAttention, DeepGEMM, and FlashInfer. Hardware documentation complements these implementations with instruction semantics and requirements.
Domain-specific compiler analysis: correctness tools
TIRx Harness provides NumSim for numerical simulation, Synccheck for synchronization analysis, and Racecheck for data-race analysis. These CPU-based tools inspect supported program behavior and complement tests on the target GPU.
KCoral benchmark server: GPU evaluation
KCoral coordinates GPU access for correctness checks, timings, and profiling captures. IKET captures kernel timelines, and Nsight Compute (NCU) supplies hardware measurements. Evaluation code defines checks and timing rules; KCoral provides the managed execution environment.
GPUs: target hardware
The GPUs execute compiled kernels and profiling captures. They can reside on a server or edge device separate from the agent’s development machine.
Self-improvement across runs
Retained traces can reveal gaps in the harness tools, while validated kernels can extend the knowledge base. Part III discusses how to review optimization results for these improvements.