Profiling
measure
Measure and diagnose
The benchmark server runs correctness checks, timing, and profiling on the target GPUs. It returns results and profiles that the agent uses to compare candidates and investigate bottlenecks.
Choose the next experiment
The agent submits candidate code, inputs, and evaluation or profiling code. It compares benchmark results and uses profiles to decide which stage or hardware behavior to investigate next.
Coordinate experiments on the target GPUs
The server manages GPU access, cleans up between runs, and returns results and artifacts. It distinguishes completed measurements from timeouts and execution failures.
Did the candidate improve?
Check the outputs, then compare candidate and baseline timings under the same protocol. The task's evaluation code defines correctness and timing rules; the server provides the execution environment.
Which stage holds up progress?
IKET records named stages inside the kernel by CTA and warp. The timeline shows long waits, stage overlap, and work imbalance, helping the agent locate the part of the schedule to investigate.
What explains the delay?
Nsight Compute reports throughput, memory traffic, occupancy, and stall metrics. These measurements help the agent investigate a bottleneck and test hypotheses such as bandwidth limits or register spills.
Execute and measure the candidate
The target GPUs execute the candidate kernels. The benchmark server coordinates their use so competing experiments do not interfere with correctness checks, timings, or profiles.