One command, one remote experiment
Each command sends a script and its inputs to the benchmark server
Development machine · the agent runs
python evolution/remote/kcoral_python.py --remote http://10.0.2.2:8901 \
--send artifacts/gate-in-gram/gputest \
-e CASES=256:standard,512:saturated,...,320:corr_saturated \
-e H=8 --timeout 600 -- test_small.pyGPU host · a worker runs, with the GPU leased
test_small.py written by the agentRuns the candidate on small shapes for each stress regime, after filling both outputs with NaN, and compares them with an fp64 token-by-token recurrence.
What comes back
T=256 H=8 standard output: finite=True rms_ratio=1.4041e-03 ... bad=0.00e+00 T=256 H=8 standard state: finite=True rms_ratio=4.5680e-04 ... bad=0.00e+00 T=512 H=8 saturated output: finite=True rms_ratio=3.1333e-03 ... bad=0.00e+00 ...
A quick gate the agent designed for itself
Twelve lines with no bad elements before the official benchmark; a hang shows up here as a timeout instead of wasting an official run.
Development machine · the agent runs
python evolution/remote/kcoral_remote.py candidates/kda/forward_b1_t8192_h96 \
scratch/gram-swp-c --remote http://10.0.2.2:8901 --timeout 1500GPU host · a worker runs, with the GPU leased
task evaluator supplied by the taskOwns the inputs, the correctness checks on ten workloads, the private holdout, and the timing method. The baseline is timed in the same request.
What comes back
holdout correctness: PASS ... baseline: 1.050103 ms kernel: 0.453369 ms speedup: 2.3162x ... passed: 10/10
The official score
The agent cannot change how it is measured, only what it submits. This run produced the best kernel of the run.
Development machine · the agent runs
python evolution/remote/kcoral_iket.py --remote http://10.0.2.2:8901 --timeout 1200 \
--send artifacts/fused96/iket --output-dir artifacts/fused96/iket_out \
-e T=2048 -e H=96 -e REPEAT=1 \
-- profile --postprocess json -- python capture.pyGPU host · a worker runs, with the GPU leased
capture.py written by the agentCompiles the kernel with IKET range markers and launches it once; the wrapper returns the trace files.
What comes back
$ python3 artifacts/fused96/iket_analyze.py iket_pid_*.trace.json steady-state chunk period (warp13 m-ks start deltas): mean 3.948 us over 96 CTAs, ... CTA spans (us): min 129.1 median 131.5 max 133.0
Summarized locally
The trace comes back as files; the agent’s own script turns it into the time per chunk and per stage, without another GPU run.
Development machine · the agent runs
python evolution/remote/kcoral_ncu.py --remote http://10.0.2.2:8901 \
--send candidates/kda/forward_b1_t8192_h96/scratch/profile-overlap-diag-inverse \
-o artifacts/overlap-diag-inverse-gram-tail/kda-full.ncu-rep \
--set full --launch-count 1 --kernel-name kda_fwd_kernel -- python capture_ncu.pyGPU host · a worker runs, with the GPU leased
capture_ncu.py written by the agentLaunches the target kernel once; NCU profiles the launch whose name matches and the wrapper returns the report file.
What comes back
$ ncu --import kda-full.ncu-rep --page details | rg -i 'Duration|Throughput|...'
Memory Throughput % 35.52
Duration us 135.33
Compute (SM) Throughput % 30.99Read as often as needed
The report is opened locally with ncu --import; later sessions reread the same file instead of profiling again.