One command, one remote experiment

Each command sends a script and its inputs to the benchmark server

Development machine · the agent runs

python evolution/remote/kcoral_python.py --remote http://10.0.2.2:8901 \
    --send artifacts/gate-in-gram/gputest \
    -e CASES=256:standard,512:saturated,...,320:corr_saturated \
    -e H=8 --timeout 600 -- test_small.py

GPU host · a worker runs, with the GPU leased

test_small.py written by the agent

Runs the candidate on small shapes for each stress regime, after filling both outputs with NaN, and compares them with an fp64 token-by-token recurrence.

What comes back

T=256 H=8 standard     output: finite=True rms_ratio=1.4041e-03 ... bad=0.00e+00
T=256 H=8 standard     state: finite=True rms_ratio=4.5680e-04 ... bad=0.00e+00
T=512 H=8 saturated    output: finite=True rms_ratio=3.1333e-03 ... bad=0.00e+00
...

A quick gate the agent designed for itself

Twelve lines with no bad elements before the official benchmark; a hang shows up here as a timeout instead of wasting an official run.