EvEMTBench is a protocol-driven ML benchmark for power-system protection on EMT waveform windows. Labels are exact by construction — every window's label derives from the simulation script that injected the event — so the honest risks are leakage, held-out contamination, stale published tables and trivial-task artefacts. Each has a mechanical guard:
src/evemtbench/splitssrc/evemtbench/coverage.pydocs/determinism.mdhpc/check_results_fresh.py| Field | What it gives the reviewer |
|---|---|
| Value | the measured number, read from the committed result file at build time |
| Script | link to the exact source that produces it |
| Command | the verbatim invocation to reproduce it |
| Artifact | the output (result.json / rendered file) it writes |
| Runtime & Hardware | wall-clock and the machine it ran on |
| Verified by | the test / gate that guards it |
| Status | measured vs in progress vs deferred — a plan is never dressed up as a result |
Reference docs: protocol — what is scored and how · determinism — what is and is not reproducible · submission — artifact schemas & repro requirements · metrics — the frozen metric contract · grids — the four reference topologies