EvEMTBench — results audit

benchmark version 1.1.0 · held_out results v1.0.0 · splits v1.0/held_out, v1.1/multi_grid
updated 2026-09-29 13:18 UTC · git 7243f310ddbc

EvEMTBench is a protocol-driven ML benchmark for power-system protection on EMT waveform windows. Labels are exact by construction — every window's label derives from the simulation script that injected the event — so the honest risks are leakage, held-out contamination, stale published tables and trivial-task artefacts. Each has a mechanical guard:

Frozen, audited splitsEpisode-disjoint, family-pure train/val/test/benchmark splits are committed to the repo and re-audited by named tests on every CI run — no training window shares an episode with an evaluation window (held_out v1.0; multi_grid pretrain v1.1, topology-grouped).
code tour → · src/evemtbench/splits
Analytic floor self-checkThe majority baseline must land on the closed-form 1/k chance floor and no learned baseline may fall below it — a known-answer check over the whole results matrix, re-runnable in seconds.
code tour → · src/evemtbench/coverage.py
Never-trained benchmark family + attestationThe benchmark family is scored but never fitted or model-selected on; every run commits a machine-readable no-tuning attestation alongside git SHA, lockfile hash and seeds in its run record.
docs/determinism.md
Byte-fresh rendered viewsRESULTS.md, leaderboard.html and this audit site are pure functions of outputs/ — a pre-commit gate fails any commit where the committed pages differ from a fresh rebuild, so a stale number cannot be published.
code tour → · hpc/check_results_fresh.py
580of 580 held_out cells done
1876of 1876 cells across all protocols
1876provenance records
24tasks × 7 baselines
Claim ledger →
every paper claim resolved to script, command, artifact, runtime, hardware and verifying gate
Paper numbers →
every number the manuscript reports, traced to output file, eval code, training log, training code, config and W&B run
Coverage matrix →
every cell of the results matrix: done / pending / failed / N-A-by-rule, with its reproduce command
Verification gates →
a guided code tour of the five gates that guard the numbers
Environment →
hardware, pinned software, seeds, and input-data hashes
Reproduce →
cheapest gate first: from a 5-second self-check to regenerating any single cell
Provenance →
tests, CI, hooks and the change log

How to read a claim

FieldWhat it gives the reviewer
Valuethe measured number, read from the committed result file at build time
Scriptlink to the exact source that produces it
Commandthe verbatim invocation to reproduce it
Artifactthe output (result.json / rendered file) it writes
Runtime & Hardwarewall-clock and the machine it ran on
Verified bythe test / gate that guards it
Statusmeasured vs in progress vs deferred — a plan is never dressed up as a result

Reference docs: protocol — what is scored and how · determinism — what is and is not reproducible · submission — artifact schemas & repro requirements · metrics — the frozen metric contract · grids — the four reference topologies