Cheapest gate first — a reviewer can stop at any level of effort.
# 0 · seconds, DATALESS (pip install only): the scoring harness proves itself
PYTHONPATH=src python -m evemtbench.evaluation.cli --self-test
# 1 · seconds, read-only: coverage + the analytic floor self-check
PYTHONPATH=src python -m evemtbench.coverage --summary
PYTHONPATH=src python -m evemtbench.coverage --check-floor
# 2 · seconds: are the committed rendered views (incl. this site) byte-fresh?
python hpc/check_results_fresh.py
# 3 · minutes (SLURM — never on the login node): the full test gate
srun --cpus-per-task=4 --mem=8G --time=15:00 bash -c 'PYTHONPATH=src pytest'
# 4 · regenerate any single cell (its exact arg line is on the coverage page / snapshot)
PYTHONPATH=src python -m evemtbench.baselines.runner \
--task <task> --protocol <protocol> --grid <grid> --baseline <baseline> --seeds 0 1 2 3 4
# 5 · rebuild every rendered view from outputs/
PYTHONPATH=src python hpc/make_results.py && PYTHONPATH=src evemtbench-leaderboard \
&& PYTHONPATH=src python hpc/make_audit.pythe exact invocation behind every ledger row
| Claim | Command | Artifact |
|---|---|---|
| Fault detection (FD, binary): learnable in-distribution and scored on the never-trained benchmark family | PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4 | outputs/<version>/fault_detection_local/held_out/double_line/<baseline>/*/result.json |
| Fault-type classification (FC, 9-class): learnable in-distribution and scored on the never-trained benchmark family | PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_classification_local --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4 | outputs/<version>/fault_classification_local/held_out/double_line/<baseline>/*/result.json |
| Fault localisation (FL, regression): learnable in-distribution and scored on the never-trained benchmark family | PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_location_local --protocol held_out --grid double_line --baseline gru --seeds 0 1 2 3 4 | outputs/<version>/fault_location_local/held_out/double_line/<baseline>/*/result.json |
| Event detection (ED, binary): learnable in-distribution and scored on the never-trained benchmark family | PYTHONPATH=src python -m evemtbench.baselines.runner --task event_detection_global --protocol held_out --grid double_line --baseline gru --seeds 0 1 2 3 4 | outputs/<version>/event_detection_global/held_out/double_line/<baseline>/*/result.json |
| Event classification (EC, 23-class): learnable in-distribution and scored on the never-trained benchmark family | PYTHONPATH=src python -m evemtbench.baselines.runner --task event_classification_global --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4 | outputs/<version>/event_classification_global/held_out/double_line/<baseline>/*/result.json |
| Cross-grid transfer from the multi_grid pretrain corpora (zero-shot vs fine-tune) under the corrected topology-grouped source splits (v1.1) | PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol transfer_finetune --grid testgrid_110kv --baseline resnet --seeds 0 1 2 3 4 | outputs/<version>/*/transfer_*/*/*/*/result.json |
| Protocol-driven benchmark: every result is conditional evidence, never one score | PYTHONPATH=src python hpc/make_results.py | RESULTS.md |
| Five tasks with frozen label derivations; one scoring rule per window | PYTHONPATH=src python -c "from evemtbench.tasks.registry import BENCHMARK_TASKS; print(len(BENCHMARK_TASKS))" | src/evemtbench/tasks/labels/ |
| Three nested observability views on four reference grids spanning 20–345 kV | PYTHONPATH=src python -c "from evemtbench.data.roots import GRIDS; print(list(GRIDS))" | src/evemtbench/data/roots.py |
| Episode-disjoint, family-pure 70/15/15 committed splits with leakage controls | pytest tests/integration/test_leakage.py tests/integration/test_splits_audit.py -q | src/evemtbench/splits/v1.0/ · src/evemtbench/splits/v1.1/ |
| Transparent, beatable baseline suite from trivial floor to deep learning | PYTHONPATH=src python -m evemtbench.baselines.runner --help | src/evemtbench/baselines/ |
| The benchmark family is scored but never fitted or model-selected on | jq .attestation outputs/<version>/<task>/<protocol>/<grid>/<baseline>/run_record.json | outputs/<version>/*/*/*/*/run_record.json |
| Analytic floor self-check over the whole matrix | PYTHONPATH=src python -m evemtbench.coverage --check-floor | outputs/<version>/ (read-only check) |
| Rendered views are byte-fresh: RESULTS.md, leaderboard.html and this audit site | python hpc/check_results_fresh.py | RESULTS.md · leaderboard.html · audit/ · coverage_snapshot.json |
| Frozen window contract: 50 ms windows, 5 ms stride, 480 samples @ 9.6 kHz | pytest tests/integration/test_contract_conformance.py -q | src/evemtbench/contract/contract.json |
| Every run emits a machine-readable provenance record | jq '.git_sha,.lockfile_sha256,.host' outputs/<version>/<task>/<protocol>/<grid>/<baseline>/run_record.json | outputs/<version>/*/*/*/*/run_record.json |
| Paper result tables are machine-generated from result.json (a live gated view) | PYTHONPATH=src python hpc/make_paper_tables.py | results/ |
| Coverage: 580 of 580 held_out cells measured | PYTHONPATH=src python -m evemtbench.coverage --summary | coverage_snapshot.json |
| The scoring harness is self-testing — dataless controls + a per-run oracle check | evemtbench-evaluate --self-test | outputs/<version>/*/*/*/*/run_record.json |
| Headline results survive alternative aggregation | PYTHONPATH=src python hpc/make_audit.py | audit/robustness.html |
| First reference evaluation, released as an extensible reproducible foundation | PYTHONPATH=src python hpc/make_results_bundle.py | results/ |