Reproduce

benchmark version 1.1.0 · held_out results v1.0.0 · splits v1.0/held_out, v1.1/multi_grid
updated 2026-09-29 13:18 UTC · git 7243f310ddbc

Cheapest gate first — a reviewer can stop at any level of effort.

# 0 · seconds, DATALESS (pip install only): the scoring harness proves itself
PYTHONPATH=src python -m evemtbench.evaluation.cli --self-test

# 1 · seconds, read-only: coverage + the analytic floor self-check
PYTHONPATH=src python -m evemtbench.coverage --summary
PYTHONPATH=src python -m evemtbench.coverage --check-floor

# 2 · seconds: are the committed rendered views (incl. this site) byte-fresh?
python hpc/check_results_fresh.py

# 3 · minutes (SLURM — never on the login node): the full test gate
srun --cpus-per-task=4 --mem=8G --time=15:00 bash -c 'PYTHONPATH=src pytest'

# 4 · regenerate any single cell (its exact arg line is on the coverage page / snapshot)
PYTHONPATH=src python -m evemtbench.baselines.runner \
    --task <task> --protocol <protocol> --grid <grid> --baseline <baseline> --seeds 0 1 2 3 4

# 5 · rebuild every rendered view from outputs/
PYTHONPATH=src python hpc/make_results.py && PYTHONPATH=src evemtbench-leaderboard \
    && PYTHONPATH=src python hpc/make_audit.py

Per-claim commands

the exact invocation behind every ledger row

ClaimCommandArtifact
Fault detection (FD, binary): learnable in-distribution and scored on the never-trained benchmark familyPYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4outputs/<version>/fault_detection_local/held_out/double_line/<baseline>/*/result.json
Fault-type classification (FC, 9-class): learnable in-distribution and scored on the never-trained benchmark familyPYTHONPATH=src python -m evemtbench.baselines.runner --task fault_classification_local --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4outputs/<version>/fault_classification_local/held_out/double_line/<baseline>/*/result.json
Fault localisation (FL, regression): learnable in-distribution and scored on the never-trained benchmark familyPYTHONPATH=src python -m evemtbench.baselines.runner --task fault_location_local --protocol held_out --grid double_line --baseline gru --seeds 0 1 2 3 4outputs/<version>/fault_location_local/held_out/double_line/<baseline>/*/result.json
Event detection (ED, binary): learnable in-distribution and scored on the never-trained benchmark familyPYTHONPATH=src python -m evemtbench.baselines.runner --task event_detection_global --protocol held_out --grid double_line --baseline gru --seeds 0 1 2 3 4outputs/<version>/event_detection_global/held_out/double_line/<baseline>/*/result.json
Event classification (EC, 23-class): learnable in-distribution and scored on the never-trained benchmark familyPYTHONPATH=src python -m evemtbench.baselines.runner --task event_classification_global --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4outputs/<version>/event_classification_global/held_out/double_line/<baseline>/*/result.json
Cross-grid transfer from the multi_grid pretrain corpora (zero-shot vs fine-tune) under the corrected topology-grouped source splits (v1.1)PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol transfer_finetune --grid testgrid_110kv --baseline resnet --seeds 0 1 2 3 4outputs/<version>/*/transfer_*/*/*/*/result.json
Protocol-driven benchmark: every result is conditional evidence, never one scorePYTHONPATH=src python hpc/make_results.pyRESULTS.md
Five tasks with frozen label derivations; one scoring rule per windowPYTHONPATH=src python -c "from evemtbench.tasks.registry import BENCHMARK_TASKS; print(len(BENCHMARK_TASKS))"src/evemtbench/tasks/labels/
Three nested observability views on four reference grids spanning 20–345 kVPYTHONPATH=src python -c "from evemtbench.data.roots import GRIDS; print(list(GRIDS))"src/evemtbench/data/roots.py
Episode-disjoint, family-pure 70/15/15 committed splits with leakage controlspytest tests/integration/test_leakage.py tests/integration/test_splits_audit.py -qsrc/evemtbench/splits/v1.0/ · src/evemtbench/splits/v1.1/
Transparent, beatable baseline suite from trivial floor to deep learningPYTHONPATH=src python -m evemtbench.baselines.runner --helpsrc/evemtbench/baselines/
The benchmark family is scored but never fitted or model-selected onjq .attestation outputs/<version>/<task>/<protocol>/<grid>/<baseline>/run_record.jsonoutputs/<version>/*/*/*/*/run_record.json
Analytic floor self-check over the whole matrixPYTHONPATH=src python -m evemtbench.coverage --check-flooroutputs/<version>/ (read-only check)
Rendered views are byte-fresh: RESULTS.md, leaderboard.html and this audit sitepython hpc/check_results_fresh.pyRESULTS.md · leaderboard.html · audit/ · coverage_snapshot.json
Frozen window contract: 50 ms windows, 5 ms stride, 480 samples @ 9.6 kHzpytest tests/integration/test_contract_conformance.py -qsrc/evemtbench/contract/contract.json
Every run emits a machine-readable provenance recordjq '.git_sha,.lockfile_sha256,.host' outputs/<version>/<task>/<protocol>/<grid>/<baseline>/run_record.jsonoutputs/<version>/*/*/*/*/run_record.json
Paper result tables are machine-generated from result.json (a live gated view)PYTHONPATH=src python hpc/make_paper_tables.pyresults/
Coverage: 580 of 580 held_out cells measuredPYTHONPATH=src python -m evemtbench.coverage --summarycoverage_snapshot.json
The scoring harness is self-testing — dataless controls + a per-run oracle checkevemtbench-evaluate --self-testoutputs/<version>/*/*/*/*/run_record.json
Headline results survive alternative aggregationPYTHONPATH=src python hpc/make_audit.pyaudit/robustness.html
First reference evaluation, released as an extensible reproducible foundationPYTHONPATH=src python hpc/make_results_bundle.pyresults/