Claim ledger

benchmark version 1.1.0 · held_out results v1.0.0 · splits v1.0/held_out, v1.1/multi_grid
updated 2026-09-29 13:18 UTC · git 7243f310ddbc

One row per claim of the paper. Every number below is a placeholder filled from the committed artifacts at build time — the manifest source contains no digits (see hpc/audit_manifest.py). Rows whose cells are still being swept show in progress with live counts; planned work is deferred, never a result.

Fault detection (FD, binary): learnable in-distribution and scored on the never-trained benchmark family0.883 ± 0.030 test / 0.989 ± 0.008 benchmark (balanced accuracy; best learned baseline resnet; floor 0.500)measured

Representative cell: fault_detection_local on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.

Script
src/evemtbench/baselines/runner.py
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4
Artifact
outputs/<version>/fault_detection_local/held_out/double_line/<baseline>/*/result.json
Runtime
~7.5m fit per seed (compute.fit_seconds_mean)
Hardware
NHR@FAU Alex, 1× A40 · seeds 0–4
Verified by
floor gate · leakage gate
Fault-type classification (FC, 9-class): learnable in-distribution and scored on the never-trained benchmark family0.855 ± 0.007 test / 0.896 ± 0.045 benchmark (balanced accuracy; best learned baseline resnet; floor 0.143)measured

Representative cell: fault_classification_local on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.

Script
src/evemtbench/baselines/runner.py
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_classification_local --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4
Artifact
outputs/<version>/fault_classification_local/held_out/double_line/<baseline>/*/result.json
Runtime
~7.5m fit per seed (compute.fit_seconds_mean)
Hardware
NHR@FAU Alex, 1× A40 · seeds 0–4
Verified by
floor gate · leakage gate
Fault localisation (FL, regression): learnable in-distribution and scored on the never-trained benchmark family18.600 ± 0.213 test / 11.695 ± 0.530 benchmark (MAE, lower better; best learned baseline gru; floor 24.416)measured

Representative cell: fault_location_local on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.

Script
src/evemtbench/baselines/runner.py
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_location_local --protocol held_out --grid double_line --baseline gru --seeds 0 1 2 3 4
Artifact
outputs/<version>/fault_location_local/held_out/double_line/<baseline>/*/result.json
Runtime
~1.5m fit per seed (compute.fit_seconds_mean)
Hardware
NHR@FAU Alex, 1× A40 · seeds 0–4
Verified by
floor gate · leakage gate
Event detection (ED, binary): learnable in-distribution and scored on the never-trained benchmark family0.911 ± 0.012 test / 0.990 ± 0.003 benchmark (balanced accuracy; best learned baseline gru; floor 0.500)measured

Representative cell: event_detection_global on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.

Script
src/evemtbench/baselines/runner.py
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --task event_detection_global --protocol held_out --grid double_line --baseline gru --seeds 0 1 2 3 4
Artifact
outputs/<version>/event_detection_global/held_out/double_line/<baseline>/*/result.json
Runtime
~8.9m fit per seed (compute.fit_seconds_mean)
Hardware
NHR@FAU Alex, 1× A40 · seeds 0–4
Verified by
floor gate · leakage gate
Event classification (EC, 23-class): learnable in-distribution and scored on the never-trained benchmark family0.761 ± 0.006 test / 0.833 ± 0.016 benchmark (balanced accuracy; best learned baseline resnet; floor 0.062)measured

Representative cell: event_classification_global on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.

Script
src/evemtbench/baselines/runner.py
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --task event_classification_global --protocol held_out --grid double_line --baseline resnet --seeds 0 1 2 3 4
Artifact
outputs/<version>/event_classification_global/held_out/double_line/<baseline>/*/result.json
Runtime
~20.5m fit per seed (compute.fit_seconds_mean)
Hardware
NHR@FAU Alex, 1× A40 · seeds 0–4
Verified by
floor gate · leakage gate
Cross-grid transfer from the multi_grid pretrain corpora (zero-shot vs fine-tune) under the corrected topology-grouped source splits (v1.1)1152 of 1152 transfer cells measured (+ 144/144 pretrain cells)measured

Deep baselines pretrained on the multi_grid template corpus are evaluated zero-shot and fine-tuned on the four reference grids. Transfer protocols are defined only for baselines with transferable weights — classical baselines are N/A by rule, not missing.

Script
hpc/sweep.sh
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol transfer_finetune --grid testgrid_110kv --baseline resnet --seeds 0 1 2 3 4
Artifact
outputs/<version>/*/transfer_*/*/*/*/result.json
Runtime
≤ 24 h wall per grid sweep job
Hardware
NHR@FAU Alex, 1× A40
Verified by
tests/integration/test_transfer_protocols.py · floor gate
Protocol-driven benchmark: every result is conditional evidence, never one scoreevery value keyed by (task × observability × grid × protocol × test set × metric)enforced

The dataset (waveforms, labels, metadata) is separated from the protocol that fixes how it may be used. RESULTS.md, the leaderboard and the paper tables all render per-cell values; no aggregate across tasks or grids exists anywhere in the pipeline.

Script
hpc/make_results.py
Command
PYTHONPATH=src python hpc/make_results.py
Artifact
RESULTS.md
Runtime
seconds (pure read of outputs/)
Hardware
any
Verified by
freshness gate
Five tasks with frozen label derivations; one scoring rule per window24 registered task ids: 12 objective functions × nested observabilityenforced

ED, EC, FD, FC, FL (plus fault-attribute and switching tasks) each carry a frozen label derivation and valid-sample mask, so a window is scored by exactly one rule per task. Changing a label derivation is a major-version event.

Script
src/evemtbench/tasks/registry.py
Command
PYTHONPATH=src python -c "from evemtbench.tasks.registry import BENCHMARK_TASKS; print(len(BENCHMARK_TASKS))"
Artifact
src/evemtbench/tasks/labels/
Runtime
—
Hardware
—
Verified by
tests/integration/test_registry_coherence.py · tests/integration/test_new_tasks_load.py
Three nested observability views on four reference grids spanning 20–345 kVlocal (6 ch) ⊂ line (12 ch) ⊂ global (per-grid width) × 4 gridsenforced

Single-terminal local, two-terminal line-differential and grid-wide global views are evaluated on double_line (110 kV), cigre_mv (20 kV), testgrid_110kv (110 kV) and ieee39 (345 kV, 60 Hz), plus the template_20kv + template_110kv pretrain corpora.

Script
src/evemtbench/data/roots.py
Command
PYTHONPATH=src python -c "from evemtbench.data.roots import GRIDS; print(list(GRIDS))"
Artifact
src/evemtbench/data/roots.py
Runtime
—
Hardware
—
Verified by
tests/integration/test_line_view.py · leakage gate
Episode-disjoint, family-pure 70/15/15 committed splits with leakage controlscommitted split files for 4 grids × 4 partitions, re-audited on every CI runenforced

Splits are stratified at episode level and committed as plain text, so every run reads exactly the same partitions. The leakage tests re-derive and re-audit them; the splits auditor is itself tested to catch planted leaks.

Script
src/evemtbench/splits/loader.py
Command
pytest tests/integration/test_leakage.py tests/integration/test_splits_audit.py -q
Artifact
src/evemtbench/splits/v1.0/ · src/evemtbench/splits/v1.1/
Runtime
seconds
Hardware
any (CI re-runs it)
Verified by
leakage gate · splits_audit gate
Transparent, beatable baseline suite from trivial floor to deep learning7 baselines: majority · threshold · random_forest · mlp · gru · cnn · resnetenforced

The majority floor anchors chance level analytically; the overcurrent threshold represents conventional protection practice; RF and four deep architectures span the learned spectrum. Splits, leakage checks and evaluation scripts are released with the repo.

Script
src/evemtbench/baselines/
Command
PYTHONPATH=src python -m evemtbench.baselines.runner --help
Artifact
src/evemtbench/baselines/
Runtime
—
Hardware
—
Verified by
floor gate · tests/integration/test_runner_cli.py
The benchmark family is scored but never fitted or model-selected on1876 of 1876 run records carry the no-tuning attestationmeasured

Every run record commits fit_on / model_selection_on / held_out_untouched plus a plain-language statement. The audit page is a production reader of these records: the count on the left is computed from them at build time.

Script
src/evemtbench/baselines/runner.py
Command
jq .attestation outputs/<version>/<task>/<protocol>/<grid>/<baseline>/run_record.json
Artifact
outputs/<version>/*/*/*/*/run_record.json
Runtime
—
Hardware
—
Verified by
tests/integration/test_runner_cli.py
Analytic floor self-check over the whole matrix564 comparisons across 96 (task × grid) groups: 2 violations, 0 skipped (pending)measured

The majority baseline must land on the closed-form 1/k floor and no learned baseline may fall below it on the test split. A violation would indicate a broken label mapping or metric wiring — the benchmark's known-answer test.

Script
src/evemtbench/coverage.py
Command
PYTHONPATH=src python -m evemtbench.coverage --check-floor
Artifact
outputs/<version>/ (read-only check)
Runtime
seconds
Hardware
any machine with outputs/
Verified by
floor gate
Rendered views are byte-fresh: RESULTS.md, leaderboard.html and this audit sitepre-commit gate + CI-tested builders; one normalized volatile lineenforced

All rendered views are pure functions of outputs/. The pre-commit hook runs the freshness gate, so a commit with stale numbers is rejected; this row is self-referential — the page you are reading is covered by the gate it describes.

Script
hpc/check_results_fresh.py
Command
python hpc/check_results_fresh.py
Artifact
RESULTS.md · leaderboard.html · audit/ · coverage_snapshot.json
Runtime
seconds
Hardware
loop host (skips where outputs/ is absent)
Verified by
freshness gate · tests/integration/test_check_results_fresh.py
Frozen window contract: 50 ms windows, 5 ms stride, 480 samples @ 9.6 kHzcontract v1.0.0, SHA-256 76a8f6eb814c…enforced

The window geometry and channel order are vendored with a pinned SHA-256; each windowed input file carries a sidecar manifest whose feature hash is verified at load time. Changing the contract is a major-version event.

Script
src/evemtbench/contract/__init__.py
Command
pytest tests/integration/test_contract_conformance.py -q
Artifact
src/evemtbench/contract/contract.json
Runtime
seconds
Hardware
any
Verified by
contract_pin gate · hpc/check_contract_sync.py
Every run emits a machine-readable provenance record1876 run records · 146 distinct git SHAs · 1 lockfile hash(es) · 21 recorded hostsmeasured

run_record.json per cell: git commit, dependency-lockfile SHA-256, library versions, seeds, per-seed compute, checkpoint SHA-256s and the no-tuning attestation; newer records add the executing host, SLURM job id and input-data hashes. The counts here are read from the records at build time.

Script
src/evemtbench/baselines/runner.py
Command
jq '.git_sha,.lockfile_sha256,.host' outputs/<version>/<task>/<protocol>/<grid>/<baseline>/run_record.json
Artifact
outputs/<version>/*/*/*/*/run_record.json
Runtime
written at the end of each cell run
Hardware
the machine that ran the cell (self-recorded)
Verified by
integrity gate · tests/integration/test_checkpoints_telemetry.py
Paper result tables are machine-generated from result.json (a live gated view)5 main-text fragments + 15 supplementary tables, all rendered from outputs/enforced

hpc/make_paper_tables.py and hpc/make_supplement_tables.py render every numeric table cell from result.json files; the committed fragments are a live gated view (byte-checked and --fix-regenerated like RESULTS.md), with five-state markers (\pending / \notapp / \undefm / \oos) standing in for any cell without a result. The public results bundle (results/<version>/tables/) mirrors the same tables from the same builders, byte-identically.

Script
hpc/make_paper_tables.py
Command
PYTHONPATH=src python hpc/make_paper_tables.py
Artifact
results/
Runtime
seconds
Hardware
any
Verified by
tests/integration/test_paper_tables.py · freshness gate
Coverage: 580 of 580 held_out cells measured580 done · 0 pending · 0 failed (held_out); 1876/1876 across all four protocolsmeasured

A cell is one (task × protocol × grid × baseline) runner invocation; it counts as done only when every expected test-set result.json is valid. The full matrix with per-cell status and the exact reproduce command is on the coverage page and in the committed coverage_snapshot.json.

Script
src/evemtbench/coverage.py
Command
PYTHONPATH=src python -m evemtbench.coverage --summary
Artifact
coverage_snapshot.json
Runtime
seconds
Hardware
any machine with outputs/
Verified by
freshness gate · tests/unit/test_coverage.py
The scoring harness is self-testing — dataless controls + a per-run oracle check12 ceiling/floor/alignment controls (pip-install only) + a positive control before every persisted resultenforced

evemtbench-evaluate --self-test proves the evaluator scores known inputs correctly with no data: perfect predictions hit the metric ceiling, constant predictions land exactly on the analytic 1/k floor, malformed prediction frames raise. Every sweep run additionally feeds the TRUE labels through the same evaluate() path and refuses to write results unless they score the ceiling (run_record.harness_control).

Script
src/evemtbench/evaluation/selftest.py
Command
evemtbench-evaluate --self-test
Artifact
outputs/<version>/*/*/*/*/run_record.json
Runtime
~2 s (self-test) · one extra evaluate per test set (runner control)
Hardware
any (self-test needs no data root)
Verified by
tests/unit/test_eval_selftest.py · tests/integration/test_runner_cli.py::test_broken_harness_aborts_the_cell
Headline results survive alternative aggregationper-task re-aggregation under 4 schemes + seeded bootstrap CIs — see the robustness pagemeasured

The same committed per-seed numbers, never re-measured, re-aggregated under equal-weight-by-grid mean, median, leave-one-grid-out and a seeded percentile bootstrap; tasks whose best baseline changes under any scheme are flagged rank-sensitive on audit/robustness.html rather than hidden.

Script
hpc/make_audit.py
Command
PYTHONPATH=src python hpc/make_audit.py
Artifact
audit/robustness.html
Runtime
seconds (pure read of outputs/)
Hardware
any machine with outputs/
Verified by
freshness gate · tests/unit/test_significance.py
Topology-aware GNN baselinedeferreddeferred

A graph baseline that consumes the grid topology is planned but was not run for this version — listed for honesty, not dressed up as a result. Design notes: docs/design/cross_topology.md.

Script
—
Command
—
Artifact
—
Runtime
—
Hardware
—
Verified by
—
Cross-topology protocol for topology-agnostic modelsdeferreddeferred

Training on one set of grids and evaluating on an unseen topology is designed but not part of this version's scored matrix. Design notes: docs/design/cross_topology.md.

Script
—
Command
—
Artifact
—
Runtime
—
Hardware
—
Verified by
—
Open-challenge tracks: robustness, selectivity, cross-cluster, pretrain-transferdesigned, not builtdeferred

Four extension tracks are specified as design documents only (docs/design/robustness.md, selectivity.md, pretrain_transfer.md; the cross-cluster federation design lives in the development repo). No numbers exist and none are claimed.

Script
—
Command
—
Artifact
docs/design/
Runtime
—
Hardware
—
Verified by
—
ieee39 adapt_grid dataset — finalfinalnote

The ieee39 adapt_grid dataset refresh landed in the v1.0.0 re-simulation; all grids' results (ieee39 included) are final for this version.

Script
—
Command
—
Artifact
—
Runtime
—
Hardware
—
Verified by
—
First reference evaluation, released as an extensible reproducible foundationbaselines + splits + leakage checks + evaluation scripts, versioned and gatedenforced

The reference results are rendered from committed result.json by the gated builders and traced on the audit dashboard; the released splits, leakage checks, baseline implementations, and evaluation scripts let others reproduce them and extend the benchmark with new methods, grids, and a future challenge set.

Script
hpc/make_results_bundle.py
Command
PYTHONPATH=src python hpc/make_results_bundle.py
Artifact
results/
Runtime
seconds
Hardware
any
Verified by
tests/integration/test_results_bundle.py · tests/integration/test_paper_numbers_covered.py