One row per claim of the paper. Every number below is a placeholder filled from the committed artifacts at build time — the manifest source contains no digits (see hpc/audit_manifest.py). Rows whose cells are still being swept show in progress with live counts; planned work is deferred, never a result.
Fault detection (FD, binary): learnable in-distribution and scored on the never-trained benchmark family0.883 ± 0.030 test / 0.989 ± 0.008 benchmark (balanced accuracy; best learned baseline resnet; floor 0.500)measured
Representative cell: fault_detection_local on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.
Fault-type classification (FC, 9-class): learnable in-distribution and scored on the never-trained benchmark family0.855 ± 0.007 test / 0.896 ± 0.045 benchmark (balanced accuracy; best learned baseline resnet; floor 0.143)measured
Representative cell: fault_classification_local on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.
Fault localisation (FL, regression): learnable in-distribution and scored on the never-trained benchmark family18.600 ± 0.213 test / 11.695 ± 0.530 benchmark (MAE, lower better; best learned baseline gru; floor 24.416)measured
Representative cell: fault_location_local on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.
Event detection (ED, binary): learnable in-distribution and scored on the never-trained benchmark family0.911 ± 0.012 test / 0.990 ± 0.003 benchmark (balanced accuracy; best learned baseline gru; floor 0.500)measured
Representative cell: event_detection_global on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.
Event classification (EC, 23-class): learnable in-distribution and scored on the never-trained benchmark family0.761 ± 0.006 test / 0.833 ± 0.016 benchmark (balanced accuracy; best learned baseline resnet; floor 0.062)measured
Representative cell: event_classification_global on double_line under the held_out protocol, 5 seeds, mean ± 95% CI. Trained on the adapt_grid train split, model-selected on val; the benchmark family was never fitted or tuned on (attestation in every run record). Every grid × view combination is in the coverage matrix and the per-task page.
Cross-grid transfer from the multi_grid pretrain corpora (zero-shot vs fine-tune) under the corrected topology-grouped source splits (v1.1)1152 of 1152 transfer cells measured (+ 144/144 pretrain cells)measured
Deep baselines pretrained on the multi_grid template corpus are evaluated zero-shot and fine-tuned on the four reference grids. Transfer protocols are defined only for baselines with transferable weights — classical baselines are N/A by rule, not missing.
Protocol-driven benchmark: every result is conditional evidence, never one scoreevery value keyed by (task × observability × grid × protocol × test set × metric)enforced
The dataset (waveforms, labels, metadata) is separated from the protocol that fixes how it may be used. RESULTS.md, the leaderboard and the paper tables all render per-cell values; no aggregate across tasks or grids exists anywhere in the pipeline.
Five tasks with frozen label derivations; one scoring rule per window24 registered task ids: 12 objective functions × nested observabilityenforced
ED, EC, FD, FC, FL (plus fault-attribute and switching tasks) each carry a frozen label derivation and valid-sample mask, so a window is scored by exactly one rule per task. Changing a label derivation is a major-version event.
Three nested observability views on four reference grids spanning 20–345 kVlocal (6 ch) ⊂ line (12 ch) ⊂ global (per-grid width) × 4 gridsenforced
Single-terminal local, two-terminal line-differential and grid-wide global views are evaluated on double_line (110 kV), cigre_mv (20 kV), testgrid_110kv (110 kV) and ieee39 (345 kV, 60 Hz), plus the template_20kv + template_110kv pretrain corpora.
Episode-disjoint, family-pure 70/15/15 committed splits with leakage controlscommitted split files for 4 grids × 4 partitions, re-audited on every CI runenforced
Splits are stratified at episode level and committed as plain text, so every run reads exactly the same partitions. The leakage tests re-derive and re-audit them; the splits auditor is itself tested to catch planted leaks.
Transparent, beatable baseline suite from trivial floor to deep learning7 baselines: majority · threshold · random_forest · mlp · gru · cnn · resnetenforced
The majority floor anchors chance level analytically; the overcurrent threshold represents conventional protection practice; RF and four deep architectures span the learned spectrum. Splits, leakage checks and evaluation scripts are released with the repo.
The benchmark family is scored but never fitted or model-selected on1876 of 1876 run records carry the no-tuning attestationmeasured
Every run record commits fit_on / model_selection_on / held_out_untouched plus a plain-language statement. The audit page is a production reader of these records: the count on the left is computed from them at build time.
Analytic floor self-check over the whole matrix564 comparisons across 96 (task × grid) groups: 2 violations, 0 skipped (pending)measured
The majority baseline must land on the closed-form 1/k floor and no learned baseline may fall below it on the test split. A violation would indicate a broken label mapping or metric wiring — the benchmark's known-answer test.
Rendered views are byte-fresh: RESULTS.md, leaderboard.html and this audit sitepre-commit gate + CI-tested builders; one normalized volatile lineenforced
All rendered views are pure functions of outputs/. The pre-commit hook runs the freshness gate, so a commit with stale numbers is rejected; this row is self-referential — the page you are reading is covered by the gate it describes.
Frozen window contract: 50 ms windows, 5 ms stride, 480 samples @ 9.6 kHzcontract v1.0.0, SHA-256 76a8f6eb814c…enforced
The window geometry and channel order are vendored with a pinned SHA-256; each windowed input file carries a sidecar manifest whose feature hash is verified at load time. Changing the contract is a major-version event.
Every run emits a machine-readable provenance record1876 run records · 146 distinct git SHAs · 1 lockfile hash(es) · 21 recorded hostsmeasured
run_record.json per cell: git commit, dependency-lockfile SHA-256, library versions, seeds, per-seed compute, checkpoint SHA-256s and the no-tuning attestation; newer records add the executing host, SLURM job id and input-data hashes. The counts here are read from the records at build time.
Paper result tables are machine-generated from result.json (a live gated view)5 main-text fragments + 15 supplementary tables, all rendered from outputs/enforced
hpc/make_paper_tables.py and hpc/make_supplement_tables.py render every numeric table cell from result.json files; the committed fragments are a live gated view (byte-checked and --fix-regenerated like RESULTS.md), with five-state markers (\pending / \notapp / \undefm / \oos) standing in for any cell without a result. The public results bundle (results/<version>/tables/) mirrors the same tables from the same builders, byte-identically.
Coverage: 580 of 580 held_out cells measured580 done · 0 pending · 0 failed (held_out); 1876/1876 across all four protocolsmeasured
A cell is one (task × protocol × grid × baseline) runner invocation; it counts as done only when every expected test-set result.json is valid. The full matrix with per-cell status and the exact reproduce command is on the coverage page and in the committed coverage_snapshot.json.
The scoring harness is self-testing — dataless controls + a per-run oracle check12 ceiling/floor/alignment controls (pip-install only) + a positive control before every persisted resultenforced
evemtbench-evaluate --self-test proves the evaluator scores known inputs correctly with no data: perfect predictions hit the metric ceiling, constant predictions land exactly on the analytic 1/k floor, malformed prediction frames raise. Every sweep run additionally feeds the TRUE labels through the same evaluate() path and refuses to write results unless they score the ceiling (run_record.harness_control).
Headline results survive alternative aggregationper-task re-aggregation under 4 schemes + seeded bootstrap CIs — see the robustness pagemeasured
The same committed per-seed numbers, never re-measured, re-aggregated under equal-weight-by-grid mean, median, leave-one-grid-out and a seeded percentile bootstrap; tasks whose best baseline changes under any scheme are flagged rank-sensitive on audit/robustness.html rather than hidden.
A graph baseline that consumes the grid topology is planned but was not run for this version — listed for honesty, not dressed up as a result. Design notes: docs/design/cross_topology.md.
Script
—
Command
—
Artifact
—
Runtime
—
Hardware
—
Verified by
—
Cross-topology protocol for topology-agnostic modelsdeferreddeferred
Training on one set of grids and evaluating on an unseen topology is designed but not part of this version's scored matrix. Design notes: docs/design/cross_topology.md.
Script
—
Command
—
Artifact
—
Runtime
—
Hardware
—
Verified by
—
Open-challenge tracks: robustness, selectivity, cross-cluster, pretrain-transferdesigned, not builtdeferred
Four extension tracks are specified as design documents only (docs/design/robustness.md, selectivity.md, pretrain_transfer.md; the cross-cluster federation design lives in the development repo). No numbers exist and none are claimed.
The ieee39 adapt_grid dataset refresh landed in the v1.0.0 re-simulation; all grids' results (ieee39 included) are final for this version.
Script
—
Command
—
Artifact
—
Runtime
—
Hardware
—
Verified by
—
First reference evaluation, released as an extensible reproducible foundationbaselines + splits + leakage checks + evaluation scripts, versioned and gatedenforced
The reference results are rendered from committed result.json by the gated builders and traced on the audit dashboard; the released splits, leakage checks, baseline implementations, and evaluation scripts let others reproduce them and extend the benchmark with new methods, grids, and a future challenge set.