Verification gates — code tour

benchmark version 1.1.0 · held_out results v1.0.0 · splits v1.0/held_out, v1.1/multi_grid
updated 2026-09-29 13:18 UTC · git 7243f310ddbc

A reviewer should be able to confirm what each gate asserts without reading the whole codebase. The two diagrams show where each gate sits in the pipeline; each card then walks the gate's real call stack — every step linked at the exact (build-time-derived) source line — down to the single assertion that matters.

The data path

PowerFactory EMT simulation            (producer; labels exact by construction)
   |  writes X_<topology>.raw + y_<topology>.parquet + *.raw.manifest.json (feature_hash)
   v
data/windows.load_windows      -- _verify_manifest: status/topology/geometry  [contract_pin]
   |                              sha256(feature_names) == manifest.feature_hash
   v
splits/loader.read_held_out_splits -> committed v1.0/held_out/<grid>/*.txt   [leakage]
   |  episode-disjoint, family-pure                                          [splits_audit]
   |  (multi_grid_pretrain: v1.1/multi_grid_pretrain/, topology-grouped)
   v
baselines/runner  (fit on train, select on val; benchmark family predict-only)
   v
outputs/<ver>/<task>/<protocol>/<grid>/<baseline>/<test_set>/result.json
        + predictions_seed*.parquet + run_record.json                        [integrity]

The publishing path

outputs/<ver>/**/result.json + run_record.json
   |-- make_results.build_results        -> RESULTS.md
   |-- leaderboard.render_html           -> leaderboard.html                 [floor]
   |-- make_audit.fresh_pages            -> audit/*.html
   '-- coverage.render_snapshot          -> coverage_snapshot.json
   v
committed views  ->  hpc/check_results_fresh.py (byte-compare)               [freshness]
   v
.githooks/pre-commit blocks any drifting commit; CI re-runs the same gates

The six gates

Leakage gate — episode-disjoint, family-pure committed splits

Asserts: No episode contributes windows to more than one partition, and every partition draws from its declared dataset family only.

Where: tests/integration/test_leakage.py:62 — def test_committed_splits_are_disjoint_and_family_pure(grid): (all line numbers derived at build time)

Step in the call stackWhat is actually computed
test_leakage.py:62
def test_committed_splits_are_disjoint_and_family_pure(grid):
Entry point, parametrized over every registered grid; data-free — it reads only the committed, packaged split files.
loader.py:56
path = split_dir(protocol, grid, version=version, splits_root=splits_root) / f"{split}.txt"
Resolves the concrete committed artifact: splits/v<ver>/held_out/<grid>/<split>.txt.
loader.py:62
for line in path.read_text().split():
The actual read — each line is one (family, grid, sim_idx) episode triple; these files are authoritative, never recomputed at run time.
leakage.py:35
raise LeakageError(f"{t} appears in both {seen[t]!r} and {name!r} partitions")
THE core check: raises the moment any episode appears in two partitions (single dict scan over all four partitions).
test_leakage.py:73
assert all(fam == "adapt_grid" for fam, _g, _i in triples[name])
Family purity: train/val/test may draw only from the adapt_grid family (benchmark-family purity asserted separately just below).

On failure: pytest fails naming the grid and the episode id that appears in two partitions (or the window whose family is wrong).

How to read it: The assertion that matters is assert_disjoint_triples() over the committed split files of each grid — the exact files every run reads. CI re-runs this test by name on every push so a leaky split can never merge silently.

Run it: pytest tests/integration/test_leakage.py -q

Splits audit — the validator is itself validated

Asserts: The split auditor detects a planted duplicate episode and a missing held-out class — proving the leakage checks can actually catch what they claim to.

Where: tests/integration/test_splits_audit.py:84 — def test_audit_flags_episode_leakage(): (all line numbers derived at build time)

Step in the call stackWhat is actually computed
test_splits_audit.py:84
def test_audit_flags_episode_leakage():
The mirror gate's entry: proves the auditor is non-vacuous by making it catch a deliberately planted leak.
test_splits_audit.py:88
"val": [0, *range(14, 17)],
Plants the leak — episode 0 is put in val while it already sits in train.
audit.py:52
not (sets["train"] & sets["val"])
The disjointness computation (pairwise set intersections across partitions); the planted duplicate makes this intersection non-empty.
audit.py:80
"ok": disjoint and fractions_ok and coverage_ok,
The report: ok only when disjointness, split fractions AND per-class held-out coverage all hold.
test_splits_audit.py:92
assert not rep["disjoint"] and not rep["ok"]
THE assertion: the planted leak must flip both flags — if this ever passes with a healthy report, the auditor itself regressed.

On failure: pytest fails because a deliberately corrupted split was NOT flagged — i.e. the audit machinery itself regressed.

How to read it: This is the mirror of the leakage gate: it corrupts a split on purpose and asserts the auditor raises. A green run means the leakage gate's 'pass' is meaningful, not vacuous.

Run it: pytest tests/integration/test_splits_audit.py -q

Analytic floor gate — majority ≈ 1/k, learned ≥ floor

Asserts: Per (task × grid): the majority baseline's balanced accuracy equals the analytic 1/k chance floor (tolerance 0.02), and no learned baseline scores below the floor on the test split — direction-aware, pending cells skipped (never a false fail).

Where: src/evemtbench/coverage.py:379 — def check_floor( (all line numbers derived at build time)

Step in the call stackWhat is actually computed
coverage.py:379
def check_floor(
Entry: iterates the coverage cells grouped by (task, protocol, grid) on the in-distribution test set; pending/undefined cells are skipped, never failed.
coverage.py:421
maj_doc = _load_result(cell_dir(out, maj_cell, version=version) / test_set / "result.json")
Reads the majority baseline's committed result.json for the group.
coverage.py:436
floor = 1.0 / k
The analytic chance floor — computable by hand: 1/k over the k classes present in the eval set.
coverage.py:437
if abs(maj - floor) > tol:
Plausibility: the majority baseline's balanced accuracy must sit within tol (0.02) of 1/k, else labels/vocabulary/eval alignment are broken.
coverage.py:457
if direction * (val - maj) < -band:
Dominance: no learned baseline may score below the majority floor — direction-aware, with a relative band for regression scales.

On failure: exit 1 listing each violating (task, grid, baseline) with its value vs the floor — e.g. a mis-mapped label set or an evaluation bug pushing majority off 1/k.

How to read it: A known-answer test over the whole results matrix: the floor is computable by hand (1/k for k classes), so agreement is evidence the metric pipeline, label derivations and aggregation are wired correctly end to end.

Current report: 564 comparisons across 96 groups, 2 violations, 0 skipped (pending cells).

Run it: PYTHONPATH=src python -m evemtbench.coverage --check-floor

Freshness gate — committed views byte-match a rebuild

Asserts: RESULTS.md, leaderboard.html, audit/ and coverage_snapshot.json equal a fresh rebuild from outputs/ byte-for-byte (one volatile 'updated …' line normalized).

Where: hpc/check_results_fresh.py:268 — if _normalize(committed) != _normalize(fresh): (all line numbers derived at build time)

Step in the call stackWhat is actually computed
check_results_fresh.py:70
def _fresh_leaderboard() -> str:
Fresh-build calls: every view is rebuilt from outputs/ in-process, all with generated_at=None so the rebuild is deterministic.
make_audit.py:1914
def fresh_pages(
Builder purity: the single audit-site entry point the gate and the writer share — one code path, so verify and regenerate can never diverge.
leaderboard.py:530
def render_html(
Same for the leaderboard: the caller stamps time, the renderer is pure.
check_results_fresh.py:268
if _normalize(committed) != _normalize(fresh):
THE comparison: committed vs fresh, byte-for-byte, after the single volatile 'updated …' line is normalized away.
check_results_fresh.py:198
if p.name not in audit_fresh and p.name not in _UNGATED
Stray-page check: a committed audit page a rebuild no longer produces is stale too (progress.html is the one documented, history-derived exemption).

On failure: exit 1 with a unified diff of the stale file; the pre-commit hook blocks the commit until `--fix` regenerates it.

How to read it: The one comparison that matters is committed-vs-fresh equality per rendered file. Because the builders are pure functions of outputs/, every number on this site is mechanically traceable to a result.json — hand-editing a published number is impossible without the gate failing.

Run it: python hpc/check_results_fresh.py

Artifact integrity — the outputs corpus itself is checked

Asserts: Every done cell carries the full seed set, a schema-valid provenance record with non-null git/lockfile pins, run_record.seeds == result.seeds, and existing checkpoints; --deep additionally verifies every predictions parquet's row count against n_test (truncated-write detection).

Where: src/evemtbench/integrity.py:162 — def check_outputs( (all line numbers derived at build time)

Step in the call stackWhat is actually computed
integrity.py:162
def check_outputs(
Entry: iterates every applicable cell and aggregates per-cell problems plus the distinct (git sha, lockfile hash) provenance pairs.
integrity.py:181
if cell_status(out, cell, version=version) != "done":
Skip-pending: incomplete cells are counted as skipped, never failed — coverage owns those.
integrity.py:81
if expected_seeds is not None and seeds != expected_seeds:
Seed-set check: every done cell must carry the full production seed set [0..4] in each test set's result.json.
integrity.py:101
for err in validator.iter_errors(record):
run_record.json validates against the committed JSON Schema (baselines/run_record.schema.json).
integrity.py:154
if n_rows != doc.get("n_test"):
--deep: each predictions parquet's metadata row count must equal n_test — catches truncated NFS writes no code test can see.

On failure: exit 1 listing each problem as '<cell>: <what is inconsistent>'; pending cells are skipped, never failed (coverage owns those).

How to read it: CI tests validate the code; this gate validates the artifacts the sweep left behind, which only exist on the loop host. It also reports the distinct (git sha, lockfile hash) pairs across the corpus — the evidence behind the paper's 'attributable to a frozen configuration' sentence. Run at every freeze.

Run it: PYTHONPATH=src python -m evemtbench.integrity

Contract pin — SHA-256-pinned window contract and inputs

Asserts: The vendored window contract matches its pinned SHA-256 and version; every run record pins the dependency lockfile hash; every windowed input carries a sidecar manifest whose feature hash is verified at load time.

Where: src/evemtbench/contract/__init__.py:51 — def verify_pin() -> list[str]: (all line numbers derived at build time)

Step in the call stackWhat is actually computed
__init__.py:51
def verify_pin() -> list[str]:
Offline pin entry — returns a list of problems (empty = OK).
__init__.py:61
if actual != CONTRACT_SHA256:
The SHA compare: live sha256(contract.json) must equal the pinned constant — the window-format contract cannot drift unreviewed.
windows.py:156
if manifest.get("status") != "complete":
Load-time manifest gate: status, topology and window geometry are all verified before a single byte of waveform data is mapped.
windows.py:252
if h.hexdigest() != fhash:
Feature-hash consistency: sha256 over the channel-name list must equal the producer manifest's feature_hash — channel order cannot silently shift.
runner.py:193
def _results_env_hash(version: str) -> str | None:
The environment pin: sha256(env/results-v<version>.lock) — the captured conda prefix incl. torch + CUDA — recorded into every new run_record.json. The older lockfile_sha256 pins the CI dependency set, which contains no torch and so never described the results environment.

On failure: verify_pin() returns a non-empty problem list (version or SHA mismatch) and the contract-conformance tests fail.

How to read it: Three pins close the input side: contract SHA (window geometry can't drift), the per-file feature hash checked when windows are memory-mapped (channel order can't drift), and the environment. Read the environment pin precisely: a lockfile SHA in a run record proves only which file sat in the repo, NOT which interpreter ran — so the real check is that each record's independently captured numpy/pandas/scikit-learn/scipy/torch versions agree with the captured results environment (env/results-v<version>.lock, verified by hpc/verify_results_env.py). Two independent captures agreeing is the guarantee; the hash alone is a pointer.

Run it: pytest tests/integration/test_contract_conformance.py -q