A reviewer should be able to confirm what each gate asserts without reading the whole codebase. The two diagrams show where each gate sits in the pipeline; each card then walks the gate's real call stack — every step linked at the exact (build-time-derived) source line — down to the single assertion that matters.
PowerFactory EMT simulation (producer; labels exact by construction)
| writes X_<topology>.raw + y_<topology>.parquet + *.raw.manifest.json (feature_hash)
v
data/windows.load_windows -- _verify_manifest: status/topology/geometry [contract_pin]
| sha256(feature_names) == manifest.feature_hash
v
splits/loader.read_held_out_splits -> committed v1.0/held_out/<grid>/*.txt [leakage]
| episode-disjoint, family-pure [splits_audit]
| (multi_grid_pretrain: v1.1/multi_grid_pretrain/, topology-grouped)
v
baselines/runner (fit on train, select on val; benchmark family predict-only)
v
outputs/<ver>/<task>/<protocol>/<grid>/<baseline>/<test_set>/result.json
+ predictions_seed*.parquet + run_record.json [integrity]outputs/<ver>/**/result.json + run_record.json
|-- make_results.build_results -> RESULTS.md
|-- leaderboard.render_html -> leaderboard.html [floor]
|-- make_audit.fresh_pages -> audit/*.html
'-- coverage.render_snapshot -> coverage_snapshot.json
v
committed views -> hpc/check_results_fresh.py (byte-compare) [freshness]
v
.githooks/pre-commit blocks any drifting commit; CI re-runs the same gatesAsserts: No episode contributes windows to more than one partition, and every partition draws from its declared dataset family only.
Where: tests/integration/test_leakage.py:62 — def test_committed_splits_are_disjoint_and_family_pure(grid): (all line numbers derived at build time)
| Step in the call stack | What is actually computed |
|---|---|
test_leakage.py:62def test_committed_splits_are_disjoint_and_family_pure(grid): | Entry point, parametrized over every registered grid; data-free — it reads only the committed, packaged split files. |
loader.py:56path = split_dir(protocol, grid, version=version, splits_root=splits_root) / f"{split}.txt" | Resolves the concrete committed artifact: splits/v<ver>/held_out/<grid>/<split>.txt. |
loader.py:62for line in path.read_text().split(): | The actual read — each line is one (family, grid, sim_idx) episode triple; these files are authoritative, never recomputed at run time. |
leakage.py:35raise LeakageError(f"{t} appears in both {seen[t]!r} and {name!r} partitions") | THE core check: raises the moment any episode appears in two partitions (single dict scan over all four partitions). |
test_leakage.py:73assert all(fam == "adapt_grid" for fam, _g, _i in triples[name]) | Family purity: train/val/test may draw only from the adapt_grid family (benchmark-family purity asserted separately just below). |
On failure: pytest fails naming the grid and the episode id that appears in two partitions (or the window whose family is wrong).
How to read it: The assertion that matters is assert_disjoint_triples() over the committed split files of each grid — the exact files every run reads. CI re-runs this test by name on every push so a leaky split can never merge silently.
Run it: pytest tests/integration/test_leakage.py -q
Asserts: The split auditor detects a planted duplicate episode and a missing held-out class — proving the leakage checks can actually catch what they claim to.
Where: tests/integration/test_splits_audit.py:84 — def test_audit_flags_episode_leakage(): (all line numbers derived at build time)
| Step in the call stack | What is actually computed |
|---|---|
test_splits_audit.py:84def test_audit_flags_episode_leakage(): | The mirror gate's entry: proves the auditor is non-vacuous by making it catch a deliberately planted leak. |
test_splits_audit.py:88"val": [0, *range(14, 17)], | Plants the leak — episode 0 is put in val while it already sits in train. |
audit.py:52not (sets["train"] & sets["val"]) | The disjointness computation (pairwise set intersections across partitions); the planted duplicate makes this intersection non-empty. |
audit.py:80"ok": disjoint and fractions_ok and coverage_ok, | The report: ok only when disjointness, split fractions AND per-class held-out coverage all hold. |
test_splits_audit.py:92assert not rep["disjoint"] and not rep["ok"] | THE assertion: the planted leak must flip both flags — if this ever passes with a healthy report, the auditor itself regressed. |
On failure: pytest fails because a deliberately corrupted split was NOT flagged — i.e. the audit machinery itself regressed.
How to read it: This is the mirror of the leakage gate: it corrupts a split on purpose and asserts the auditor raises. A green run means the leakage gate's 'pass' is meaningful, not vacuous.
Run it: pytest tests/integration/test_splits_audit.py -q
Asserts: Per (task × grid): the majority baseline's balanced accuracy equals the analytic 1/k chance floor (tolerance 0.02), and no learned baseline scores below the floor on the test split — direction-aware, pending cells skipped (never a false fail).
Where: src/evemtbench/coverage.py:379 — def check_floor( (all line numbers derived at build time)
| Step in the call stack | What is actually computed |
|---|---|
coverage.py:379def check_floor( | Entry: iterates the coverage cells grouped by (task, protocol, grid) on the in-distribution test set; pending/undefined cells are skipped, never failed. |
coverage.py:421maj_doc = _load_result(cell_dir(out, maj_cell, version=version) / test_set / "result.json") | Reads the majority baseline's committed result.json for the group. |
coverage.py:436floor = 1.0 / k | The analytic chance floor — computable by hand: 1/k over the k classes present in the eval set. |
coverage.py:437if abs(maj - floor) > tol: | Plausibility: the majority baseline's balanced accuracy must sit within tol (0.02) of 1/k, else labels/vocabulary/eval alignment are broken. |
coverage.py:457if direction * (val - maj) < -band: | Dominance: no learned baseline may score below the majority floor — direction-aware, with a relative band for regression scales. |
On failure: exit 1 listing each violating (task, grid, baseline) with its value vs the floor — e.g. a mis-mapped label set or an evaluation bug pushing majority off 1/k.
How to read it: A known-answer test over the whole results matrix: the floor is computable by hand (1/k for k classes), so agreement is evidence the metric pipeline, label derivations and aggregation are wired correctly end to end.
Current report: 564 comparisons across 96 groups, 2 violations, 0 skipped (pending cells).
Run it: PYTHONPATH=src python -m evemtbench.coverage --check-floor
Asserts: RESULTS.md, leaderboard.html, audit/ and coverage_snapshot.json equal a fresh rebuild from outputs/ byte-for-byte (one volatile 'updated …' line normalized).
Where: hpc/check_results_fresh.py:268 — if _normalize(committed) != _normalize(fresh): (all line numbers derived at build time)
| Step in the call stack | What is actually computed |
|---|---|
check_results_fresh.py:70def _fresh_leaderboard() -> str: | Fresh-build calls: every view is rebuilt from outputs/ in-process, all with generated_at=None so the rebuild is deterministic. |
make_audit.py:1914def fresh_pages( | Builder purity: the single audit-site entry point the gate and the writer share — one code path, so verify and regenerate can never diverge. |
leaderboard.py:530def render_html( | Same for the leaderboard: the caller stamps time, the renderer is pure. |
check_results_fresh.py:268if _normalize(committed) != _normalize(fresh): | THE comparison: committed vs fresh, byte-for-byte, after the single volatile 'updated …' line is normalized away. |
check_results_fresh.py:198if p.name not in audit_fresh and p.name not in _UNGATED | Stray-page check: a committed audit page a rebuild no longer produces is stale too (progress.html is the one documented, history-derived exemption). |
On failure: exit 1 with a unified diff of the stale file; the pre-commit hook blocks the commit until `--fix` regenerates it.
How to read it: The one comparison that matters is committed-vs-fresh equality per rendered file. Because the builders are pure functions of outputs/, every number on this site is mechanically traceable to a result.json — hand-editing a published number is impossible without the gate failing.
Run it: python hpc/check_results_fresh.py
Asserts: Every done cell carries the full seed set, a schema-valid provenance record with non-null git/lockfile pins, run_record.seeds == result.seeds, and existing checkpoints; --deep additionally verifies every predictions parquet's row count against n_test (truncated-write detection).
Where: src/evemtbench/integrity.py:162 — def check_outputs( (all line numbers derived at build time)
| Step in the call stack | What is actually computed |
|---|---|
integrity.py:162def check_outputs( | Entry: iterates every applicable cell and aggregates per-cell problems plus the distinct (git sha, lockfile hash) provenance pairs. |
integrity.py:181if cell_status(out, cell, version=version) != "done": | Skip-pending: incomplete cells are counted as skipped, never failed — coverage owns those. |
integrity.py:81if expected_seeds is not None and seeds != expected_seeds: | Seed-set check: every done cell must carry the full production seed set [0..4] in each test set's result.json. |
integrity.py:101for err in validator.iter_errors(record): | run_record.json validates against the committed JSON Schema (baselines/run_record.schema.json). |
integrity.py:154if n_rows != doc.get("n_test"): | --deep: each predictions parquet's metadata row count must equal n_test — catches truncated NFS writes no code test can see. |
On failure: exit 1 listing each problem as '<cell>: <what is inconsistent>'; pending cells are skipped, never failed (coverage owns those).
How to read it: CI tests validate the code; this gate validates the artifacts the sweep left behind, which only exist on the loop host. It also reports the distinct (git sha, lockfile hash) pairs across the corpus — the evidence behind the paper's 'attributable to a frozen configuration' sentence. Run at every freeze.
Run it: PYTHONPATH=src python -m evemtbench.integrity
Asserts: The vendored window contract matches its pinned SHA-256 and version; every run record pins the dependency lockfile hash; every windowed input carries a sidecar manifest whose feature hash is verified at load time.
Where: src/evemtbench/contract/__init__.py:51 — def verify_pin() -> list[str]: (all line numbers derived at build time)
| Step in the call stack | What is actually computed |
|---|---|
__init__.py:51def verify_pin() -> list[str]: | Offline pin entry — returns a list of problems (empty = OK). |
__init__.py:61if actual != CONTRACT_SHA256: | The SHA compare: live sha256(contract.json) must equal the pinned constant — the window-format contract cannot drift unreviewed. |
windows.py:156if manifest.get("status") != "complete": | Load-time manifest gate: status, topology and window geometry are all verified before a single byte of waveform data is mapped. |
windows.py:252if h.hexdigest() != fhash: | Feature-hash consistency: sha256 over the channel-name list must equal the producer manifest's feature_hash — channel order cannot silently shift. |
runner.py:193def _results_env_hash(version: str) -> str | None: | The environment pin: sha256(env/results-v<version>.lock) — the captured conda prefix incl. torch + CUDA — recorded into every new run_record.json. The older lockfile_sha256 pins the CI dependency set, which contains no torch and so never described the results environment. |
On failure: verify_pin() returns a non-empty problem list (version or SHA mismatch) and the contract-conformance tests fail.
How to read it: Three pins close the input side: contract SHA (window geometry can't drift), the per-file feature hash checked when windows are memory-mapped (channel order can't drift), and the environment. Read the environment pin precisely: a lockfile SHA in a run record proves only which file sat in the repo, NOT which interpreter ran — so the real check is that each record's independently captured numpy/pandas/scikit-learn/scipy/torch versions agree with the captured results environment (env/results-v<version>.lock, verified by hpc/verify_results_env.py). Two independent captures agreeing is the guarantee; the hash alone is a pointer.
Run it: pytest tests/integration/test_contract_conformance.py -q