Where the ground truth comes from. Every window's label derives from the PowerFactory EMT simulation script that injected the event — the fault type, location, phases or switching action are simulation inputs, not annotations. Labels are therefore exact by construction: there is no annotation noise and no external label source to verify. What CAN go wrong — and what the six gates exist to catch — is leakage between partitions, contamination of the held-out sets, published tables drifting from artifacts, and class-imbalance artefacts.
The three dataset families
adapt_grid — the in-distribution familySimulated at each grid's nominal operating point. Split 70/15/15 into train / val / test at the episode level: every window of one simulated event stays in one partition, so a model can never see a different slice of an evaluation episode during training.
benchmark — the never-trained familyA second simulation campaign at a shifted operating point. It is scored for every cell but never fitted or model-selected on — every run record carries a no-tuning attestation. Generalisation here is the benchmark's headline claim.
multi_grid — the pre-training corpusMany synthetic 20 kV topologies in a portable 12-channel line view; the source domain for the transfer protocols (pretrain → zero-shot / fine-tune on unseen grids).
How results are graded
the metric set is a frozen contract (src/evemtbench/evaluation/metrics.py); changing it is a major version bump — results across versions are never comparable
Balanced accuracy / macro-F1Classification headline: every class weighs equally, so the majority class cannot buy a good score — and the majority baseline must land exactly on the analytic 1/k floor, which the floor gate checks over the whole matrix.
Missed-fault & false-alarm rateProtection-specific asymmetry for binary tasks: a missed fault (FN) and a needless trip (FP) have very different real-world costs, so both are always reported next to the headline.
MAE / RMSE / R²Fault-localisation regression, in the grid's own distance units; the majority analogue (train-mean predictor) is the floor.
The grids
4 reference grids spanning 20–345 kV plus the pre-training corpus; class balances read from the committed split metadata (src/evemtbench/splits/v1.0)