Task — switch_detection_global

benchmark version 1.1.0 · held_out results v1.0.0 · splits v1.0/held_out, v1.1/multi_grid
updated 2026-09-29 13:18 UTC · git 7243f310ddbc

tier 1 · switch_detection · global view · 2 classes · all cells measured

What this task is. Decide whether the window contains a switching action (breaker or disconnector operating). Switching transients look superficially like faults — this task isolates that confusion pair. Global view: every cubicle in the grid at once (width differs per grid), the wide-area upper bound on observability.

How it's scored

Headline metric balanced_accuracy (↑ higher is better); full metric set: balanced_accuracy · macro_f1 · missed_event_rate · false_event_trip_rate · accuracy. Metrics are a frozen contract (src/evemtbench/evaluation/metrics.py); label derivation: src/evemtbench/tasks/labels.

Example window

example waveform window

event switch_cap_on · grid double_line (adapt_grid/test split) · sample 181, window #4163, cubicle 3 (max-RMS) · 50 ms / 480 samples @ 9.6 kHz · top: 3-phase current, bottom: 3-phase voltage · figure provenance in assets/manifest.json

Results (held_out)

switch_detection_global — every grid × baseline; click headers to sort
taskgridbaselineheadlinetestbenchmarkparamsfit/seedtrace
switch_detection_globalcigre_mvmajoritybalanced_accuracy ↑0.500 ± 0.000 (n=5)0.500 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_detection_globalcigre_mvrandom_forestbalanced_accuracy ↑0.754 ± 0.002 (n=5)0.816 ± 0.005 (n=5)—6.7mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalcigre_mvmlpbalanced_accuracy ↑0.688 ± 0.002 (n=5)0.710 ± 0.001 (n=5)42.9M36.6mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalcigre_mvgrubalanced_accuracy ↑0.718 ± 0.127 (n=5)0.758 ± 0.102 (n=5)117k28.3mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalcigre_mvcnnbalanced_accuracy ↑0.781 ± 0.012 (n=5)0.822 ± 0.014 (n=5)124k37.9mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalcigre_mvresnetbalanced_accuracy ↑0.865 ± 0.011 (n=5)0.914 ± 0.014 (n=5)602k37.2mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaldouble_linemajoritybalanced_accuracy ↑0.500 ± 0.000 (n=5)0.500 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_detection_globaldouble_linerandom_forestbalanced_accuracy ↑0.721 ± 0.001 (n=5)0.630 ± 0.016 (n=5)—1.5mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaldouble_linemlpbalanced_accuracy ↑0.701 ± 0.000 (n=5)0.730 ± 0.002 (n=5)12.0M10.3mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaldouble_linegrubalanced_accuracy ↑0.769 ± 0.029 (n=5)0.826 ± 0.032 (n=5)69k9.3mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaldouble_linecnnbalanced_accuracy ↑0.729 ± 0.025 (n=5)0.816 ± 0.017 (n=5)68k10.6mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaldouble_lineresnetbalanced_accuracy ↑0.780 ± 0.004 (n=5)0.808 ± 0.017 (n=5)530k26.4mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaltestgrid_110kvmajoritybalanced_accuracy ↑0.500 ± 0.000 (n=5)0.500 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_detection_globaltestgrid_110kvrandom_forestbalanced_accuracy ↑0.682 ± 0.002 (n=5)0.592 ± 0.006 (n=5)—2.5mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaltestgrid_110kvmlpbalanced_accuracy ↑0.711 ± 0.001 (n=5)0.734 ± 0.003 (n=5)26.7M19.6mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaltestgrid_110kvgrubalanced_accuracy ↑0.725 ± 0.057 (n=5)0.805 ± 0.042 (n=5)92k17.5mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaltestgrid_110kvcnnbalanced_accuracy ↑0.694 ± 0.021 (n=5)0.823 ± 0.002 (n=5)94k21.8mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globaltestgrid_110kvresnetbalanced_accuracy ↑0.786 ± 0.006 (n=5)0.785 ± 0.031 (n=5)564k31.5mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalieee39majoritybalanced_accuracy ↑0.500 ± 0.000 (n=5)0.500 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_detection_globalieee39random_forestbalanced_accuracy ↑0.922 ± 0.000 (n=5)0.778 ± 0.055 (n=5)—11.5mtrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalieee39mlpbalanced_accuracy ↑0.754 ± 0.015 (n=5)0.760 ± 0.016 (n=5)103.4M1.6htrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalieee39grubalanced_accuracy ↑0.720 ± 0.041 (n=5)0.789 ± 0.021 (n=5)211k1.7htrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalieee39cnnbalanced_accuracy ↑0.783 ± 0.011 (n=5)0.769 ± 0.044 (n=5)234k1.7htrace: output · eval-code · log · train-code · config · W&B
switch_detection_globalieee39resnetbalanced_accuracy ↑0.888 ± 0.013 (n=5)0.654 ± 0.092 (n=5)744k1.6htrace: output · eval-code · log · train-code · config · W&B

Metric definitions

headline & reported metrics for this view — the metric set is a frozen contract; each links to the exact function that computes it

metricdefinitionentry
balanced_accuracymean of per-class recall — the majority class cannot buy a good scoreevaluate()
macro_f1unweighted mean of per-class F1 (zero_division=0)evaluate()
missed_event_rateFN-rate among positives (missed events)evaluate()
false_event_trip_rateFP-rate among negatives (tripping on a benign event)evaluate()
accuracyfraction of exact-match predictionsevaluate()

Reproduce one cell

PYTHONPATH=src python -m evemtbench.baselines.runner --task switch_detection_global --protocol held_out \
    --grid <grid> --baseline <baseline> --seeds 0 1 2 3 4