Task — switch_type_global

benchmark version 1.1.0 · held_out results v1.0.0 · splits v1.0/held_out, v1.1/multi_grid
updated 2026-09-29 13:18 UTC · git 7243f310ddbc

tier 2 · switch_type · global view · 8 classes · all cells measured

What this task is. Given a switching window, name what kind of switching action it is. Distinguishing planned topology changes from protection trips matters for post-event analysis. Global view: every cubicle in the grid at once (width differs per grid), the wide-area upper bound on observability.

How it's scored

Headline metric balanced_accuracy (↑ higher is better); full metric set: balanced_accuracy · macro_f1 · accuracy. Metrics are a frozen contract (src/evemtbench/evaluation/metrics.py); label derivation: src/evemtbench/tasks/labels.

Example window

example waveform window

event switch_inrush_lv · grid double_line (adapt_grid/test split) · sample 1, window #23, cubicle 3 (max-RMS) · 50 ms / 480 samples @ 9.6 kHz · top: 3-phase current, bottom: 3-phase voltage · figure provenance in assets/manifest.json

Results (held_out)

switch_type_global — every grid × baseline; click headers to sort
taskgridbaselineheadlinetestbenchmarkparamsfit/seedtrace
switch_type_globalcigre_mvmajoritybalanced_accuracy ↑0.143 ± 0.000 (n=5)0.143 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_type_globalcigre_mvrandom_forestbalanced_accuracy ↑0.605 ± 0.005 (n=5)0.751 ± 0.003 (n=5)—54strace: output · eval-code · log · train-code · config · W&B
switch_type_globalcigre_mvmlpbalanced_accuracy ↑0.592 ± 0.002 (n=5)0.687 ± 0.013 (n=5)42.9M10.7mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalcigre_mvgrubalanced_accuracy ↑0.739 ± 0.012 (n=5)0.781 ± 0.020 (n=5)118k26.6mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalcigre_mvcnnbalanced_accuracy ↑0.726 ± 0.005 (n=5)0.782 ± 0.016 (n=5)125k12.4mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalcigre_mvresnetbalanced_accuracy ↑0.628 ± 0.064 (n=5)0.683 ± 0.074 (n=5)603k3.3mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaldouble_linemajoritybalanced_accuracy ↑0.250 ± 0.000 (n=5)0.250 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_type_globaldouble_linerandom_forestbalanced_accuracy ↑0.729 ± 0.002 (n=5)0.304 ± 0.000 (n=5)—19strace: output · eval-code · log · train-code · config · W&B
switch_type_globaldouble_linemlpbalanced_accuracy ↑0.705 ± 0.009 (n=5)0.732 ± 0.011 (n=5)12.0M4.5mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaldouble_linegrubalanced_accuracy ↑0.696 ± 0.014 (n=5)0.720 ± 0.012 (n=5)69k2.1mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaldouble_linecnnbalanced_accuracy ↑0.710 ± 0.006 (n=5)0.678 ± 0.036 (n=5)68k3.9mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaldouble_lineresnetbalanced_accuracy ↑0.791 ± 0.050 (n=5)0.738 ± 0.007 (n=5)531k7.2mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaltestgrid_110kvmajoritybalanced_accuracy ↑0.200 ± 0.000 (n=5)0.200 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_type_globaltestgrid_110kvrandom_forestbalanced_accuracy ↑0.658 ± 0.005 (n=5)0.247 ± 0.001 (n=5)—37strace: output · eval-code · log · train-code · config · W&B
switch_type_globaltestgrid_110kvmlpbalanced_accuracy ↑0.694 ± 0.002 (n=5)0.692 ± 0.026 (n=5)26.7M6.2mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaltestgrid_110kvgrubalanced_accuracy ↑0.747 ± 0.006 (n=5)0.766 ± 0.029 (n=5)92k5.7mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaltestgrid_110kvcnnbalanced_accuracy ↑0.729 ± 0.003 (n=5)0.739 ± 0.007 (n=5)95k5.6mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globaltestgrid_110kvresnetbalanced_accuracy ↑0.835 ± 0.056 (n=5)0.865 ± 0.084 (n=5)565k7.3mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalieee39majoritybalanced_accuracy ↑0.200 ± 0.000 (n=5)0.200 ± 0.000 (n=5)—0strace: output · eval-code · log · train-code · config · W&B
switch_type_globalieee39random_forestbalanced_accuracy ↑0.872 ± 0.002 (n=5)0.664 ± 0.023 (n=5)—2.4mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalieee39mlpbalanced_accuracy ↑0.719 ± 0.005 (n=5)0.270 ± 0.005 (n=5)103.4M20.0mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalieee39grubalanced_accuracy ↑0.669 ± 0.016 (n=5)0.251 ± 0.017 (n=5)212k13.3mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalieee39cnnbalanced_accuracy ↑0.682 ± 0.011 (n=5)0.231 ± 0.030 (n=5)235k12.9mtrace: output · eval-code · log · train-code · config · W&B
switch_type_globalieee39resnetbalanced_accuracy ↑0.689 ± 0.022 (n=5)0.234 ± 0.021 (n=5)745k8.2mtrace: output · eval-code · log · train-code · config · W&B

Metric definitions

headline & reported metrics for this view — the metric set is a frozen contract; each links to the exact function that computes it

metricdefinitionentry
balanced_accuracymean of per-class recall — the majority class cannot buy a good scoreevaluate()
macro_f1unweighted mean of per-class F1 (zero_division=0)evaluate()
accuracyfraction of exact-match predictionsevaluate()

Reproduce one cell

PYTHONPATH=src python -m evemtbench.baselines.runner --task switch_type_global --protocol held_out \
    --grid <grid> --baseline <baseline> --seeds 0 1 2 3 4