tier 2 · switch_type · global view · 8 classes · all cells measured
Headline metric balanced_accuracy (↑ higher is better); full metric set: balanced_accuracy · macro_f1 · accuracy. Metrics are a frozen contract (src/evemtbench/evaluation/metrics.py); label derivation: src/evemtbench/tasks/labels.
event switch_inrush_lv · grid double_line (adapt_grid/test split) · sample 1, window #23, cubicle 3 (max-RMS) · 50 ms / 480 samples @ 9.6 kHz · top: 3-phase current, bottom: 3-phase voltage · figure provenance in assets/manifest.json
| task | grid | baseline | headline | test | benchmark | params | fit/seed | trace |
|---|---|---|---|---|---|---|---|---|
| switch_type_global | cigre_mv | majority | balanced_accuracy ↑ | 0.143 ± 0.000 (n=5) | 0.143 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | cigre_mv | random_forest | balanced_accuracy ↑ | 0.605 ± 0.005 (n=5) | 0.751 ± 0.003 (n=5) | — | 54s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | cigre_mv | mlp | balanced_accuracy ↑ | 0.592 ± 0.002 (n=5) | 0.687 ± 0.013 (n=5) | 42.9M | 10.7m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | cigre_mv | gru | balanced_accuracy ↑ | 0.739 ± 0.012 (n=5) | 0.781 ± 0.020 (n=5) | 118k | 26.6m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | cigre_mv | cnn | balanced_accuracy ↑ | 0.726 ± 0.005 (n=5) | 0.782 ± 0.016 (n=5) | 125k | 12.4m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | cigre_mv | resnet | balanced_accuracy ↑ | 0.628 ± 0.064 (n=5) | 0.683 ± 0.074 (n=5) | 603k | 3.3m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | double_line | majority | balanced_accuracy ↑ | 0.250 ± 0.000 (n=5) | 0.250 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | double_line | random_forest | balanced_accuracy ↑ | 0.729 ± 0.002 (n=5) | 0.304 ± 0.000 (n=5) | — | 19s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | double_line | mlp | balanced_accuracy ↑ | 0.705 ± 0.009 (n=5) | 0.732 ± 0.011 (n=5) | 12.0M | 4.5m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | double_line | gru | balanced_accuracy ↑ | 0.696 ± 0.014 (n=5) | 0.720 ± 0.012 (n=5) | 69k | 2.1m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | double_line | cnn | balanced_accuracy ↑ | 0.710 ± 0.006 (n=5) | 0.678 ± 0.036 (n=5) | 68k | 3.9m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | double_line | resnet | balanced_accuracy ↑ | 0.791 ± 0.050 (n=5) | 0.738 ± 0.007 (n=5) | 531k | 7.2m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | testgrid_110kv | majority | balanced_accuracy ↑ | 0.200 ± 0.000 (n=5) | 0.200 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | testgrid_110kv | random_forest | balanced_accuracy ↑ | 0.658 ± 0.005 (n=5) | 0.247 ± 0.001 (n=5) | — | 37s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | testgrid_110kv | mlp | balanced_accuracy ↑ | 0.694 ± 0.002 (n=5) | 0.692 ± 0.026 (n=5) | 26.7M | 6.2m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | testgrid_110kv | gru | balanced_accuracy ↑ | 0.747 ± 0.006 (n=5) | 0.766 ± 0.029 (n=5) | 92k | 5.7m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | testgrid_110kv | cnn | balanced_accuracy ↑ | 0.729 ± 0.003 (n=5) | 0.739 ± 0.007 (n=5) | 95k | 5.6m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | testgrid_110kv | resnet | balanced_accuracy ↑ | 0.835 ± 0.056 (n=5) | 0.865 ± 0.084 (n=5) | 565k | 7.3m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | ieee39 | majority | balanced_accuracy ↑ | 0.200 ± 0.000 (n=5) | 0.200 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | ieee39 | random_forest | balanced_accuracy ↑ | 0.872 ± 0.002 (n=5) | 0.664 ± 0.023 (n=5) | — | 2.4m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | ieee39 | mlp | balanced_accuracy ↑ | 0.719 ± 0.005 (n=5) | 0.270 ± 0.005 (n=5) | 103.4M | 20.0m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | ieee39 | gru | balanced_accuracy ↑ | 0.669 ± 0.016 (n=5) | 0.251 ± 0.017 (n=5) | 212k | 13.3m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | ieee39 | cnn | balanced_accuracy ↑ | 0.682 ± 0.011 (n=5) | 0.231 ± 0.030 (n=5) | 235k | 12.9m | trace: output · eval-code · log · train-code · config · W&B |
| switch_type_global | ieee39 | resnet | balanced_accuracy ↑ | 0.689 ± 0.022 (n=5) | 0.234 ± 0.021 (n=5) | 745k | 8.2m | trace: output · eval-code · log · train-code · config · W&B |
headline & reported metrics for this view — the metric set is a frozen contract; each links to the exact function that computes it
| metric | definition | entry |
|---|---|---|
balanced_accuracy | mean of per-class recall — the majority class cannot buy a good score | evaluate() |
macro_f1 | unweighted mean of per-class F1 (zero_division=0) | evaluate() |
accuracy | fraction of exact-match predictions | evaluate() |
PYTHONPATH=src python -m evemtbench.baselines.runner --task switch_type_global --protocol held_out \
--grid <grid> --baseline <baseline> --seeds 0 1 2 3 4