tier 1 · fault_detection · local view · 2 classes · all cells measured
Headline metric balanced_accuracy (↑ higher is better); full metric set: balanced_accuracy · macro_f1 · missed_fault_rate · false_alarm_rate · accuracy. Metrics are a frozen contract (src/evemtbench/evaluation/metrics.py); label derivation: src/evemtbench/tasks/labels.
event flt_1phg_incipient · grid double_line (adapt_grid/test split) · sample 2, window #46, cubicle 3 (max-RMS) · 50 ms / 480 samples @ 9.6 kHz · top: 3-phase current, bottom: 3-phase voltage · figure provenance in assets/manifest.json
| task | grid | baseline | headline | test | benchmark | params | fit/seed | trace |
|---|---|---|---|---|---|---|---|---|
| fault_detection_local | cigre_mv | majority | balanced_accuracy ↑ | 0.500 ± 0.000 (n=5) | 0.500 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | cigre_mv | threshold | balanced_accuracy ↑ | 0.660 ± 0.000 (n=5) | 0.810 ± 0.000 (n=5) | — | 2s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | cigre_mv | random_forest | balanced_accuracy ↑ | 0.836 ± 0.002 (n=5) | 0.983 ± 0.001 (n=5) | — | 17s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | cigre_mv | mlp | balanced_accuracy ↑ | 0.772 ± 0.002 (n=5) | 0.958 ± 0.003 (n=5) | 1.6M | 2.9m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | cigre_mv | gru | balanced_accuracy ↑ | 0.825 ± 0.036 (n=5) | 0.975 ± 0.013 (n=5) | 52k | 2.7m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | cigre_mv | cnn | balanced_accuracy ↑ | 0.826 ± 0.006 (n=5) | 0.970 ± 0.004 (n=5) | 49k | 2.4m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | cigre_mv | resnet | balanced_accuracy ↑ | 0.831 ± 0.043 (n=5) | 0.969 ± 0.008 (n=5) | 506k | 8.2m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | majority | balanced_accuracy ↑ | 0.500 ± 0.000 (n=5) | 0.500 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | threshold | balanced_accuracy ↑ | 0.729 ± 0.000 (n=5) | 0.916 ± 0.000 (n=5) | — | 2s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | random_forest | balanced_accuracy ↑ | 0.881 ± 0.001 (n=5) | 0.499 ± 0.000 (n=5) | — | 16s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | mlp | balanced_accuracy ↑ | 0.850 ± 0.001 (n=5) | 0.981 ± 0.000 (n=5) | 1.6M | 1.1m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | gru | balanced_accuracy ↑ | 0.864 ± 0.017 (n=5) | 0.984 ± 0.004 (n=5) | 52k | 2.2m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | cnn | balanced_accuracy ↑ | 0.872 ± 0.014 (n=5) | 0.987 ± 0.001 (n=5) | 49k | 2.0m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | double_line | resnet | balanced_accuracy ↑ | 0.883 ± 0.030 (n=5) | 0.989 ± 0.008 (n=5) | 506k | 7.5m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | majority | balanced_accuracy ↑ | 0.500 ± 0.000 (n=5) | 0.500 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | threshold | balanced_accuracy ↑ | 0.755 ± 0.000 (n=5) | 0.886 ± 0.000 (n=5) | — | 2s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | random_forest | balanced_accuracy ↑ | 0.882 ± 0.001 (n=5) | 0.672 ± 0.002 (n=5) | — | 15s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | mlp | balanced_accuracy ↑ | 0.847 ± 0.001 (n=5) | 0.982 ± 0.000 (n=5) | 1.6M | 1.1m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | gru | balanced_accuracy ↑ | 0.846 ± 0.063 (n=5) | 0.963 ± 0.035 (n=5) | 52k | 2.4m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | cnn | balanced_accuracy ↑ | 0.875 ± 0.014 (n=5) | 0.987 ± 0.002 (n=5) | 49k | 1.7m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | testgrid_110kv | resnet | balanced_accuracy ↑ | 0.949 ± 0.016 (n=5) | 0.995 ± 0.004 (n=5) | 506k | 9.3m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | majority | balanced_accuracy ↑ | 0.500 ± 0.000 (n=5) | 0.500 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | threshold | balanced_accuracy ↑ | 0.793 ± 0.000 (n=5) | 0.914 ± 0.000 (n=5) | — | 2s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | random_forest | balanced_accuracy ↑ | 0.931 ± 0.001 (n=5) | 0.886 ± 0.000 (n=5) | — | 15s | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | mlp | balanced_accuracy ↑ | 0.863 ± 0.001 (n=5) | 0.984 ± 0.000 (n=5) | 1.6M | 1.2m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | gru | balanced_accuracy ↑ | 0.879 ± 0.049 (n=5) | 0.979 ± 0.016 (n=5) | 52k | 2.7m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | cnn | balanced_accuracy ↑ | 0.899 ± 0.005 (n=5) | 0.987 ± 0.001 (n=5) | 49k | 2.6m | trace: output · eval-code · log · train-code · config · W&B |
| fault_detection_local | ieee39 | resnet | balanced_accuracy ↑ | 0.902 ± 0.006 (n=5) | 0.989 ± 0.001 (n=5) | 506k | 10.2m | trace: output · eval-code · log · train-code · config · W&B |
headline & reported metrics for this view — the metric set is a frozen contract; each links to the exact function that computes it
| metric | definition | entry |
|---|---|---|
balanced_accuracy | mean of per-class recall — the majority class cannot buy a good score | evaluate() |
macro_f1 | unweighted mean of per-class F1 (zero_division=0) | evaluate() |
missed_fault_rate | FN / (TP+FN) — fraction of true faults the model missed | evaluate() |
false_alarm_rate | FP / (FP+TN) — fraction of non-faults flagged | evaluate() |
accuracy | fraction of exact-match predictions | evaluate() |
PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_detection_local --protocol held_out \
--grid <grid> --baseline <baseline> --seeds 0 1 2 3 4