tier 2 · fault_class · local view · 9 classes · all cells measured
Headline metric balanced_accuracy (↑ higher is better); full metric set: balanced_accuracy · macro_f1 · accuracy. Metrics are a frozen contract (src/evemtbench/evaluation/metrics.py); label derivation: src/evemtbench/tasks/labels.
event flt_1phg_incipient_w_arc · grid double_line (adapt_grid/test split) · sample 197, window #4531, cubicle 3 (max-RMS) · 50 ms / 480 samples @ 9.6 kHz · top: 3-phase current, bottom: 3-phase voltage · figure provenance in assets/manifest.json
| task | grid | baseline | headline | test | benchmark | params | fit/seed | trace |
|---|---|---|---|---|---|---|---|---|
| fault_classification_local | cigre_mv | majority | balanced_accuracy ↑ | 0.111 ± 0.000 (n=5) | 0.111 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | cigre_mv | random_forest | balanced_accuracy ↑ | 0.477 ± 0.005 (n=5) | 0.526 ± 0.004 (n=5) | — | 9s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | cigre_mv | mlp | balanced_accuracy ↑ | 0.527 ± 0.005 (n=5) | 0.546 ± 0.002 (n=5) | 1.6M | 54s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | cigre_mv | gru | balanced_accuracy ↑ | 0.519 ± 0.009 (n=5) | 0.542 ± 0.001 (n=5) | 53k | 2.9m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | cigre_mv | cnn | balanced_accuracy ↑ | 0.566 ± 0.028 (n=5) | 0.589 ± 0.031 (n=5) | 50k | 1.9m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | cigre_mv | resnet | balanced_accuracy ↑ | 0.656 ± 0.011 (n=5) | 0.701 ± 0.011 (n=5) | 507k | 6.6m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | double_line | majority | balanced_accuracy ↑ | 0.143 ± 0.000 (n=5) | 0.143 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | double_line | random_forest | balanced_accuracy ↑ | 0.606 ± 0.005 (n=5) | 0.539 ± 0.002 (n=5) | — | 10s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | double_line | mlp | balanced_accuracy ↑ | 0.679 ± 0.004 (n=5) | 0.704 ± 0.001 (n=5) | 1.6M | 55s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | double_line | gru | balanced_accuracy ↑ | 0.679 ± 0.010 (n=5) | 0.710 ± 0.002 (n=5) | 53k | 1.9m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | double_line | cnn | balanced_accuracy ↑ | 0.795 ± 0.012 (n=5) | 0.812 ± 0.010 (n=5) | 50k | 2.0m | trace: output · eval-code · log · train-code · config |
| fault_classification_local | double_line | resnet | balanced_accuracy ↑ | 0.855 ± 0.007 (n=5) | 0.896 ± 0.045 (n=5) | 507k | 7.5m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | testgrid_110kv | majority | balanced_accuracy ↑ | 0.143 ± 0.000 (n=5) | 0.143 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | testgrid_110kv | random_forest | balanced_accuracy ↑ | 0.628 ± 0.003 (n=5) | 0.586 ± 0.006 (n=5) | — | 9s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | testgrid_110kv | mlp | balanced_accuracy ↑ | 0.656 ± 0.010 (n=5) | 0.705 ± 0.001 (n=5) | 1.6M | 51s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | testgrid_110kv | gru | balanced_accuracy ↑ | 0.667 ± 0.012 (n=5) | 0.709 ± 0.002 (n=5) | 53k | 1.8m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | testgrid_110kv | cnn | balanced_accuracy ↑ | 0.685 ± 0.005 (n=5) | 0.727 ± 0.004 (n=5) | 50k | 1.8m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | testgrid_110kv | resnet | balanced_accuracy ↑ | 0.825 ± 0.006 (n=5) | 0.848 ± 0.007 (n=5) | 507k | 6.7m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | ieee39 | majority | balanced_accuracy ↑ | 0.143 ± 0.000 (n=5) | 0.143 ± 0.000 (n=5) | — | 0s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | ieee39 | random_forest | balanced_accuracy ↑ | 0.665 ± 0.002 (n=5) | 0.667 ± 0.002 (n=5) | — | 10s | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | ieee39 | mlp | balanced_accuracy ↑ | 0.675 ± 0.003 (n=5) | 0.711 ± 0.000 (n=5) | 1.6M | 1.5m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | ieee39 | gru | balanced_accuracy ↑ | 0.678 ± 0.010 (n=5) | 0.703 ± 0.003 (n=5) | 53k | 2.2m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | ieee39 | cnn | balanced_accuracy ↑ | 0.679 ± 0.011 (n=5) | 0.709 ± 0.002 (n=5) | 50k | 1.8m | trace: output · eval-code · log · train-code · config · W&B |
| fault_classification_local | ieee39 | resnet | balanced_accuracy ↑ | 0.681 ± 0.013 (n=5) | 0.704 ± 0.008 (n=5) | 507k | 3.3m | trace: output · eval-code · log · train-code · config · W&B |
headline & reported metrics for this view — the metric set is a frozen contract; each links to the exact function that computes it
| metric | definition | entry |
|---|---|---|
balanced_accuracy | mean of per-class recall — the majority class cannot buy a good score | evaluate() |
macro_f1 | unweighted mean of per-class F1 (zero_division=0) | evaluate() |
accuracy | fraction of exact-match predictions | evaluate() |
PYTHONPATH=src python -m evemtbench.baselines.runner --task fault_classification_local --protocol held_out \
--grid <grid> --baseline <baseline> --seeds 0 1 2 3 4