LAB/009 · Running · Self-directed, not client work

Spacecraft Telemetry Anomaly Detection

Five anomaly detectors on real spacecraft telemetry from NASA and ESA, judged the way an operations team would: how many real problems they catch, how many false alarms they raise, and how late.

Satellites send back thousands of telemetry channels: voltages, temperatures, currents, wheel speeds. Operations engineers cannot watch them all, so ground systems raise alarms. The question for any automated detector is not its benchmark score but what it does to the people on shift: does it catch the problems that matter, how many false alarms does it add, and how late does it fire?

This Lab measures five detectors on two public sets of real spacecraft telemetry:

  • NASA SMAP and MSL (the SMAP soil-moisture satellite and the Curiosity rover; Hundman et al., 2018): 81 channels and 104 labelled anomalies, the most widely used dataset of its kind.
  • ESA Mission1 from the ESA Anomalies Dataset (Kotowski et al., 2024): 14 years of a real ESA mission, annotated by ESA’s spacecraft operations engineers. The six channels and the train/test split are the ones ESA’s benchmark (ESA-ADB) recommends: trained on 2000–2006, tested on 2007–2013.

The detectors

  1. Static limits: alarm when a value leaves the range seen in training. This is what most ground systems do today.
  2. Rolling z-score: alarm when a value is far from the recent average, relative to recent variation.
  3. Isolation Forest: a standard machine-learning outlier detector, on short windows of the signal.
  4. Matrix profile: alarm when a stretch of the signal looks like nothing in the training data.
  5. Telemanom: the method published with the NASA dataset, an LSTM network that forecasts each channel and alarms when the forecast error is unusually high.

Detectors 1–4 set their threshold from training data alone: the highest score seen on known-good data. Telemanom sets its own threshold without labels. No setting was tuned on the test labels.

How alarms are counted: an anomaly is caught if any alarm falls inside it. Alarms close together are merged into one episode, one notification for the operator. An episode that touches no labelled anomaly is a false alarm.

NASA SMAP and MSL

NASA SMAP and MSL: 81 telemetry channels, 104 labelled anomalies. Thresholds from training data only.
Events caughtCaughtFalse alarm episodesAlarms on a labelled eventMedian delay (samples)
Static limits47 / 10445.2%12532.1%17
Rolling z-score55 / 10452.9%16338.5%44
Isolation Forest46 / 10444.2%35516.7%44
Matrix profile58 / 10455.8%13845.9%26
Telemanom (LSTM)75 / 10472.1%17943.2%57
An alarm episode is one notification (alarms less than 50 samples apart are merged) · LAB/009 · measured 2026-10-08 · RTX 4070 12 GB
NASA SMAP/MSL: anomalies caughtScale 0–100% · training-data thresholds
  1. Static limits45.2%
  2. Rolling z-score52.9%
  3. Isolation Forest44.2%
  4. Matrix profile55.8%
  5. Telemanom (LSTM)72.1%

Telemanom catches the most: 75 of 104 anomalies (85% on SMAP, matching the published figure). But every detector raises more false alarm episodes than correct ones with these thresholds. For the operations team that matters more than the catch rate: a detector that is wrong more often than right gets ignored.

The thresholds can be moved, and that is where the real decision lies:

The alarm budget: the same simple detectors with their training-derived threshold multiplied by k. Events caught / false alarm episodes.
k = 1.0k = 1.25k = 1.5k = 2.0k = 3.0
Static limits47 / 12543 / 1442 / 637 / 135 / 1
Rolling z-score55 / 16346 / 5145 / 2841 / 1841 / 14
Isolation Forest46 / 3558 / 06 / 00 / 00 / 0
Matrix profile58 / 13837 / 4631 / 3527 / 3326 / 18
Out of 104 events. Reported as a curve; no setting was chosen from it · LAB/009 · measured 2026-10-08 · RTX 4070 12 GB

Widen the plain limit check by a quarter and the false alarms drop from 125 to 14, while it still catches 43 anomalies. Twice the training range: 37 caught, one false alarm. A simple limit check with a well-chosen margin catches over a third of these anomalies almost for free. That matches a known criticism of this dataset (Wu and Keogh, 2021): many of its anomalies are trivially easy, which flatters complex methods. It is also why the second dataset matters.

ESA Mission1

ESA Mission1, the six channels of ESA-ADB's lightweight subset, 2007–2013 test period (1,841,040 two-minute steps), 65 annotated events. Trained on 2000–2006.
Events caughtCaughtFalse alarm episodesAlarms on a labelled eventAnomalies caughtRare events caught
Static limits41 / 6563.1%1,5273.7%26 / 2915 / 36
Rolling z-score1 / 651.5%0100.0%1 / 290 / 36
Isolation Forest7 / 6510.8%0100.0%4 / 293 / 36
Matrix profile3 / 654.6%3850.0%2 / 291 / 36
Telemanom (LSTM)34 / 6552.3%3167.7%26 / 298 / 36
Alarms from the six channels combined. Not ESA-ADB's official metrics · LAB/009 · measured 2026-10-08 · RTX 4070 12 GB
ESA Mission1: anomalies caught (of 29)Scale 0–100% · rare events and false alarms in the table
  1. Static limits89.7%
  2. Rolling z-score3.4%
  3. Isolation Forest13.8%
  4. Matrix profile6.9%
  5. Telemanom (LSTM)89.7%

On seven years of real mission data the picture changes:

  • Static limits fail. They catch 41 events but raise 1,527 false alarm episodes, because a long mission’s normal levels drift: values never seen in training keep appearing without anything being wrong.
  • Telemanom is the only detector that is both useful and trusted: it caught 26 of the 29 anomalies, and two thirds of its alarm episodes fell on annotated events.
  • It caught only 8 of 36 rare events: unusual but planned activities that ESA labels separately from anomalies. A forecaster that learned the normal pattern also misses some deliberate departures from it, which for anomaly detection is arguably the right behaviour.

The two simple detectors that look useless at their default threshold are a calibration problem, not a modelling one:

ESA: thresholds at quantiles of each detector's score on nominal training data. Events caught / false alarm episodes.
99.9%99.99%99.999%max
Static limits58 / 444549 / 338445 / 330141 / 1527
Rolling z-score50 / 381527 / 323 / 01 / 0
Isolation Forest14 / 25458 / 07 / 07 / 0
Matrix profile3 / 1663 / 973 / 473 / 38
Out of 65 events. A higher quantile means a stricter threshold · LAB/009 · measured 2026-10-08 · RTX 4070 12 GB

Over seven years of training data, “never exceed the highest score ever seen” sets the bar too high. Set at the 99.99th percentile of nominal training scores instead, the plain rolling z-score caught 27 events with only 3 false alarms. How the threshold is set mattered as much as which model was used.

What this means for an operations team

  • Start by measuring the existing limit checks on your own archived anomalies. With a sensible margin they may already catch a useful share at almost no cost.
  • Judge any detector by false alarms per real catch, on your mission’s data, before it goes on shift. A good benchmark score can hide hundreds of false alarms.
  • Expect drift. Fixed limits learned years ago will fire constantly; anything deployed needs recalibration, or a method that models normal behaviour, as telemanom does.
  • Treat the threshold as an operational decision, chosen with the people who will answer the alarms, and revisit it.
Method and environment
nasa data
SMAP/MSL (Hundman et al., 2018), 81 channels, 104 events after merging the duplicated P-2 row; train/test as published
esa data
ESA Anomalies Dataset Mission1 (Kotowski et al., 2024), channels 41–46; 2000–2006 train, 2007–2013 test; 2min bins keeping each bin's most extreme sample; annotated training events (±1 h) replaced by straight lines (training labels only)
detectors
static limits (training min/max); rolling z-score (100); Isolation Forest (50-sample windows); matrix profile (m = 100; ESA at 10-min bins); telemanom LSTM + nonparametric dynamic threshold (ported)
thresholds
detectors 1–4: maximum of the score on training data (sweeps report alternatives as curves); telemanom: its own unsupervised threshold
scoring
event caught if any alarm inside it; alarms merged into episodes across gaps ≤ 50 samples; episode overlapping no event = false alarm; delay = first alarm − event start
environment
gpu: RTX 4070 12 GB · cpu: i5-13400F
measured
2026-10-08

Limits

Two missions’ worth of public data and five detectors with standard settings. Detection delay is measured in samples, and telemanom’s threshold is computed over windows that include later data (as in the original method), so it is not strictly real-time as evaluated. On ESA the data was reduced to 2-minute steps (keeping each step’s most extreme value), matrix profile ran at 10-minute steps, and alarms from the six channels were combined; the metrics are this Lab’s operator-oriented ones, not ESA-ADB’s official set, so the numbers cannot be compared directly with its published tables. Annotated training-period events were removed before fitting, using training labels only. The code is in labs/telemetry-anomaly in this site’s repository.

All data is public; this Lab covers detection and analysis of spacecraft health only.