Spacecraft Telemetry Anomaly Detection
Five anomaly detectors on real spacecraft telemetry from NASA and ESA, judged the way an operations team would: how many real problems they catch, how many false alarms they raise, and how late.
Satellites send back thousands of telemetry channels: voltages, temperatures, currents, wheel speeds. Operations engineers cannot watch them all, so ground systems raise alarms. The question for any automated detector is not its benchmark score but what it does to the people on shift: does it catch the problems that matter, how many false alarms does it add, and how late does it fire?
This Lab measures five detectors on two public sets of real spacecraft telemetry:
- NASA SMAP and MSL (the SMAP soil-moisture satellite and the Curiosity rover; Hundman et al., 2018): 81 channels and 104 labelled anomalies, the most widely used dataset of its kind.
- ESA Mission1 from the ESA Anomalies Dataset (Kotowski et al., 2024): 14 years of a real ESA mission, annotated by ESA’s spacecraft operations engineers. The six channels and the train/test split are the ones ESA’s benchmark (ESA-ADB) recommends: trained on 2000–2006, tested on 2007–2013.
The detectors
- Static limits: alarm when a value leaves the range seen in training. This is what most ground systems do today.
- Rolling z-score: alarm when a value is far from the recent average, relative to recent variation.
- Isolation Forest: a standard machine-learning outlier detector, on short windows of the signal.
- Matrix profile: alarm when a stretch of the signal looks like nothing in the training data.
- Telemanom: the method published with the NASA dataset, an LSTM network that forecasts each channel and alarms when the forecast error is unusually high.
Detectors 1–4 set their threshold from training data alone: the highest score seen on known-good data. Telemanom sets its own threshold without labels. No setting was tuned on the test labels.
How alarms are counted: an anomaly is caught if any alarm falls inside it. Alarms close together are merged into one episode, one notification for the operator. An episode that touches no labelled anomaly is a false alarm.
NASA SMAP and MSL
| Events caught | Caught | False alarm episodes | Alarms on a labelled event | Median delay (samples) | |
|---|---|---|---|---|---|
| Static limits | 47 / 104 | 45.2% | 125 | 32.1% | 17 |
| Rolling z-score | 55 / 104 | 52.9% | 163 | 38.5% | 44 |
| Isolation Forest | 46 / 104 | 44.2% | 355 | 16.7% | 44 |
| Matrix profile | 58 / 104 | 55.8% | 138 | 45.9% | 26 |
| Telemanom (LSTM) | 75 / 104 | 72.1% | 179 | 43.2% | 57 |
- Static limits45.2%
- Rolling z-score52.9%
- Isolation Forest44.2%
- Matrix profile55.8%
- Telemanom (LSTM)72.1%
Telemanom catches the most: 75 of 104 anomalies (85% on SMAP, matching the published figure). But every detector raises more false alarm episodes than correct ones with these thresholds. For the operations team that matters more than the catch rate: a detector that is wrong more often than right gets ignored.
The thresholds can be moved, and that is where the real decision lies:
| k = 1.0 | k = 1.25 | k = 1.5 | k = 2.0 | k = 3.0 | |
|---|---|---|---|---|---|
| Static limits | 47 / 125 | 43 / 14 | 42 / 6 | 37 / 1 | 35 / 1 |
| Rolling z-score | 55 / 163 | 46 / 51 | 45 / 28 | 41 / 18 | 41 / 14 |
| Isolation Forest | 46 / 355 | 8 / 0 | 6 / 0 | 0 / 0 | 0 / 0 |
| Matrix profile | 58 / 138 | 37 / 46 | 31 / 35 | 27 / 33 | 26 / 18 |
Widen the plain limit check by a quarter and the false alarms drop from 125 to 14, while it still catches 43 anomalies. Twice the training range: 37 caught, one false alarm. A simple limit check with a well-chosen margin catches over a third of these anomalies almost for free. That matches a known criticism of this dataset (Wu and Keogh, 2021): many of its anomalies are trivially easy, which flatters complex methods. It is also why the second dataset matters.
ESA Mission1
| Events caught | Caught | False alarm episodes | Alarms on a labelled event | Anomalies caught | Rare events caught | |
|---|---|---|---|---|---|---|
| Static limits | 41 / 65 | 63.1% | 1,527 | 3.7% | 26 / 29 | 15 / 36 |
| Rolling z-score | 1 / 65 | 1.5% | 0 | 100.0% | 1 / 29 | 0 / 36 |
| Isolation Forest | 7 / 65 | 10.8% | 0 | 100.0% | 4 / 29 | 3 / 36 |
| Matrix profile | 3 / 65 | 4.6% | 38 | 50.0% | 2 / 29 | 1 / 36 |
| Telemanom (LSTM) | 34 / 65 | 52.3% | 31 | 67.7% | 26 / 29 | 8 / 36 |
- Static limits89.7%
- Rolling z-score3.4%
- Isolation Forest13.8%
- Matrix profile6.9%
- Telemanom (LSTM)89.7%
On seven years of real mission data the picture changes:
- Static limits fail. They catch 41 events but raise 1,527 false alarm episodes, because a long mission’s normal levels drift: values never seen in training keep appearing without anything being wrong.
- Telemanom is the only detector that is both useful and trusted: it caught 26 of the 29 anomalies, and two thirds of its alarm episodes fell on annotated events.
- It caught only 8 of 36 rare events: unusual but planned activities that ESA labels separately from anomalies. A forecaster that learned the normal pattern also misses some deliberate departures from it, which for anomaly detection is arguably the right behaviour.
The two simple detectors that look useless at their default threshold are a calibration problem, not a modelling one:
| 99.9% | 99.99% | 99.999% | max | |
|---|---|---|---|---|
| Static limits | 58 / 4445 | 49 / 3384 | 45 / 3301 | 41 / 1527 |
| Rolling z-score | 50 / 3815 | 27 / 3 | 23 / 0 | 1 / 0 |
| Isolation Forest | 14 / 2545 | 8 / 0 | 7 / 0 | 7 / 0 |
| Matrix profile | 3 / 166 | 3 / 97 | 3 / 47 | 3 / 38 |
Over seven years of training data, “never exceed the highest score ever seen” sets the bar too high. Set at the 99.99th percentile of nominal training scores instead, the plain rolling z-score caught 27 events with only 3 false alarms. How the threshold is set mattered as much as which model was used.
What this means for an operations team
- Start by measuring the existing limit checks on your own archived anomalies. With a sensible margin they may already catch a useful share at almost no cost.
- Judge any detector by false alarms per real catch, on your mission’s data, before it goes on shift. A good benchmark score can hide hundreds of false alarms.
- Expect drift. Fixed limits learned years ago will fire constantly; anything deployed needs recalibration, or a method that models normal behaviour, as telemanom does.
- Treat the threshold as an operational decision, chosen with the people who will answer the alarms, and revisit it.
Method and environment
- nasa data
SMAP/MSL (Hundman et al., 2018), 81 channels, 104 events after merging the duplicated P-2 row; train/test as published- esa data
ESA Anomalies Dataset Mission1 (Kotowski et al., 2024), channels 41–46; 2000–2006 train, 2007–2013 test; 2min bins keeping each bin's most extreme sample; annotated training events (±1 h) replaced by straight lines (training labels only)- detectors
static limits (training min/max); rolling z-score (100); Isolation Forest (50-sample windows); matrix profile (m = 100; ESA at 10-min bins); telemanom LSTM + nonparametric dynamic threshold (ported)- thresholds
detectors 1–4: maximum of the score on training data (sweeps report alternatives as curves); telemanom: its own unsupervised threshold- scoring
event caught if any alarm inside it; alarms merged into episodes across gaps ≤ 50 samples; episode overlapping no event = false alarm; delay = first alarm − event start- environment
gpu: RTX 4070 12 GB · cpu: i5-13400F- measured
- 2026-10-08
Limits
Two missions’ worth of public data and five detectors with standard settings. Detection delay is measured in samples, and telemanom’s threshold is computed over windows that include later data (as in the original method), so it is not strictly real-time as evaluated. On ESA the data was reduced to 2-minute steps (keeping each step’s most extreme value), matrix profile ran at 10-minute steps, and alarms from the six channels were combined; the metrics are this Lab’s operator-oriented ones, not ESA-ADB’s official set, so the numbers cannot be compared directly with its published tables. Annotated training-period events were removed before fitting, using training labels only. The code is in labs/telemetry-anomaly in this site’s repository.
All data is public; this Lab covers detection and analysis of spacecraft health only.