LAB/012 · Running · Self-directed, not client work

Collision-Risk Forecasting

Can machine learning tell a satellite operator, two days ahead, how risky a close approach will turn out? On ESA's public conjunction data, the honest answer was no: the simplest forecast won.

The question
Two days before a close approach, can a model predict the final collision risk better than simply trusting the latest estimate?
What it showed
No. With only 66 high-risk training events, every gradient-boosting variant lost to the "latest known risk" forecast, in cross-validation and on the test set. That forecast scored 0.694, exactly the published baseline.
What it means for you
Before building a model, measure the simple rule your operators already use. Sometimes the right deliverable is that baseline, with a known error, plus the evidence for why a model is not worth deploying yet.

Want this measured on your own data? AI feasibility diagnostic, €1,900 · about one week.

When two objects in orbit are predicted to pass close to each other, the operator receives a series of conjunction data messages (CDMs), each with an updated collision probability. A manoeuvre has to be decided roughly two days before the closest approach. The question this Lab asks is the one an operations team would: at that point, can a model predict the final risk better than simply trusting the latest estimate?

The data and the rules

ESA published real, anonymised CDMs from 2015–2019 for its Spacecraft Collision Avoidance Challenge (Kelvins dataset, CC BY 4.0). The task, the test set and the metric are the competition’s own:

  • Input: the CDMs released at least 2 days before closest approach.
  • Target: the risk in the last CDM, less than a day before closest approach.
  • Score: mean squared error on the events that end up high-risk (above 10⁻⁶), divided by the F2 score for spotting them. Lower is better.

Two published reference points: the “latest known risk” forecast scored 0.694, and the best of 97 teams scored 0.556. Only 12 teams beat the simple forecast.

Results

Test set: 2167 real ESA conjunction events, 150 of them high-risk. Lower L is better.
Score L (lower is better)MSE, high-riskF2PrecisionRecall
Constant risk 10⁻⁵ (sanity baseline)2.5040.6790.2716.9%100.0%
Latest known risk (selected by cross-validation)0.6940.5130.73964.6%76.7%
Gradient boosting, best in cross-validation (not selected)2.2341.3010.58265.4%56.7%
Published: latest-risk baseline (Uriot et al. 2020)0.6940.5130.739——
Published: winning team of 970.556————
Our latest-risk score reproduces the published one to three decimals, which validates the scoring · LAB/012 · measured 2026-10-08 · not used

My scorer reproduces the published baseline exactly (0.694, with the same error and F2 to three decimals), so the numbers below are on the competition’s own scale.

No model beat the simple forecast. Every method was chosen by cross-validation on training events only, with the simple forecast as one of the candidates, and it won every time:

Cross-validation on training events: score LLower is better · used to choose the method
  1. Latest known risk0.80
  2. Gradient boosting, raw target (off scale: never predicted a high-risk event)8.00
  3. Gradient boosting, clipped target6.33
  4. Gradient boosting, gated target2.06

The reason is in the data: only 66 high-risk events in the training set that match the test conditions. A flexible model trained mostly on harmless events learns to nudge thousands of them over the alarm threshold. Restricting the model to events that were already high-risk (the “gated” variant, added after a first test evaluation and labelled as such) helped, but not enough. The competition’s top teams did beat it, by 0.07 to 0.14 on the score; with standard tools, and without ever choosing a method on the test set, this Lab did not.

What this means in practice

  • Measure the incumbent first. The rule operators already use (trust the latest estimate) is often a strong baseline, and it costs nothing to run.
  • Count your positive examples. With a few dozen of the events that matter, a model’s apparent sophistication mostly adds variance.
  • A negative result is a deliverable. Knowing that a model is not yet worth deploying, and by how much, is exactly what a feasibility assessment is for.
Method and environment
data
ESA Spacecraft Collision Avoidance Challenge (Kelvins), anonymised CDMs 2015–2019, CC BY 4.0; 8293 eligible training events, 2167 test events
task
from CDMs released at least 2 days before closest approach, predict log10 of the risk in the last CDM (less than 1 day before)
metric
L = MSE on high-risk events (final risk ≥ 10⁻⁶) / F2 over all events; predictions below −6 clipped to −6.001 (paper Eq. 1)
selection
5-fold cross-validation on training events, grouped by mission; the latest-risk baseline was a candidate; the test set was used for reporting only
models
scikit-learn HistGradientBoostingRegressor on latest-CDM and trend features, predicting the change from the latest risk; raw, clipped and gated variants (gated added after a first test evaluation)
environment
cpu: i5-13400F · gpu: not used
measured
2026-10-08

Limits

One public dataset, anonymised and from 2015–2019; standard gradient boosting rather than the competitions’ tuned pipelines. The score metric weights high-risk events heavily, as an operator would. The code is in labs/collision-risk in this site’s repository.

Public data only; this Lab covers collision-risk analysis for spacecraft safety.

New Lab results by email

A short email when a new measurement is published, a few times a year. No newsletter fluff, reply to unsubscribe.