Data Lab / First light for a thermosphere sensor made of falling satellites

First light on a real storm: a fleet-scale drag sensor meets Dst −150

Dek: 10,950 decaying Starlinks observed the deepest geomagnetic storm of the period in near-real time. The measured drag increase is positive and resampling-robust but not significant against a quiet-epoch null — a sharper null that moves the limiting factor from storm availability to the two-line-element noise floor.


Abstract

When a geomagnetic storm heats the upper atmosphere, the thermosphere expands, air density at satellite altitude rises, and low-orbiting spacecraft feel more drag. An earlier TerraPulse pilot built a fleet-scale drag sensor from the pooled orbital decay of the Starlink fleet, could only test a shallow substitute storm (Dst −88 nT), returned a non-detection, and closed by predicting the sensor would "gain power when a strong storm lands inside coverage." A strong storm has since landed: on 2026-07-04 the Dst index reached −150 nT, the deepest excursion of the period — deeper than even the originally pre-registered −101 nT target — and squarely inside our now-extended coverage. This paper takes up the pilot's own falsifiable prediction and re-points the test at it.

From 111 days of CelesTrak orbital elements (2026-03-27 to 2026-07-15, 10,950 Starlinks, 930,196 object-days) we built a per-object daily decay rate and aggregated the non-thrusting population into a single fleet-median time series. Around the storm the fleet-median decay rate rose from 8.5 to 11.1 × 10⁻⁶ rev/day², a difference of +2.5 × 10⁻⁶. The rise excludes zero under object resampling (95% interval [1.6, 3.4] × 10⁻⁶) and maneuvering satellites move the opposite way (−3.3 × 10⁻⁶), so it is a real, thruster-clean rise. But it does not clear the honest bar: a quiet control date drifts +1.0 × 10⁻⁶ on its own, the storm response sits at about 1.1 standard deviations of the quiet-epoch background (z = 1.06, one-sided p = 0.12), and the continuous decay-rate versus storm-intensity correlation is flat over the full record (Pearson r = 0.02, n = 101 days).

The decisive result is comparative. A storm 70% deeper than the shallow substitute moved the response magnitude and its significance almost not at all (+2.5 vs +2.6 × 10⁻⁶; p = 0.12 vs 0.19). The limiting factor is therefore not storm availability, as the pilot supposed, but the two-line-element noise floor of a fleet median dominated by near-static shell satellites. We report a sharper, earned non-detection with a working instrument that has finally met a genuinely deep storm.

Ingestion note. The storm is recent — eleven days into the archive. The post-storm impulse window (days 0 to +5) is complete and the headline result is fully determined, but the fuller superposed-epoch tail (offsets +12 to +15) and the fixed-cohort survivorship check are still accreting as the live feed ingests each day. They firm the estimate up over the coming days and do not change its sign or significance.


1. Why ask this

The thermosphere is the only part of the climate system that breathes on a timescale of hours. A solar flare or a geomagnetic storm dumps energy into the upper atmosphere, it puffs outward, and the air that low-Earth-orbit satellites fly through gets denser. The effect is large and well documented. In February 2022 a moderate storm raised drag enough to pull roughly 38 freshly launched Starlinks back into the atmosphere before they reached their operating altitude.

Satellite drag has been used as a thermospheric density proxy for decades. What is new is not the physics but the sample. There are now more than 10,000 Starlinks in a narrow altitude shell, each reporting orbital elements many times a day — in principle a dense, fleet-scale instrument for watching the upper atmosphere expand and contract. The pilot version of this monitor asked the most basic question that instrument has to answer before any ambitious version is worth attempting: can it see a single, ordinary geomagnetic storm? It could not answer cleanly, because the one deep storm of the era pre-dated the catalog and the only in-coverage storm was a shallow one. It closed with a concrete, falsifiable prediction — the sensor would gain power once a strong storm landed in coverage — and it named the target: a Dst below about −150 nT, with a clean onset and coverage on both sides.

That is precisely the storm that arrived on 2026-07-04. This paper is the re-pointing the prediction invited.

2. Data

Starlink orbital elements (via the satellite_decay dex). Each satellite reports a mean motion, the number of orbits it completes per day. As an orbit decays the satellite drops, speeds up, and its mean motion rises. The rate of that rise is the drag signal and therefore a thermospheric density proxy. The panel is drawn from the TerraPulse satellite_decay dex — a per-object catalog whose slot carries the measured CelesTrak element track of one catalogued object. Regrouping the Starlink slots into a daily panel gives 10,950 distinct satellites, 930,196 object-days, spanning 2026-03-27 to 2026-07-15 (111 days).

Space-weather drivers (daily). Dst index (storm intensity, the most negative daily excursion), DSCOVR solar-wind speed, SILSO sunspot number (a proxy for the slow solar-EUV heating baseline), and GOES X-ray flux (flare context), each resolved through the TerraPulse provenance registry to its geophysical-index source.

The storm the pilot awaited. On 2026-07-04 the Dst index reached −150 nT, the deepest excursion of the record and deeper than the originally pre-registered −101 nT target. It falls well inside coverage, with months of pre-storm baseline. The disturbed sequence is (−49, −150, −65, −51, −42, −54) on 07-03 through 07-08, so the onset is comparatively sharp: a drop from −2 to −150 across 07-02 → 07-04. We re-point the superposed-epoch test at 2026-07-04 by the same pre-declared rule the pilot used — "deepest Dst minimum fully inside coverage" — now satisfied by a genuinely deep storm rather than a shallow substitute.

Ingestion in progress, stated plainly. The storm is recent; coverage extends to 2026-07-15, eleven days past epoch zero. The post-storm impulse window (days 0 to +5, i.e. 07-04 through 07-09) is complete, and the headline result is fully determined. What is not yet ingested is the fuller epoch tail (offsets +12 to +15) and, because the storm sits near the trailing edge of the archive, the fixed-cohort survivorship check thins (to 57 satellites tracked on every day of the window, versus 833 in the pilot's mid-window test). These accrue automatically as the live CelesTrak feed adds each day. We publish now because the monitor caught a great storm in near-real time and that is the point of a standing monitor; the estimate will firm up, not flip.

3. Method

Per-object decay rate. For each satellite we took the daily-median mean motion (to suppress the jitter between the several orbital-element updates per day) and estimated a daily decay rate as the slope of a centered seven-day (±3 day) linear fit. Robust slopes, not single-pair differences, because two-line-element sets are individually noisy.

Clean-drag filter. Operational Starlinks raise their own orbits with ion thrusters, which masks or inverts the drag signal. We kept only satellites whose mean motion never stepped sharply downward (no orbit raise above 0.002 rev/day day-to-day), whose net change over the window was positive (decaying), and which were tracked at least 20 days: 5,295 clean-drag objects against 5,422 maneuvering ones. As in the pilot, this filter is leaky toward inertness — a large share of "clean" objects barely move and sit at the noise floor, and their presence in the fleet median is the single biggest reason that median is noisy.

Fleet aggregate. The daily median decay rate over the clean population, median rather than mean to resist residual maneuver outliers.

Tests.

  • H1, storm impulse: superposed-epoch of the fleet-median rate on days −10 to +15 around 2026-07-04, comparing a pre-storm baseline (days −10 to −2) to a post-storm window (days 0 to +5). Significance against a shuffle null of random quiet epoch days, plus an object-resampling bootstrap.
  • H2, continuous driver: daily fleet rate against storm intensity (−Dst), sunspots, and solar-wind speed, at lags 0 to 2, Pearson and Spearman, with a partial correlation removing the sunspot baseline.
  • H3, altitude gradient: the storm response stratified by altitude (mean-motion band).
  • H0, null: the fleet rate is within background variability, with no driver correlation that survives honest significance testing.

4. Results

4.1 H1: a real rise that does not clear the quiet-time background

Around the 07-04 storm the fleet-median decay rate rose from a pre-storm baseline of 8.5 × 10⁻⁶ to a post-storm 11.1 × 10⁻⁶ rev/day², a difference of +2.5 × 10⁻⁶. Resampling the clean objects gives a 95% bootstrap interval of [1.6, 3.4] × 10⁻⁶ on the difference, which excludes zero: the rise is stable across which satellites we sample, not carried by a handful.

Two things make this the most honest positive the monitor has produced, and one thing takes it away:

  • The clean control passes. Maneuvering satellites — the ones we excluded as thrusting, which should not show a drag response — move in the opposite direction, −3.3 × 10⁻⁶. The positive clean-drag shift is not a thruster artifact.
  • The quiet control does not. Running the identical superposed-epoch on a quiet date, 2026-05-25, still yields a +1.0 × 10⁻⁶ "response." The instrument produces a positive shift of about 40% of the storm's size with no storm present.
  • The shuffle null. Against the pool of quiet epoch days, the 07-04 response lands at the 86th percentile (z = 1.06, one-sided p = 0.12) — short of the 95th-percentile bar. The storm-specific excess over the quiet baseline is about +1.5 × 10⁻⁶, roughly half a standard deviation of the day-to-day scatter.
  • The timing is noisy. The epoch curve is actually negative on the storm day itself, rises to its post-window peak at offsets +4 and +5, but the pre-storm baseline is itself elevated, so the before/after contrast is muddied rather than impulsive.

The honest reading of H1 is a positive but non-significant response. The fleet rate rose in the predicted direction and the rise is thruster-clean, but a rise of this size is still something the instrument does on quiet days.

4.2 H2: the continuous driver correlation is null

Over the full record the daily fleet rate against storm intensity (−Dst) is flat: Pearson r = 0.02 at lag 0 (n = 101 days), r = 0.16 at lag 2, none significant, and the partial correlation removing the sunspot baseline is r = 0.14 (p = 0.20). The strongest cell anywhere in the nine-way table is solar-wind speed at lag 2 (r = 0.18, p = 0.08), which does not survive multiplicity. The pilot's earlier r = 0.40 was a short-window, leverage-driven artifact; over three times the coverage it does not persist. H2 is null.

4.3 H3: the altitude gradient is an artifact of degenerate binning

The clean population is concentrated in a razor-thin mean-motion shell, so quartile altitude bands are degenerate. The band responses (335, 2, −3, 1 × 10⁻⁶) are non-monotonic and dominated by a single band contaminated by a few end-of-life objects with baseline rates two orders of magnitude above the rest. We do not claim H3.

5. Why this is still a non-detection — and what it now means

The pilot named four suspects for its null. The re-pointing convicts one and clears another.

  1. Storm availability — cleared. A deep, sharply-onset storm (−150 nT) is now in coverage, and the fleet response (+2.5 × 10⁻⁶, p = 0.12) is statistically indistinguishable from the response to the shallow −88 nT substitute (+2.6 × 10⁻⁶, p = 0.19). A storm 70% deeper bought essentially no additional detectability.
  2. The two-line-element noise floor — convicted. The fleet median is dominated by near-static shell satellites whose day-to-day scatter (about 3 × 10⁻⁶ rev/day²) is comparable to the storm effect itself, so no single storm, however deep, can lift the raw fleet median cleanly above it.
  3. The effective sample is days, not satellites. Thousands of satellites collapse into one daily series with a limited number of independent points; power lives in independent days.
  4. The storm shape. Even this comparatively sharp onset is not the idealized impulse a superposed-epoch design wants.

The headline lesson inverts the pilot's closing expectation. The limiting resource is not the sky but the instrument's precision. That redirects the program: the path to detection runs through decay-magnitude-weighted aggregation and a hard restriction to vigorous decayers — lifting the signal off the shell-satellite noise floor — not through waiting for a bigger storm.

6. Limitations

  • Exploratory re-pointing. Taking up the pilot's prediction is legitimate but it is an exploratory re-point, not a fresh pre-registration.
  • One storm. A single event, however deep, cannot establish or rule out a population-level drag response.
  • Ingestion in progress. The storm is eleven days into the archive; the epoch tail and fixed-cohort check are provisional and still accreting.
  • No direct EUV index. We proxy solar heating with sunspot number; F10.7 and Kp are not ingested. The solar-wind and sunspot series also lag the satellite panel by about two weeks, weakening the driver side of H2 but not the H1 null.
  • TLE provenance. All decay rates derive from public two-line elements, whose per-element accuracy and cadence we do not control.

7. What would change the verdict

This is a standing monitor, and it now knows its own limit. The next gains are instrumental, not meteorological:

  • A decay-weighted aggregate. Weighting the fleet statistic by decay magnitude, or restricting hard to vigorous decayers, would lift the signal off the noise floor. This is now the single most promising move.
  • Autocorrelation-aware inference as standard, so the continuous-driver test is honest from the start.
  • More independent storm days, accruing automatically as the archive grows.
  • Solar-cycle horizon. With years of history, the slow expansion and contraction of the thermosphere across the solar cycle becomes the target, and eventually the decadal carbon-dioxide cooling and contraction of the upper atmosphere — the genuine climate signal. Both are explicitly out of reach now and held as dream-tier seeds.

8. Conclusion

We built a fleet-scale thermospheric-drag instrument from thousands of decaying Starlinks and, taking up the pilot's own falsifiable prediction, pointed it at the deepest geomagnetic storm of the era. It returned a positive, resampling-robust rise in the predicted direction that nonetheless does not clear a quiet-epoch null — and, decisively, a storm far deeper than the previous substitute barely changed the result. The instrument works, the pipeline is reproducible, and the reason for the non-detection is now pinned to the two-line-element noise floor rather than to the absence of a strong storm. The standing monitor gains a sharper entry: an honest null with a working sensor that has finally seen a great storm and measured, precisely, that it is not yet precise enough.


Data and reproducibility

  • scripts/extract.py: regroups the Starlink daily panel from the satellite_decay dex and pulls the daily drivers via the provenance registry.
  • scripts/analyze.py: clean-drag filter, decay rates, superposed-epoch, correlations, sensitivity. Writes data/results.json.
  • scripts/viz.py: the storm-vs-quiet superposed-epoch figure. Writes www/sea-curve.png.
  • Data products: data/starlink_daily.parquet, data/starlink_decay.parquet, data/drivers_daily.parquet. The v1 published result is frozen at data/frozen/results.v1_published.json.

References

  • Emmert, J. T. (2015). Thermospheric mass density: A review.
  • Picone, J. M. et al. (2002). NRLMSISE-00 empirical model of the atmosphere.
  • Hapgood, M. et al. (2022); Dang, T. et al. Analyses of the February 2022 Starlink storm loss.
  • Roble, R. G. and Dickinson, R. E. (1989); Emmert, J. T. (2021). Carbon-dioxide cooling and secular contraction of the upper atmosphere.
  • TerraPulse: standing-monitor pilot v1 (the shallow-substitute non-detection this paper re-points).

Published: · Updated:

← Back to Data Lab