TerraPulse Forecast Scoreboard — Product Plan
The wedge: Everyone in AI weather sells forecasts. TerraPulse sells the scoring of forecasts — against station-measured reality, at the surface, where money and safety live. The existing benchmark (WeatherBench 2) is built by a model-maker and grades models against reanalysis — that is, against other models' output. TerraPulse never forecasts, holds no team on the field, and grades against instruments. The referee's whistle, not the jersey.
The builders: Two brothers, one garage, 370 live sources, 488M observations, and a public no-models stance that took years to earn and cannot be bought — which is the entire moat.
Why now: GraphCast, GenCast, Pangu, NeuralGCM, AIFS, and a wave of startups are racing; capital is pouring in; buyers (utilities, insurers, traders, agencies) cannot tell whose skill claims are real. The standards seat in a new field gets claimed once, early. The incumbent benchmark's three known gaps are the product: reanalysis truth instead of station truth, global averages instead of local decisions, and weakness on extremes — the exact events with dollars attached.
Build Order
1. The Public Scoreboard — the credibility engine · free
One page, one sentence: "Yesterday's forecasts vs. what the instruments actually measured."
- Ingest public forecast feeds (NWS/HRRR guidance, ECMWF open data, published AI-model outputs where licensing allows); freeze each forecast at issue time — the frozen archive IS the asset, and every day not archiving is a day of history lost forever.
- Score against station observations only — never reanalysis — at user-relevant surface variables: temperature, precipitation occurrence, freeze events, wind.
- Leaderboard + honest methodology page + per-model report cards. Updated daily, automatically.
- Build cost: forecast-feed ingest + the scoring job. The observation side already exists; this is the first product on any list that monetizes the archive's depth rather than its breadth.
2. Event Verification Bulletins — the content engine · free → sponsored
After every major event: "Who called the derecho? Who missed the freeze?" — a Lab-style bulletin scoring every archived forecast against the measured outcome.
- Each bulletin is press bait, scoreboard marketing, and a sales artifact in one.
- The Fruit Ridge frost season is the flagship series: local, high-stakes, and it cross-sells the ag product to the same readers.
3. Verification Reports for Model Makers — the first revenue · $5–25k/report
Labs and weather-AI startups buy detailed, independent verification: where their model wins, where it fails, scored against station truth, citable in papers, marketing, and fundraising ("independently verified by TerraPulse").
- Sequenced after a season of public scoreboard credibility — the free board is why the paid report means something.
- Includes private pre-publication scoring for models not yet released.
4. Buyer-Side Model Selection Reports — the bigger market · $10–50k/engagement
The mirror customer: utilities, co-ops, insurers, and trading desks choosing which forecast vendor to buy. "For YOUR territory and YOUR decision thresholds (freeze at the Ridge, wind at your turbines), here is who actually performs."
- Nobody else can sell this without conflict; every vendor's own skill claims are the reason the buyer needs a referee.
- Falls out of #1's per-station scoring machinery for free.
5. Benchmark Dataset Licensing — the compounding asset · contracts
The frozen forecast archive + matched station outcomes, licensed as training/eval data to anyone building weather AI. Grows more valuable every single day it runs; impossible for a late entrant to recreate.
Honest Risks (this plan's job is to be pressure-tested)
- Occupied ground: WeatherBench 2 owns academic mindshare; ECMWF publishes its own AI scores. Do NOT fight them — position as complementary: "WB2 for research metrics, TerraPulse for surface station-truth and extremes." If the community reads it as a clone, it dies.
- Methodology is genuinely hard: point-station vs. gridded-forecast comparison has real literature and real pitfalls (representativeness, station siting). Getting it wrong publicly burns the credibility being sold. Mitigation: publish methodology first, invite criticism before launching scores; recruit one academic advisor (an MSU meteorology contact would do double duty with the ag plan's Extension relationships).
- Forecast licensing: verify redistribution/scoring rights per feed (NWS public; ECMWF open data has terms; commercial AI outputs vary — scoring may be fine where republishing isn't; legal nuance for the attorney agenda).
- Will labs pay, or just cite the free board? Unknown. The buyer-side product (#4) is the hedge — those customers pay for decisions, not vanity.
Pricing Tiers
| Tier | Price | Includes |
|---|---|---|
| Scoreboard | Free | Public leaderboard, methodology, event bulletins |
| Bulletin sponsor | $1–5k/series | Clearly-labeled sponsorship of verification content |
| Model-maker report | $5–25k | Independent verification, citable, private pre-release option |
| Buyer selection report | $10–50k | Territory- and threshold-specific vendor scoring |
| Benchmark license | Contract | Frozen forecast + outcome archive for training/eval |
Garage Roadmap
- This month (costs one weekend + disk): start freezing forecast feeds NOW, before any product decision — the archive only grows forward.
- Write and publish the methodology as a Lab piece; invite the weather community to tear it apart. The criticism is free peer review and free marketing.
- Ship the scoreboard for ONE region and ONE variable first — Michigan surface temperature and freeze events — where local credibility, the ag product, and frost season all reinforce it. Expand variables and regions only after the method survives contact.
- Run bulletin #1 on the first big scored event; pitch it to weather Twitter/Bluesky and one trade publication.
- First paid report: offer it at cost to one friendly weather-AI startup in exchange for a citation. The citation is the product's birth certificate.
- Let #4 and #5 emerge from a year of the archive doing what archives do.
Appendix — Field Notes
What the landscape check found (Jul 2026)
The territory is partly occupied, in a way that sharpens the wedge rather than killing it. WeatherBench 2 — run by Google Research with DeepMind and ECMWF collaborators — is the established open-source evaluation framework, with a continuously updated public site tracking state-of-the-art models. A generic "we score AI weather models" pitch is therefore dead on arrival. But three cracks in the incumbent are exactly TerraPulse-shaped:
- Reanalysis truth. WB2 evaluates models against reanalysis datasets (ERA5) — models graded against other models' output. Scoring against raw station measurements is philosophically the entire TerraPulse opening.
- Extremes are the known weakness. The literature's consistent finding is that AI models struggle with extreme events — tropical-storm winds, extreme precipitation — precisely the events the severe-weather dexes capture and the ones with money attached.
- The referee fields a team. The dominant benchmark is operated by an organization that also builds the models being ranked (GraphCast, GenCast). Nobody can press that conflict-of-interest point as loudly as a shop that fields no team.
Why the roadmap starts tiny (and other strategic notes)
- The only real deadline is the archive. Every other product across all five planning documents can be built whenever; a forecast archive only grows forward. Freezing feeds costs disk space and one ingest job. The frozen forecasts-plus-matched-outcomes corpus is what products #3–5 eventually monetize and what no late entrant can ever recreate. Start this weekend regardless of every other decision.
- Complementary, not competitive. WB2 owns research metrics and academic mindshare; fighting it loses. The open lane is station-truth, surface variables, local decision thresholds, and extremes — which happen to be both the AI models' known weak spot and the only things the paying customers from Plans 1–2 care about.
- One region, one variable first. Michigan surface temperature and freeze events is where the methodology risk is containable and where the scoreboard, the ag product, and frost season all advertise each other. Expand only after the method survives contact.
- The risks section is load-bearing, not decoration. Point-station vs. gridded-forecast comparison is real science with real failure modes — hence methodology published for public criticism before any scores. If the method survives weather Twitter, this is a company-defining position; if it doesn't, the cost was a weekend and the lesson was cheap. That asymmetry is the whole reason this idea earned a one-pager.