Methodology

How the numbers are made.

For each analytical capability: what the model actually computes, which data feed it, how it has been validated, and β€” deliberately β€” its known limits.

rev 2026-08-29This page is versioned by the date of its last correction, so a citation points at a text that cannot change under it.
Traceability. Every figure on this page is traceable to the production codebase or a named validation run. Where no defensible figure exists, none is shown β€” a limit stated plainly beats a number invented confidently.
  1. 01 Shock Severity Index
  2. 02 Earthquake impact engine
  3. 03 Per-building damage
  4. 04 Food-security nowcast
  5. 05 Forecasts & the ledger
  6. 06 Convergence watch
  7. 07 Data boundaries & attribution
  8. 08 The honesty contract
01 Β· Shock severity

Shock Severity Index β€” detect, then quantify.

The SSI is a two-stage pipeline. Stage 1 monitors every watched country for the presence of an acute shock and scores it on six weighted factors β€” scope, criticality, duration, elasticity, interconnectedness, and policy-response capacity. Scope and duration are derived live from the event itself (hazard alert level, conflict status, IPC phase); the remaining four come from per-shock-type analyst priors. Between the six-factor core and the published headline sits a bounded country channel β€” an economy-size logistic and an income-group tilt, squashed through an exponential soft-cap β€” and per-source ceilings (GDACS alert colour, IPC assessment age). A capped row publishes its uncapped value and the rule that bit (ssiUncapped / cappedBy), so the gap between factors and headline is checkable on the wire, never inferred. Stage 2 projects the economic impact of a detected shock: per-sector severities drive GDP, inflation, fiscal, trade, currency and employment components through sector elasticities and per-country baselines, with cross-sector interaction terms.

Data sources

GDACS disaster alerts, a verified armed-conflict catalogue, IPC food-security assessments, WHO disease-outbreak notifications, and per-country macroeconomic baselines.

Validation

Detection is wired end-to-end for four shock types today β€” natural disasters, geopolitical conflict, food-security crises and disease outbreaks. The authoritative detected-versus-specified split is published live by the platform itself, so it can never go stale in this paragraph. Stage-2 projections are being calibrated against documented historical shocks β€” the empirical benchmark currently covers earthquake, hurricane and flood classes only; until that calibration is signed, Stage-2 outputs are presented as projections, not measurements.

Known limits
  • Several factor priors are provisional pending analyst sign-off. The SSI's constants are under active review; expect revisions.
  • The remaining specified shock types are not yet wired to live detectors; the platform's live coverage read is the authority on that split at any moment.
  • Stage 2 currently runs on deterministic elasticity means; formal uncertainty bands are a staged increment, not yet shipped.
  • The framing is economic and supply-chain by design β€” the SSI is not a humanitarian impact estimate.
02 Β· Seismic

Earthquake impact engine β€” real footprints first.

For a real earthquake with a published USGS ShakeMap, the engine intersects the actual intensity contours with administrative districts (geoBoundaries ADM2) server-side and reports per-district shaking plus population exposure by intensity band, using GHSL gridded population. For the window before a ShakeMap exists β€” and for hypothetical scenarios β€” it synthesises a footprint from an intensity-prediction equation, Allen, Wald & Worden (2012), extended with a finite-fault effective-distance correction built from Wells & Coppersmith (1994) rupture scaling and Thompson & Worden (2018) unknown-strike averaging.

Data sources

USGS ShakeMap intensity contours and PAGER; GDACS event footprints; geoBoundaries ADM2 districts; GHSL population.

Validation

The finite-fault correction was fitted against the published USGS ShakeMap for the M7.1 Venezuela earthquake: mean contour error fell from 21.6 km to 8.8 km, and the corrected model recovers the innermost MMI 8 ring that a point-source form misses. District exposure bucketing was checked against USGS PAGER for the M7.5 Venezuela event: the engine's MMI 6+ exposed population matched PAGER's published figure (8.0M vs 8.0M).

Casualty ranges

Fatality brackets use the Jaiswal-Wald (2010) lognormal lethality form with a global vulnerability bracket β€” the span of the range is the uncertainty. They are labeled modeled, order-of-magnitude: a planning signal, never a prediction. When USGS publishes PAGER for an event, PAGER is authoritative β€” the internal bracket is shown alongside it, never blended with it, never as a competing headline.

Known limits
  • The casualty model is not calibrated per country β€” it uses a global vulnerability bracket, which is why it reports a wide range rather than a point estimate.
  • Scenario footprints are idealised concentric contours from the intensity-prediction equation; a real published ShakeMap supersedes them the moment it exists.
  • Exposure counts derive from modeled gridded population, not census rolls; they are exposure to shaking, not affected-person counts.
  • The engine emits only what it actually computed β€” surfaces without a defensible value stay empty rather than showing a synthesised one.
03 Β· Building scale

Per-building damage β€” observed beats modeled.

At close zoom the globe estimates damage probability per building: Google Open Buildings footprints, the real ShakeMap ground-motion value sampled under each footprint, and one number per 150 m block: a single logistic rate model in ln(PGA), fitted on twelve Copernicus EMS activations, giving the expected share of that block which is damaged (fitted coefficients available to clients under agreement). Every building inside a block carries the block's number, because that is the resolution the evidence has. The two-part model this paragraph used to describe — a lognormal fragility curve plus a gradient-boosted ranking of buildings by footprint size, density and micro-terrain — was removed from the engine on 2 August 2026 and is no longer computed anywhere; the ranking it performed measures AUC 0.490, chance, once the undamaged class is drawn from the standing building stock. Where satellite damage assessments exist, each assessed building's status is displayed as a fact, not a probability β€” and inside those assessed areas, observed data always wins; the model fills only the gaps between them.

Data sources

Google Open Buildings footprints; USGS ShakeMap ground-motion grids; Copernicus Emergency Management Service rapid-mapping gradings; radar-measured ground deformation (InSAR line-of-sight displacement) as a regional observed layer.

© European Union, 2026, Copernicus Emergency Management Service (© 2026 European Union), [EMSR884] — Venezuela Earthquake. Copernicus EMS rapid-mapping products are published under CC BY 4.0, which permits commercial use with attribution; the terms grant reproduction, distribution and adaptation with no non-commercial clause.

Validation

The fragility curve was re-fitted by maximum likelihood on three satellite-observed damage tiles β€” 46,564 buildings spanning peak ground accelerations of 0.25–0.84 g. An earlier coastal-amplification heuristic over-predicted observed damage roughly threefold and was removed: where the observed data contradicted the model, the observed data won.

Corrected 2 August 2026 β€” a claim withdrawn. This paragraph previously reported the ranking model at “AUC 0.65–0.80 within an imaged zone, β‰ˆ0.67 transferred to an unimaged area, above the bare fragility curve’s β‰ˆ0.58”. Retested against 12 Copernicus EMS activations β€” 401,250 buildings across 11 earthquakes β€” none of the three figures survives. The 0.67 was leave-one-tile-out inside a single earthquake, which is not transfer to a new event. The β‰ˆ0.58 corresponded to no validation run in the codebase. And once the undamaged class is drawn from the standing building stock rather than from an analyst’s selection, per-building skill is AUC 0.490 β€” indistinguishable from chance. Seven covariates were swept (footprint area, height, built density, neighbour count, spectral shape, construction epoch at 1 km and 100 m); each works inside one area with a coefficient specific to that area, and those coefficients disagree across areas.

What is measured, and what it is for. Shaking predicts the damage rate of a neighbourhood, not the fate of a building. Refitted on 2 August 2026 over 12 earthquakes and 579,357 graded buildings: at ~150 m cells, Spearman ρ +0.362 leave-one-EVENT-out (p <0.0001) β€” pooled across the held-out earthquakes. That is a ranking claim, and it is the only accuracy figure Watchfloor publishes for this model. Qualified 5 August 2026, and read this before using it. Holding a whole earthquake out is the right split. But the correlation is then computed over every held-out cell from every earthquake at once, so differences in LEVEL between events β€” a M7.8 damages more everywhere than a M6.0 β€” enter the number and count as skill at ordering neighbourhoods. Watchfloor never asks anyone to compare a block in TΓΌrkiye with a block in Peru; it answers within this earthquake, which blocks are worst. That is the within-event figure, and it has now been measured. Measured 6 August 2026. The twelve Copernicus EMS activations were rebuilt from scratch and the rebuild authenticates itself twice: it yields 401,250 graded buildings β€” the same count this paragraph has always cited β€” and reproduces the published pooled figure at +0.355 against +0.362. On that same rebuild, the within-event ordering is:

  • +0.104 β€” mean across the eleven earthquakes that can be scored.
  • +0.211 β€” mean across the nine with at least fifty 150 m cells, which is the fairer read.
  • +0.687 to −0.051 β€” the observed range on those nine. Four of eleven are negative: in those earthquakes the model ordered neighbourhoods backwards.
Use the range, not the average. A buyer who reads “+0.21” pictures modest but dependable performance; what the model actually delivers is one strong ordering or none, and nothing measurable tells you in advance which earthquake you have. We tested three candidate explanations for the spread and none survives: number of cells (rank correlation +0.445, p = 0.17, and the second-largest event scores −0.051), share of cells with zero damage (84.7% on average, p = 0.87), and the spread of ground motion within the event (p = 0.74). And one more thing that must travel with the number. Nine of the twelve activations grade only damaged buildings, so the undamaged class is reconstructed as “every other building inside the same mapped area”. Where Copernicus mapped a small fraction of what the footprint layer sees — coverage runs from 0.3% to 282% across areas — that reconstruction is doing more work than the data. Both figures above inherit that assumption. That assumption has now been tested against real observed undamaged buildings. Measured 24 August 2026. Microsoft’s AI for Good Lab published on HDX (CC BY 4.0) a city-wide assessment of Pereira for the 10 August Colombia M7.4: 143,472 buildings with usable imagery, 752 damaged — 0.52%. Every damaged building was reviewed by a person (166 ground truth, 302 confirmed, 295 unsure) and 754 further model candidates were rejected at review. So for the first time the undamaged class is observed rather than reconstructed, and the reconstruction caveat above does not apply. Matched building-to-building against what Watchfloor predicts — 24,767 buildings, 99.7% matched at a median 2.4 m, across the fourteen 0.008° boxes holding most of the damage — the result is:
  • Level: right. Watchfloor predicts a ~0.65% damage rate for Pereira against 0.52% observed — agreement within about 1.2×.
  • Ordering, within this city: absent. Per building the ranking measures AUC 0.445; stratified within each box, 0.482. At the model’s own 150 m cells, restricted to the 359 cells that hold the ≥25 buildings the rate was fitted on, ρ −0.087. The quartile of cells the model called worst came in at 1.07% observed damage; the quartile it called safest, 1.43%.
This is a thirteenth activation, from an independent source, and it lands just past the negative end of the range above (−0.051). Read plainly: for this earthquake the model told you correctly how much of Pereira was damaged, and told you nothing about which parts. It is one city, and a city where shaking barely varies — only seven distinct intensity values across the whole sample — so it does not refute ordering skill between regions of a larger event. It does say that inside a city, the map is a level, not a ranking. Where an observed layer exists, Watchfloor now serves it in place of the model for exactly this reason.

Which twelve β€” recorded 5 August 2026. The activations behind this figure were, until that date, written down nowhere in the codebase: the list survived only in working notes, so the figure could be rebuilt by its author and by nobody else. The list is now versioned with the engine and was verified against Copernicus rather than merely copied β€” all twelve activation pages describe an earthquake, and the two ends of the original scan range do not, which is the control that makes it a selection rather than a list of whatever answered. The file also carries the two traps that make a rebuild wrong rather than merely hard: nine of the twelve grade only damaged buildings, so their base rate is 100% by construction, and each activation must be paired to its own USGS event by date and area β€” pairing by name has previously matched Java to Sumatra 700 km away.

Which individual building fell remains chance β€” per-building AUC 0.490 once the undamaged class is drawn from the standing building stock rather than an analyst's selection β€” so Watchfloor shows individual buildings only where a satellite has actually looked. A 150 m cell holds a few dozen buildings, so what it observes is a noisy draw from the rate rather than a measurement of it: on held-out events, 88% of cells with a non-zero predicted rate saw no damage at all, which is what small samples do and is not evidence the rate is wrong.

Known limits
  • Outside an observed assessment area, every per-building value is a probability, labeled indicative β€” never a statement about an individual building's condition.
  • Calibration rests on a small number of tiles from one event region, so transfer to different building stocks is the open question. The ≈0.67 new-area figure this line used to cite belonged to the ranking model withdrawn above and is not a property of what ships; the figure that does ship is the out-of-sample ordering, Spearman +0.362 — but that is pooled across held-out earthquakes, so level differences between events count as skill and it is an upper bound. Within a single earthquake — which is the only comparison Watchfloor asks anyone to make — the ordering is ρ +0.104 mean over the eleven scorable events (+0.211 over the nine with ≥50 cells), ranging +0.687 to −0.051, with four of eleven negative. Use the range, not the average.
  • Satellite assessments are themselves model outputs of their providers, with their own error; Copernicus gradings cover only affected buildings, so their percentages read "confirmed among assessed", not a town-wide rate.
04 Β· Food security

Food-security nowcast β€” ninety days ahead.

A per-country nowcast of the 90-day trajectory of a food-insecurity proxy built from near-real-time phone-survey indicators (food-consumption and coping-strategy prevalence, normalised 0–100). The model is a three-part gradient-boosted ensemble exported to ONNX β€” a baseline regressor, a quantile model for uncertainty bounds, and a non-linear regime model β€” over 26 features computed from each country's own survey history: lags, rolling means and volatilities, trend, seasonality, and a lean-season flag.

Data sources

HungerMap LIVE survey history (the only feature source β€” no media or macro features), plus crop-calendar lean-season timing.

Validation

On its held-out test set the model reports MAE 1.6 percentage points, direction accuracy 97.7%, and RΒ² 0.98. Read those honestly: the model is autoregressive β€” it extrapolates a country's own survey history β€” so these scores measure trajectory extrapolation, not shock anticipation.

Known limits
  • Inference is skipped for countries with under 90 days of stored survey history β€” the model never runs on synthetic inputs.
  • Coverage is bounded by where phone-survey collection operates; no surveys, no nowcast.
  • Being history-only, the nowcast is blind to exogenous shocks that have not yet moved the survey series.
  • Retired from scoring on 2 August 2026. The nowcast used to raise a country's food-security signal when it forecast deterioration. It no longer moves any published score β€” on the day it was cut it was adding +8.6 to Yemen and +12.1 to Nigeria, saturating both at 100, on inputs up to 1007 days old. The model still answers on explicit request, with its staleness guard intact; what was removed is its ability to change a number you see.
05 Β· Accountability

Structural forecasts and the forecast ledger.

The structural electoral forecaster is a deliberately simple, fully inspectable logistic model over 11 structural features. It scores the probability that a governing party loses its next election, trained on 1,407 competitive national elections across 128 countries since 1990 (competitive defined as V-Dem electoral-democracy β‰₯ 0.5). It is a calibrated risk rank, not a crystal ball β€” its mean predicted probability equals the historical base rate (0.544) by construction.

Validation

Strictly out-of-sample only: forward-chaining AUC 0.714 (train on the past, predict the next) and leave-one-country-out AUC 0.721 (train on 127 countries, predict the held-out one). No random cross-validation β€” it leaks.

Qualified 5 August 2026 β€” what those figures were measured on. Both describe the eleven-feature model. The shipped forecaster feeds it eight: gdppc, sii (State Integrity Index at election time) and traj (its 180-day change) are always imputed to the training median, and alt_hist and polc join them for countries outside the curated set. gdppc is imputed for a concrete reason: the model was trained on GDP-per-capita growth while the live series returns the level, and feeding the level drives the standardised value to minus infinity. Weighed on the shipped coefficients, the three always-imputed features carry 2.8% of the linear predictor's variance (7.2% including the two curated-only ones), against 90% for incumbent tenure alone β€” so this is a small correction, not a hidden one. But it is a correction: an accuracy measured on a predictor that is not the one shipped is an upper bound. Measured 5 August 2026, and no longer an estimate. The training set was rebuilt (1,427 competitive elections, 128 countries, 54.5% incumbent-loss base rate) and both validations re-run twice on the same rows: once with all eleven features real, once imputing gdppc, sii and traj at prediction time exactly as the shipped forecaster does. With eleven features the published figures reproduce exactly β€” 0.714 and 0.721. The shipped configuration measures forward-chaining AUC 0.693 and leave-one-country-out AUC 0.713, a cost of 0.021 and 0.008. Those are the figures that describe what Watchfloor actually predicts, and they are what this page now quotes.

And a provenance repair worth naming. Until that run, the two published AUCs were literals: the export script wrote them into the model artefact by hand and printed them as fixed text, and the “countries” count was an expression that always returned 128. Nothing recomputed them, so nothing would have noticed if they had drifted away from the data. Re-running showed the values were right β€” but a correct number without provenance stays correct only until something changes underneath it. The export no longer prints an accuracy it did not compute.

The ledger

Every call runs the same loop: predict β†’ lock β†’ resolve β†’ score. A forecast is locked as an immutable snapshot the moment it is made β€” the number Watchfloor is later graded on can never be retroactively edited. When the outcome is established, the entry is resolved and Brier-scored, misses included. Forecasts built from structural priors alone are flagged as such, and every entry carries its framing β€” "calibrated relative risk Β· watch flag" β€” rendered verbatim, never as "prediction". The track record is public, no sign-in required β€” read it on the front page, or take the raw feed: /api/ledger/track-record (JSON).

Known limits
  • AUC β‰ˆ 0.71 is honest skill for structural signals β€” materially better than chance, far from certainty. Treat outputs as ranked watch flags.
  • Structural features move slowly; the model does not see fast idiosyncratic events. Narrative signals are layered separately and flagged as such.
  • The ledger is young. The resolved-forecast count and the scores live on the public feed itself β€” this page states none, so it can never overstate them.
06 Β· Convergence

Convergence watch β€” a watch, not a prediction.

For each monitored country, six deterministic danger vectors β€” conflict, economy, currency, food, governance, regime β€” are computed from the live signals Watchfloor already scores, plus V-Dem regime type. The convergence count is simply how many vectors are elevated at once. There is no model in this loop: the stacking is arithmetic, so it can be audited by counting.

Grounding

In a backtest over 4,841 country-years (1996–2024), countries with three or more simultaneously deteriorating vectors were 7.3Γ— more likely to experience an irregular government change the following year than countries with one or none. Risk stacks monotonically with the count.

Known limits
  • The lift comes from additively stacking independent signals β€” the backtest found no non-linear "convergence magic", and we do not claim one.
  • Irregular government change is a rare event: even at high convergence, most country-years see no such change. The flag marks a danger window, nothing more.
  • Each vector inherits the limits of its underlying signal β€” including the media-derived nature of the conflict signal (see below).
07 Β· Data

Data boundaries and attribution.

Watchfloor computes on top of public and licensed datasets, and the boundary of those datasets is the boundary of the product. Outside the monitored-country list, Watchfloor has no live tools β€” and says so rather than improvising.

Administrative boundaries
geoBoundaries (ADM2), William & Mary geoLab β€” CC-BY. Used exclusively for district geometry.
Population
GHSL β€” Global Human Settlement Layer, European Commission JRC.
Building footprints
Open Buildings, Google Research.
Hazards
USGS (ShakeMap, PAGER, secondary hazards), GDACS, NOAA (tropical cyclones), Copernicus EMS (rapid mapping, InSAR ground deformation).
Food security
IPC analyses; HungerMap LIVE survey indicators; World Bank indicators.
Governance & regime
V-Dem, Freedom House, and curated analyst catalogues reviewed against public watchlists.
Economy & climate
Frankfurter exchange rates, Open-Meteo precipitation, satellite vegetation indices.
Media signal
GDELT global media monitoring β€” article volumes and deviation from baseline.

On conflict: Watchfloor's conflict signal measures shifts in global media coverage. It is media-derived monitoring β€” labeled as such β€” and is not front-line reporting. No claim of ground truth is made where none exists.

08 Β· The contract

The honesty contract.

Four rules bind every surface of the product. They are enforced in the code paths that produce the numbers, not just promised here.

Modeled is never measured

Every modeled output is stamped as modeled, and order-of-magnitude estimates are labeled as such wherever they appear. A scenario can never be mistaken for an observation.

Observed wins inside its area

Wherever a satellite-observed assessment exists, it overrides the model inside its assessed area. Models fill gaps; they do not overrule evidence.

Authoritative sources lead

When an authoritative estimate exists β€” USGS PAGER for earthquake fatalities β€” it leads. Internal brackets are presented beside it as brackets, never blended into it.

Numbers are real or absent

When a feed is down or a value cannot be computed, the surface shows an honest empty state. Nothing is padded, backfilled, or invented to look complete.

Questions about a specific figure, or a limit we have not stated? admin@notamy.app β€” challenges to the methodology are welcome; they make it better.