For each analytical capability: what the model actually computes, which data feed it, how it has been validated, and β deliberately β its known limits.
The SSI is a two-stage pipeline. Stage 1 monitors every watched country for the presence of an acute shock and scores it on six weighted factors β scope, criticality, duration, elasticity, interconnectedness, and policy-response capacity. Scope and duration are derived live from the event itself (hazard alert level, conflict status, IPC phase); the remaining four come from per-shock-type analyst priors. Between the six-factor core and the published headline sits a bounded country channel β an economy-size logistic and an income-group tilt, squashed through an exponential soft-cap β and per-source ceilings (GDACS alert colour, IPC assessment age). A capped row publishes its uncapped value and the rule that bit (ssiUncapped / cappedBy), so the gap between factors and headline is checkable on the wire, never inferred. Stage 2 projects the economic impact of a detected shock: per-sector severities drive GDP, inflation, fiscal, trade, currency and employment components through sector elasticities and per-country baselines, with cross-sector interaction terms.
GDACS disaster alerts, a verified armed-conflict catalogue, IPC food-security assessments, WHO disease-outbreak notifications, and per-country macroeconomic baselines.
Detection is wired end-to-end for four shock types today β natural disasters, geopolitical conflict, food-security crises and disease outbreaks. The authoritative detected-versus-specified split is published live by the platform itself, so it can never go stale in this paragraph. Stage-2 projections are being calibrated against documented historical shocks β the empirical benchmark currently covers earthquake, hurricane and flood classes only; until that calibration is signed, Stage-2 outputs are presented as projections, not measurements.
For a real earthquake with a published USGS ShakeMap, the engine intersects the actual intensity contours with administrative districts (geoBoundaries ADM2) server-side and reports per-district shaking plus population exposure by intensity band, using GHSL gridded population. For the window before a ShakeMap exists β and for hypothetical scenarios β it synthesises a footprint from an intensity-prediction equation, Allen, Wald & Worden (2012), extended with a finite-fault effective-distance correction built from Wells & Coppersmith (1994) rupture scaling and Thompson & Worden (2018) unknown-strike averaging.
USGS ShakeMap intensity contours and PAGER; GDACS event footprints; geoBoundaries ADM2 districts; GHSL population.
The finite-fault correction was fitted against the published USGS ShakeMap for the M7.1 Venezuela earthquake: mean contour error fell from 21.6 km to 8.8 km, and the corrected model recovers the innermost MMI 8 ring that a point-source form misses. District exposure bucketing was checked against USGS PAGER for the M7.5 Venezuela event: the engine's MMI 6+ exposed population matched PAGER's published figure (8.0M vs 8.0M).
Fatality brackets use the Jaiswal-Wald (2010) lognormal lethality form with a global vulnerability bracket β the span of the range is the uncertainty. They are labeled modeled, order-of-magnitude: a planning signal, never a prediction. When USGS publishes PAGER for an event, PAGER is authoritative β the internal bracket is shown alongside it, never blended with it, never as a competing headline.
At close zoom the globe estimates damage probability per building: Google Open Buildings footprints, the real ShakeMap ground-motion value sampled under each footprint, and one number per 150 m block: a single logistic rate model in ln(PGA), fitted on twelve Copernicus EMS activations, giving the expected share of that block which is damaged (fitted coefficients available to clients under agreement). Every building inside a block carries the block's number, because that is the resolution the evidence has. The two-part model this paragraph used to describe — a lognormal fragility curve plus a gradient-boosted ranking of buildings by footprint size, density and micro-terrain — was removed from the engine on 2 August 2026 and is no longer computed anywhere; the ranking it performed measures AUC 0.490, chance, once the undamaged class is drawn from the standing building stock. Where satellite damage assessments exist, each assessed building's status is displayed as a fact, not a probability β and inside those assessed areas, observed data always wins; the model fills only the gaps between them.
Google Open Buildings footprints; USGS ShakeMap ground-motion grids; Copernicus Emergency Management Service rapid-mapping gradings; radar-measured ground deformation (InSAR line-of-sight displacement) as a regional observed layer.
© European Union, 2026, Copernicus Emergency Management Service (© 2026 European Union), [EMSR884] — Venezuela Earthquake. Copernicus EMS rapid-mapping products are published under CC BY 4.0, which permits commercial use with attribution; the terms grant reproduction, distribution and adaptation with no non-commercial clause.
The fragility curve was re-fitted by maximum likelihood on three satellite-observed damage tiles β 46,564 buildings spanning peak ground accelerations of 0.25β0.84 g. An earlier coastal-amplification heuristic over-predicted observed damage roughly threefold and was removed: where the observed data contradicted the model, the observed data won.
Corrected 2 August 2026 β a claim withdrawn. This paragraph previously reported the ranking model at “AUC 0.65β0.80 within an imaged zone, β0.67 transferred to an unimaged area, above the bare fragility curve’s β0.58”. Retested against 12 Copernicus EMS activations β 401,250 buildings across 11 earthquakes β none of the three figures survives. The 0.67 was leave-one-tile-out inside a single earthquake, which is not transfer to a new event. The β0.58 corresponded to no validation run in the codebase. And once the undamaged class is drawn from the standing building stock rather than from an analyst’s selection, per-building skill is AUC 0.490 β indistinguishable from chance. Seven covariates were swept (footprint area, height, built density, neighbour count, spectral shape, construction epoch at 1 km and 100 m); each works inside one area with a coefficient specific to that area, and those coefficients disagree across areas.
What is measured, and what it is for. Shaking predicts the damage rate of a neighbourhood, not the fate of a building. Refitted on 2 August 2026 over 12 earthquakes and 579,357 graded buildings: at ~150 m cells, Spearman Ο +0.362 leave-one-EVENT-out (p <0.0001) β pooled across the held-out earthquakes. That is a ranking claim, and it is the only accuracy figure Watchfloor publishes for this model. Qualified 5 August 2026, and read this before using it. Holding a whole earthquake out is the right split. But the correlation is then computed over every held-out cell from every earthquake at once, so differences in LEVEL between events β a M7.8 damages more everywhere than a M6.0 β enter the number and count as skill at ordering neighbourhoods. Watchfloor never asks anyone to compare a block in TΓΌrkiye with a block in Peru; it answers within this earthquake, which blocks are worst. That is the within-event figure, and it has now been measured. Measured 6 August 2026. The twelve Copernicus EMS activations were rebuilt from scratch and the rebuild authenticates itself twice: it yields 401,250 graded buildings β the same count this paragraph has always cited β and reproduces the published pooled figure at +0.355 against +0.362. On that same rebuild, the within-event ordering is:
Which twelve β recorded 5 August 2026. The activations behind this figure were, until that date, written down nowhere in the codebase: the list survived only in working notes, so the figure could be rebuilt by its author and by nobody else. The list is now versioned with the engine and was verified against Copernicus rather than merely copied β all twelve activation pages describe an earthquake, and the two ends of the original scan range do not, which is the control that makes it a selection rather than a list of whatever answered. The file also carries the two traps that make a rebuild wrong rather than merely hard: nine of the twelve grade only damaged buildings, so their base rate is 100% by construction, and each activation must be paired to its own USGS event by date and area β pairing by name has previously matched Java to Sumatra 700 km away.
Which individual building fell remains chance β per-building AUC 0.490 once the undamaged class is drawn from the standing building stock rather than an analyst's selection β so Watchfloor shows individual buildings only where a satellite has actually looked. A 150 m cell holds a few dozen buildings, so what it observes is a noisy draw from the rate rather than a measurement of it: on held-out events, 88% of cells with a non-zero predicted rate saw no damage at all, which is what small samples do and is not evidence the rate is wrong.
A per-country nowcast of the 90-day trajectory of a food-insecurity proxy built from near-real-time phone-survey indicators (food-consumption and coping-strategy prevalence, normalised 0β100). The model is a three-part gradient-boosted ensemble exported to ONNX β a baseline regressor, a quantile model for uncertainty bounds, and a non-linear regime model β over 26 features computed from each country's own survey history: lags, rolling means and volatilities, trend, seasonality, and a lean-season flag.
HungerMap LIVE survey history (the only feature source β no media or macro features), plus crop-calendar lean-season timing.
On its held-out test set the model reports MAE 1.6 percentage points, direction accuracy 97.7%, and RΒ² 0.98. Read those honestly: the model is autoregressive β it extrapolates a country's own survey history β so these scores measure trajectory extrapolation, not shock anticipation.
The structural electoral forecaster is a deliberately simple, fully inspectable logistic model over 11 structural features. It scores the probability that a governing party loses its next election, trained on 1,407 competitive national elections across 128 countries since 1990 (competitive defined as V-Dem electoral-democracy β₯ 0.5). It is a calibrated risk rank, not a crystal ball β its mean predicted probability equals the historical base rate (0.544) by construction.
Strictly out-of-sample only: forward-chaining AUC 0.714 (train on the past, predict the next) and leave-one-country-out AUC 0.721 (train on 127 countries, predict the held-out one). No random cross-validation β it leaks.
Qualified 5 August 2026 β what those figures were measured on. Both describe the eleven-feature model. The shipped forecaster feeds it eight: gdppc, sii (State Integrity Index at election time) and traj (its 180-day change) are always imputed to the training median, and alt_hist and polc join them for countries outside the curated set. gdppc is imputed for a concrete reason: the model was trained on GDP-per-capita growth while the live series returns the level, and feeding the level drives the standardised value to minus infinity. Weighed on the shipped coefficients, the three always-imputed features carry 2.8% of the linear predictor's variance (7.2% including the two curated-only ones), against 90% for incumbent tenure alone β so this is a small correction, not a hidden one. But it is a correction: an accuracy measured on a predictor that is not the one shipped is an upper bound. Measured 5 August 2026, and no longer an estimate. The training set was rebuilt (1,427 competitive elections, 128 countries, 54.5% incumbent-loss base rate) and both validations re-run twice on the same rows: once with all eleven features real, once imputing gdppc, sii and traj at prediction time exactly as the shipped forecaster does. With eleven features the published figures reproduce exactly β 0.714 and 0.721. The shipped configuration measures forward-chaining AUC 0.693 and leave-one-country-out AUC 0.713, a cost of 0.021 and 0.008. Those are the figures that describe what Watchfloor actually predicts, and they are what this page now quotes.
And a provenance repair worth naming. Until that run, the two published AUCs were literals: the export script wrote them into the model artefact by hand and printed them as fixed text, and the “countries” count was an expression that always returned 128. Nothing recomputed them, so nothing would have noticed if they had drifted away from the data. Re-running showed the values were right β but a correct number without provenance stays correct only until something changes underneath it. The export no longer prints an accuracy it did not compute.
Every call runs the same loop: predict β lock β resolve β score. A forecast is locked as an immutable snapshot the moment it is made β the number Watchfloor is later graded on can never be retroactively edited. When the outcome is established, the entry is resolved and Brier-scored, misses included. Forecasts built from structural priors alone are flagged as such, and every entry carries its framing β "calibrated relative risk Β· watch flag" β rendered verbatim, never as "prediction". The track record is public, no sign-in required β read it on the front page, or take the raw feed: /api/ledger/track-record (JSON).
For each monitored country, six deterministic danger vectors β conflict, economy, currency, food, governance, regime β are computed from the live signals Watchfloor already scores, plus V-Dem regime type. The convergence count is simply how many vectors are elevated at once. There is no model in this loop: the stacking is arithmetic, so it can be audited by counting.
In a backtest over 4,841 country-years (1996β2024), countries with three or more simultaneously deteriorating vectors were 7.3Γ more likely to experience an irregular government change the following year than countries with one or none. Risk stacks monotonically with the count.
Watchfloor computes on top of public and licensed datasets, and the boundary of those datasets is the boundary of the product. Outside the monitored-country list, Watchfloor has no live tools β and says so rather than improvising.
On conflict: Watchfloor's conflict signal measures shifts in global media coverage. It is media-derived monitoring β labeled as such β and is not front-line reporting. No claim of ground truth is made where none exists.
Four rules bind every surface of the product. They are enforced in the code paths that produce the numbers, not just promised here.
Every modeled output is stamped as modeled, and order-of-magnitude estimates are labeled as such wherever they appear. A scenario can never be mistaken for an observation.
Wherever a satellite-observed assessment exists, it overrides the model inside its assessed area. Models fill gaps; they do not overrule evidence.
When an authoritative estimate exists β USGS PAGER for earthquake fatalities β it leads. Internal brackets are presented beside it as brackets, never blended into it.
When a feed is down or a value cannot be computed, the surface shows an honest empty state. Nothing is padded, backfilled, or invented to look complete.
Questions about a specific figure, or a limit we have not stated? admin@notamy.app β challenges to the methodology are welcome; they make it better.