geo/locations.csv.docker compose -f docker/docker-compose.yml up), open a disaster project, and copy its damage-layer tile URL. See HASTE_SETUP.md.Live Updates & Conflict
Per-crisis real-time developments (Tavily news) and the structured conflict timeline (ACLED) — the same feed as the map drawer, browsable here without the map. Every item links to its primary source.
Methodology
The Evacuation Inform Index (EII) is structured as a ratio — the Risk Score for Evacuating (RSE) divided by the Risk Score for Staying (RSS) — synthesising the INFORM Severity 3-dimension structure with the IOM RICD macro/micro split.
CERAI lens — endangerment vs feasibility
The CERAI v3 framework makes one architectural argument: never collapse danger and feasibility into a single score. Under IHL the obligation to protect civilians flows from danger (GC IV Art. 49), not from operational feasibility — low feasibility does not extinguish the obligation, it intensifies the urgency of political engagement (Mariupol 2022 is the proof-of-concept). This index renders that split live, per crisis, on the 🔴 Live & Conflict tab:
| CERAI dimension | What it answers | How we compute it here |
|---|---|---|
| Endangerment (Threat Environment) | How dangerous is it to stay? | INFORM Conditions (RSS) → 0–100%, with the 75% IHL Obligation Threshold marked; trajectory from live ACLED fatalities |
| Feasibility (can people move?) | Can civilians realistically move? | INFORM Complexity (RSE), inverted → 0–100%, reduced by live Open-Meteo route weather |
| Protection gap flag | Where is escalation needed? | High endangerment (≥75%) + low feasibility (≤40%) → flags need for political, not operational, escalation |
| CERAI dimension | What it answers | How we compute it here |
|---|---|---|
| Vulnerability profile (Dimension 3) | Who is exposed? | Interactive demographic toggles set a 0.7×–1.3× multiplier that amplifies endangerment and reduces feasibility; the 0.7×–1.3× span is shown as the endangerment range across the population |
| Damage assessment (satellite) | What is physically destroyed? | Microsoft HASTE + Planet AI damage maps overlaid on the Map tab — building loss raises endangerment (infrastructure availability), blocked routes lower feasibility |
Definitions — what this index counts as "conflict"
The word does three different jobs here, and collapsing them is the most likely way to misread a score. They are kept apart deliberately.
| Sense | Definition used | What it means for the score |
|---|---|---|
| 1 · Conflict as a counted event operational | ACLED's definition: a dated, geolocated, sourced incident in one of six categories — battles, explosions/remote violence, violence against civilians, riots, protests, strategic developments. The query is country-wide and unfiltered; the timeline shows the type mix. | Only monthly fatalities enter the score (the trajectory). So: a month of mass arrests, forced relocation, checkpoint closures or looting with no deaths reads as a quiet month; fatalities measure lethality, not danger to a civilian not yet killed; and the ACLED tier used here carries a ~12-month embargo, so "recent 3 months" can mean recent as of a year ago (the cutoff date is shown). |
| 2 · Conflict as a crisis driver what's on the map | INFORM Severity's driver labels. Of the 104 crises carried here: 39 Conflict/Violence, 50 International Displacement, 25 Floods, 22 Drought, 18 Political/economic crisis, 9 Cyclone, 1 Earthquake — most crises carry several, so the counts sum to more than 104. | All 104 are scored with the same formula. The index does not restrict itself to armed conflict and does not change its arithmetic when the driver is a cyclone. The driver label is displayed but never enters the calculation. |
| 3 · Conflict as a legal classification what sets the obligations | The IHL categories below — IAC, occupation, NIAC, and other situations of violence. | Not modelled. The index has no classification field and INFORM supplies none. See the caveat under the table. |
Sense 1 · The six ACLED event types — and which of them can move the score
Real volumes: 700,135 events aggregated from the 96 of 104 crises that carry an ACLED timeline. The query is unfiltered, so all six types reach the timeline a user reads. The score reads one field only — monthly fatalities.
250,130 events — 36% of everything recorded — are non-violent by ACLED's own definition. A protest is defined as a non-violent demonstration (one that turns violent is recoded as a riot); strategic developments are explicitly contextual events such as arrests, agreements and looting. Both fill the timeline the user reads and contribute essentially nothing to the number the model computes.
Table view — ACLED event types
| ACLED event type | Events | Share | How it reaches the score |
|---|---|---|---|
| Explosions / Remote violence | 185,992 | 26.6% | Violent — fatalities expected |
| Protests | 183,943 | 26.3% | Non-violent by ACLED definition — effectively invisible to the score |
| Battles | 127,449 | 18.2% | Violent — fatalities expected |
| Violence against civilians | 98,154 | 14.0% | Violent — fatalities expected |
| Strategic developments | 66,187 | 9.5% | Non-violent by ACLED definition — effectively invisible to the score |
| Riots | 38,410 | 5.5% | Violent — fatalities variable |
| Total | 700,135 | 100% | 96 of 104 crises carry an ACLED timeline |
Colour encodes how an event type reaches the score, not its volume — bar length already carries volume. Caveat on this figure's own evidence: the aggregation in server.py counts events per type but does not break fatalities down by type, so the three classes are read from ACLED's event-type definitions rather than from measured per-type fatality data. Adding fatalities to that aggregation would let this figure be drawn from measurement instead of definition.
Sense 2 · INFORM crisis drivers — 39 of 104, one formula for all
The driver labels attached to each of the 104 crises on the map, prefix-normalised (the exported drivers strings are truncated mid-word in the source data).
Only 39 of the 104 crises carry a Conflict/Violence label — and all 104 are scored by the identical formula. The driver is displayed but never enters the arithmetic: a drought's endangerment score is constructed exactly the way a war's is. Whether that is a strength (one comparable scale) or a flaw (a category error) is the open question this figure is here to put in front of you.
Table view — INFORM crisis drivers
| Driver | Crises | Share of 104 |
|---|---|---|
| International Displacement | 50 | 48% |
| Conflict / Violence | 39 | 38% |
| Floods | 25 | 24% |
| Drought | 22 | 21% |
| Political / economic crisis | 18 | 17% |
| Cyclone | 9 | 9% |
| Earthquake | 1 | 1% |
Crises carry more than one driver, so the column sums to more than 104.
| Classification | Trigger | Consequence for evacuation |
|---|---|---|
| International armed conflict | Resort to armed force between States (Common Art. 2, GC I–IV) — no intensity threshold | Full GC IV protections; Art. 49 applies where the situation is also one of occupation |
| Belligerent occupation | Territory placed under the authority of a hostile army (Hague Regs Art. 42) | The only situation where the Art. 49 evacuation regime applies in terms: permitted for the security of the population or imperative military reasons, with return, accommodation and family-unity obligations attached |
| Non-international armed conflict | Organised armed groups + protracted armed violence (Common Art. 3; AP II where its conditions are met; ICTY Tadić 1995, §70) | Forced displacement of civilians prohibited by AP II Art. 17 unless their security or imperative military reasons demand it |
| Other situations of violence | Internal disturbance, riot, gang/criminal violence — below the armed-conflict threshold | IHL does not apply. Human rights law governs, with the Guiding Principles on Internal Displacement |
Sense 3 · The legal classification — carried by nothing
The classification above is what determines which obligations attach to a high endangerment score. The index records none of it: there is no field for it and INFORM supplies none.
0 of 104 crises are classified, yet the 75% marker is drawn on every one of them. That marker comes from GC IV Art. 49, which sits in the Convention's occupied territory section and binds an Occupying Power — the second box, and only the second box. On a crisis of gang violence the marker refers to a rule that does not apply at all. Carrying a classification field is the highest-value single addition a future version could make.
This figure is a schematic of an absence, so its table view is the classification table directly above it.
Three-layer architecture
▸ Click a layer to see its variables & proposed weights.
| Layer | Function | Precedent |
|---|---|---|
| Layer 1 — Objective Risk Score (ORS) variables → 50% | Universal factors, identical for everyone in the geography (hostilities, conflict risk, natural-hazard life risk) | INFORM Severity, ACLED, GCRI |
| Layer 2 — Infrastructure & Access (IAS) variables → 35% | Availability of evacuation routes, resources & connectivity | ACAPS Humanitarian Access, IDMC |
| Layer 3 — Personal Vulnerability Modifier (PVM) variables → 15% | Demographic / household factors applied as a multiplicative modifier (0.7×–1.3×) | CDC SVI, IOM RICD micro-level |
Aggregation — weighted geometric mean
Indicators aggregate by weighted geometric mean rather than arithmetic mean, so an extreme imbalance (e.g. all infrastructure unavailable) cannot be compensated by a low score elsewhere — the same logic used by the Human Development Index and INFORM Risk.
Recommended build sequence
| Phase | Method | Purpose |
|---|---|---|
| 1 · Variable design | Delphi + Budget Allocation (8–12 experts, 2 rounds) | Set initial layer & sub-variable weights |
| 2 · Weight validation | Fuzzy AHP (triangular numbers, Buckley's geometric mean) | Validate contested weights under uncertainty |
| 3 · Calibration | Historical case testing (Sudan '23, Ukraine '22, Kabul '21, Lebanon '06, Haiti '10) + PCA | Check the index matches real decisions; prune redundant variables |
Constraint layers (not scored)
Financial feasibility filter — a second-stage check on whether the recommended action is affordable (transport, accommodation, asset-liquidation loss, income disruption). Legal / rights (UDHR Art. 13) — flags exit-visa requirements, travel bans or closure orders that restrict self-evacuation.
Intended use — and what this is not fit for
| Built to answer | Not fit for |
|---|---|
|
• Across a portfolio of active crises, where does the risk of staying diverge most sharply from the risk of leaving? • Where does high endangerment coincide with low feasibility — i.e. where has the problem passed beyond operational reach and become political? • Which crises are deteriorating on recent conflict evidence rather than on reputation or news volume? • What does an explicitly two-sided, non-compensatory evacuation model look like when actually built? (method demonstration / teaching artefact) Readers: analysts and advocacy staff comparing crises, researchers, students of humanitarian method. Not field operations. Not affected people. |
• Advising an individual or household whether to leave. Scores are population-level; the vulnerability profile is a scenario you set, not a record of anyone — see the subgroup limitations below. • Operational go/no-go on a convoy, corridor or movement window — there is no corridor state, checkpoint state or ceasefire clock in the model. • Route selection or timing — there is no route geometry at all; the road-access signal is keyword-derived from headlines and capped at 12 pts by design. • Ranking who evacuates first — the tool does not prioritise populations and has no defensible basis to. • Any determination of a person's legal status — asylum, visa, protection claim, eligibility. • Justifying a restriction on movement. EII > 1.0 records that the model scored evacuation as riskier than staying. It is not a finding that anyone should be prevented from leaving. Freedom of movement (UDHR Art. 13; ICCPR Art. 12) is not conditioned on a risk model's output. Using this index to support a closure order, exit ban or refusal of passage inverts its purpose — the tool already treats such restrictions as constraints on evacuation, above. |
Variables & Proposed Weights
Preliminary weights synthesised from INFORM, ACLED and FSI logic — to be validated by AHP expert surveys.
Layer 1 — Objective Risk Score · 50%
| Variable | Data source | Weight | Rationale |
|---|---|---|---|
| Active hostilities | ACLED: fatalities, attack types, proximity | 20% | Immediate life threat; fastest-changing → heaviest weight |
| Likelihood of future hostility | GCRI risk score; FSI security; ICEWS | 15% | Forward-looking; less certain than observed events |
| Life risk (non-conflict) | IDMC, FEWS NET — flood, quake, fire | 15% | Natural-hazard exposure alongside conflict |
Layer 2 — Infrastructure & Access · 35%
| Variable | Data source | Weight | Rationale |
|---|---|---|---|
| Evacuation route availability | Flight seats, road/border status, satellite imagery | 12% | No route = evacuation impossible |
| Infrastructure availability (stay) | Internet, energy, food, water (IPC, FEWS NET) | 10% | Determines survivability if staying |
| Security / threat alerts | OSAC, embassy alerts, local-language news | 8% | Near-real-time signal, both directions |
| Weather | NOAA, Copernicus | 5% | Modifier on route viability & shelter |
Layer 3 — Personal Vulnerability Modifier · 15% (multiplicative 0.7×–1.3×)
| Variable | Operationalisation | Direction of effect |
|---|---|---|
| Young children (<12) | CDC SVI "age ≤17"; self-report | ↑ RSE (harder to move) & ↑ RSS (more vulnerable) |
| Elderly (65+) | CDC SVI "age 65+"; self-report | ↑ RSE (mobility) & ↑ RSS (medical risk) |
| Gender / gendered risk | UNHCR GBV risk indicators | ↑ RSS in conflict zones with GBV risk |
| Prior evacuation experience | Self-assessed preparedness | ↓ RSE (more capable evacuee) |
| Financial resources | Self-reported | ↓ RSE when high; ↑ RSS when low |
Weighting methods considered
| Method | Subjectivity | Data | Defensibility | Use |
|---|---|---|---|---|
| Equal weights | None | None | Low | Baseline & sensitivity |
| Budget allocation (BAP) | High | None | Medium | Rapid prototyping |
| AHP | Medium | Expert survey | High | Published index |
| Fuzzy AHP | Low–Med | Expert survey | Very high | Ambiguous variables |
| PCA / factor analysis | None | Historical | Medium | Validation & pruning |
Variables integrated from ETC evacuation projects
Variables and structures mapped across sibling Ethical Tech CoLab evacuation repos, folded in here (or on the roadmap) to enrich the model beyond a demographic-only vulnerability list.
| Contribution | From repo | Status here |
|---|---|---|
| Protection-based vulnerable groups — wounded/acutely sick, pregnant & new mothers, unaccompanied/separated minors, undocumented / ID-gap persons, targeted ethnic·religious·political minorities, linguistic minorities / low literacy (each with an IHL basis, e.g. AP I Arts 16, 78; GC IV Art 23; customary IHL Rules 98–99) | Evac-Sim-Melanie | Added to the Dimension-3 profile (Live tab). These are protection/legal vulnerabilities, not just mobility ones — several raise endangerment via targeting/detention rather than slowing movement. |
| Destination-readiness gatekeepers — Security, Authority consent, host Willingness, Capacity, Shelter, Food/water, Medical capacity; a confirmed host refusal hard-caps readiness | India-EvacSimulation | Roadmap — feasibility currently uses a single INFORM-Complexity score; gatekeeper caps would replace it. |
| Seven-dimension model (D1–D7) + CERAI cross-derivation (Endangerment = d1·0.45 + d2·0.20 + d6·0.35; Feasibility inverts d3,d4,d5,d7), NATO STANAG level binning | ercf · Exodus | Roadmap — a concrete recipe to decompose Endangerment/Feasibility into scored sub-dimensions. |
| Corridor / checkpoint dynamics — open/closed exit gates, ceasefire windows, congestion queues, siege "trapped" state, information-environment degradation & misinformation | Evac-Sim-Melanie | Roadmap — the static Endangerment/Feasibility pair has no time-varying corridor or information dimension yet. |
| Historical calibration harness — differential-evolution fit to 16 in-scope cases (R²=0.855, LOOCV 0.807), with documented out-of-scope failure modes | ercf | Informs the limitations below; the model-boundary honesty is the borrowed practice. |
| Non-compensatory geometric-mean aggregation + RSS floor of 0.5 | evacmodel | Already used — see Aggregation above. |
Evidence provenance & traceability
The lab's lineage in supply-chain traceability and forced-labor mitigation (director Yorke Rhodes, Microsoft) shapes how this index treats evidence. Patterns adapted from that work and two public forced-labor models:
| Principle | Adapted from | Effect on the methodology |
|---|---|---|
| Two-witness evidentiary standard — separate a verified fact from ≥2 independent corroborating reports from a single unverified one | GFEMS FLARE | Sharpens the source-credibility tiers (UN-verified 1.0× → unverified 0.7×) into an evidentiary ladder. |
| Deterministic score, not an AI score — the confidence/risk number is a fixed, auditable formula; the language model only supplies source-grounded facts | provenance-search · arts-provenance-agent | The INFORM 0–10 mapping and the D3 multiplier stay reconstructable and never overridden by a model. |
| Chain-of-custody: an evidence gap is a risk signal | arts-provenance-agent | An unverified corridor segment is penalised, not assumed safe — mirroring provenance-gap logic. |
| Labeled graceful degradation — a fallback to background/model knowledge is tagged, capped below "verified," and auto-flagged | provenance-search | Absence of a live verified source never reads as verified. |
| Behaviorally-grounded indicator scoring — risk from a weighted set of observable indicators, not one metric | Global Fishing Watch forcedlabor | Precedent for the multi-indicator endangerment/feasibility structure. |
| Decision-support, not a verdict — a human retains the high-stakes call | FLARE ("a decision support tool, not an executioner") | Governance stance stated explicitly (below). |
Research limitations
- Proxy construct. Endangerment and feasibility are derived from INFORM Conditions and Complexity sub-scores (plus live ACLED/weather), not CERAI's full 22-variable IHL engine. Treat them as a faithful architectural proxy, not a validated instrument.
- Researcher-assigned weights. All weights are best estimates pending expert validation (Delphi → Fuzzy AHP) — at the same evidentiary level as INFORM's initial weights, but not yet consensus-tested.
- No ground-truth calibration. Independent ground truth for evacuation decisions does not exist. Sibling calibration (ercf: 16 cases, R²=0.855) is face validity, not statistical generalisation, and is explicitly out of scope for genocide, large-enclave precision operations, and sieges beyond ~90 days.
- Population-level, illustrative vulnerability. The Dimension-3 profile is a user-set demographic scenario applied as a 0.7×–1.3× multiplier — not measured household data — and cannot capture individual circumstances.
- Static snapshot. The Endangerment/Feasibility pair has no corridor/checkpoint dynamics, ceasefire windows, information-environment degradation, or misinformation (all modelled in Evac-Sim-Melanie, not here). Hosted news/ACLED are captured snapshots; weather is live.
- No political-will modelling. The index cannot model actor behaviour, negotiation status, sudden shifts in belligerent intent, or consent dynamics.
- Evidence quality varies. Source-credibility tiering is a design principle; live inputs (ACLED, Tavily) differ in verification and are not yet weighted by it in code.
- Satellite & damage layers are optional. The HASTE damage overlay requires a self-hosted deployment; without it, physical-damage evidence is absent.
- Correlation, not causation. The index prioritises attention; it is decision-support and must not be the sole basis for an evacuation decision.
- Conflict is defined three ways, only two of which are modelled. See Definitions: the score counts fatalities only, the driver label never enters the arithmetic, and the legal classification (IAC / occupation / NIAC / other situations of violence) is not carried at all — so the 75% Art. 49 marker is drawn on crises where that article does not apply.
Limits specific to vulnerable subgroups
The Dimension-3 profile is the part of this tool most likely to be read as saying something about a particular person, and the part least able to. Twelve toggles each move one multiplier by ±0.06 within a 0.7–1.3 band. Every property of that construction is a limitation:
- Equal increments assert an equivalence nobody established. Being non-ambulatory and being a linguistic minority move the score by the same 0.06. No evidence supports that parity — it is a placeholder chosen to demonstrate the mechanism, not yet replaced.
- The band saturates at five factors. Five upward toggles hit the 1.3 ceiling; the 6th–10th change nothing. The households carrying the most compounded vulnerability — an elderly, disabled, undocumented, non-literate member of a targeted minority — are exactly where the model stops discriminating between cases.
- A multiplier cannot express impossibility. For some conditions evacuation is not harder but foreclosed: a non-ambulatory person with no vehicle, a woman in obstructed labour, a dialysis patient on a 3-day interval, a ventilated patient without power. These need a hard cap on feasibility — the kind the roadmap's destination-readiness gatekeepers apply to hosts, but nothing applies on behalf of a person. The law recognises the category even where the index does not: GC IV Art. 17 provides for local agreements to remove the wounded, sick, infirm, aged, children and maternity cases from besieged areas, precisely because ordinary movement is unavailable to them.
- Factors are treated as independent when they are correlated and interacting. Elderly / disabled-or-medically-dependent / wounded overlap heavily in any real population — adding 0.06 each counts one underlying condition up to three times. The error runs both ways: pregnancy plus no functioning obstetric facility is worse than the sum of the terms, and an additive form cannot represent it.
- The two pathways are forced to mirror each other. The code raises endangerment by
mand lowers feasibility by(2−m)— same magnitude, opposite sign. That is wrong for most of the twelve. Being targeted for one's ethnicity multiplies the danger of remaining while leaving mobility untouched; late-term pregnancy does close to the reverse; a wheelchair user on a flooded road suffers a feasibility collapse with no matching jump in the danger of staying. Each factor needs two coefficients, one per pathway. - Binary toggles discard the severity that decides the outcome. "Elderly (65+)" spans an independent 66-year-old and a bedbound 92-year-old. "Pregnant" spans the first trimester and the 39th week — states differing by an order of magnitude for both danger and movement. "Disabled" covers a controlled chronic condition and total dependence on assistive equipment and a carer.
- The household is the evacuating unit, not the individual. Families move at the pace of their least mobile member and frequently refuse to separate — a refusal the law supports (Art. 49 requires that members of the same family not be separated). Individual attribute toggles cannot represent one immobile member immobilising a household of eight, nor that the alternative is a separation IHL discourages.
- Care relationships are absent. Both dependent and carer are constrained; only the dependent is on the list. There is a toggle for unaccompanied minors — none for the adult whose evacuation is constrained by three children and a parent with dementia.
- No prevalence, therefore no caseload. The profile answers "how would this scenario shift the score", never "how many people in this crisis are in it". Planning figures exist and could be used — WHO estimates ~16% of the global population lives with significant disability, and inter-agency reproductive-health (MISP) planning commonly assumes ~4% of a crisis-affected population is pregnant at any time — the index carries neither. Without prevalence the multiplier describes a hypothetical person, not the population the crisis score is about.
- The legal basis is cited but not operative. The protection-based groups were added because each has a footing in law — AP I Art. 8(a) classes maternity cases, newborns, the infirm and expectant mothers as "wounded and sick"; customary IHL Rule 138 entitles the elderly, disabled and infirm to special respect and protection (Rules 134–135 for women and children); CRPD Art. 11 covers persons with disabilities in situations of risk. None of this changes the arithmetic. A strong legal footing has not produced a strong weight, and should not be read as having done so.
The law this index cites — in plain language
This page rests on 13 legal provisions, cited 22 times between them. Every citation above is clickable — the dotted underline marks them — and each opens the same explainer. No legal training is assumed.
Data Sources
Feeds that populate the index. All free unless noted.
Reference Indices & Key Papers
Methodological precedents
INFORM Severity Index
3 weighted dimensions (Impact 20% · Conditions 50% · Complexity 30%), 1–5 scale. Closest analogue.
acaps.org →ACLED Conflict Index
Deadliness 35 · Danger to civilians 25 · Diffusion 20 · Fragmentation 20; non-linear root aggregation.
acleddata.com →IOM RICD
Two-tier macro (spatial risk) + micro (community) model — the structural precedent for base score × personal modifier.
iom.int →IDMC Risk Model 2.0
Probabilistic displacement from natural hazards — feeds the "risk of staying" dimension.
internal-displacement.org →Fragile States Index
12 indicators, CAST text-analysis triangulation — template for structural / minor variables.
fragilestatesindex.org →CDC Social Vulnerability Index
16 variables, 4 equally-weighted themes, percentile ranking — the personal-vulnerability layer.
atsdr.cdc.gov →Global Conflict Risk Index
Conflict-onset risk — separates active conflict (ACLED) from forward-looking risk.
jrc.ec.europa.eu →OECD/JRC Composite Handbook
Definitive reference for normalisation, aggregation, weighting & sensitivity analysis.
publications.jrc.ec.europa.eu →Microsoft HASTE
AI framework turning satellite imagery into building/route damage maps (Azure Maps, Batch GPU, ML). No public API — deploy it, then overlay its damage tiles on the Map tab. Forked under this org for deployment.
aka.ms/HASTE → · our fork →Key papers
- Beccari, B. (2016). A Comparative Analysis of Disaster Risk, Vulnerability and Resilience Composite Indicators. PLoS Currents Disasters, PMC4807925 — reviews 106 index methodologies.
- Can severity of a humanitarian crisis be quantified? Assessment of the INFORM severity index. Globalization & Health (2023). DOI:10.1186/s12992-023-00907-y — identifies governance & access as strongest predictors.
- Al Fozaie (2022). A Guide to Integrating Expert Opinion and Fuzzy AHP When Generating Weights for Composite Indices. Advances in Fuzzy Systems. DOI:10.1155/2022/3396862.
- OECD/JRC (2008). Handbook on Constructing Composite Indicators. OECD Publishing.
- Saaty, T.L. (1990). How to Make a Decision: The Analytic Hierarchy Process. EJOR 48(1), 9–26.
- ACAPS Ukraine Severity Model Methodology Note (March 2024) — worked subnational example.