Enter scenario data to generate situation summary.
Enter data to generate gap analysis.
| Case ID | Operation | Type | Evacuated | Ground Truth Risk | Outcome Summary |
|---|
CERAI separates assessment into three conceptually distinct dimensions. The Endangerment Assessment (Dimension 1) asks: how dangerous is it to remain? The Feasibility Assessment (Dimension 2) asks: is an organised evacuation currently possible? The Decision Support Briefing asks: what does this mean for humanitarian decision-making and IHL obligations? This separation is deliberate. Under IHL, the obligation to facilitate evacuation arises from endangerment (AP I Arts. 57-58, GCIV Art. 49), not from operational feasibility. Feasibility constrains the timing and modality of evacuation but not the legal obligation itself. A high endangerment + low feasibility scenario does not extinguish the obligation; it intensifies the urgency of political engagement to create feasibility.
The central methodological contribution of CERAI is the deliberate separation of endangerment from feasibility. Most existing humanitarian indices conflate the two: INFORM Severity folds operational access into its complexity dimension, and JIAF rolls humanitarian access into the conditions analysis. CERAI takes a different position: the IHL obligation to consider evacuation flows from danger (GC IV Art. 49, AP I Art. 58), while the lawful execution of evacuation depends on a separate set of procedural and operational conditions (AP I Art. 57(2)(c) effective warning, AP II Art. 17, ICRC consent doctrine for humanitarian corridors). Conflating the two risks the appearance that low feasibility extinguishes the obligation — a conclusion that is normatively unacceptable under IHL. The Mariupol case demonstrates the operational stakes: throughout much of the siege, the danger of remaining was extreme while the feasibility of evacuation was severely constrained by security conditions, contested corridors, and disputed ceasefire guarantees. A composite index that aggregated these into a single severity score would mask the fact that the obligation to protect civilians persisted regardless of operational constraint, and would obscure precisely the type of high-danger / low-feasibility condition that requires escalated political engagement rather than algorithmic resolution.
CERAI is a decision support instrument. Its justification rests not on predictive accuracy but on three properties that are foundational to humanitarian decision-support tools: transparency (every score is decomposable to its constituent inputs and weights), structure (the three-dimension architecture forces decision-makers to consider both whether to evacuate and whether evacuation is possible), and reproducibility (the same inputs yield the same outputs, and the methodology is open to peer scrutiny and replication). The model formalises case-based reasoning — the institutional memory that experienced humanitarian coordinators apply when drawing on past operations to inform judgment about novel situations. When CERAI identifies that a current scenario most resembles Mariupol 2022 or Mosul 2016-17, it is surfacing the type of precedent-based reasoning that ICRC delegates and UN coordinators already apply in the field, making that reasoning explicit and contestable rather than implicit and individual. CERAI is not a successor to expert judgment but a structured aid that preserves what works in expert practice — pattern recognition across precedents — while making the patterns visible and auditable.
CERAI v3 uses the geometric mean rather than the arithmetic mean to aggregate sub-scores. The geometric mean penalises extreme lows more severely: a single sub-score near zero produces a near-zero aggregate score regardless of other sub-scores. This better reflects operational reality: a fully closed evacuation corridor (zone exit = 5%) makes the overall feasibility critically low regardless of how good weather and destination conditions are. The arithmetic mean would mask this single blocking constraint.
The multiplicative operation of the vulnerability multiplier on both endangerment and feasibility is not double-counting. Vulnerability acts on the same construct from two causally distinct directions. For endangerment, vulnerability amplifies the functional harm produced by a given objective hazard — a Sen-style entitlement effect, where the same artillery strike produces greater individual harm to populations with worse pre-existing health, lower mobility, and weaker social safety nets. For feasibility, vulnerability reduces the navigability of a given objective evacuation route — children, elderly, and chronically ill cannot traverse the same distances in the same times under the same conditions as able-bodied adults. These are two distinct causal pathways producing a coincident directional effect on the same population, not two measurements of the same underlying variable. The multiplicative structure correctly represents this: vulnerability is a stratifier of objective conditions, not a substitute for them.
The 0.7-1.3 bound on the Dimension 3 multiplier rests on three principles. First, empirical anchoring: observed mortality and morbidity differentials between vulnerable and non-vulnerable populations in conflict-affected settings typically run 1.5x to 2x at the individual level, but at the population level this effect is diluted by the share of the population that is vulnerable; a multiplier capped at 1.3 reflects this dilution while preserving directional sensitivity. Second, bounded leverage: the vulnerability multiplier is intended to refine, not dominate, the objective endangerment and feasibility assessments; a 30% cap ensures that no plausible vulnerability profile can shift a low-endangerment assessment into the high-endangerment band purely on demographic grounds. Third, symmetric treatment: the symmetric 0.7-1.3 range (rather than asymmetric, e.g. 0.8-1.5) reflects the modelling assumption that vulnerability operates as a stratification of risk, not as a substitution for the objective assessment, and avoids implicit weighting toward upward adjustments.
Trajectory is calculated from a rolling array of the last five endangerment scores recorded during the session. The mean change per step determines the trajectory state, which is reported in five categories: Rapidly Deteriorating (mean change > +4%), Deteriorating (+1% to +4%), Stable (−1% to +1%), Improving (−4% to −1%), and Rapidly Improving (< −4%). The "Current Trajectory" label and directional arrow on the Endangerment tab reflect this state. A minimum of two stored scores is required; fewer than two yields "Insufficient History."
When the ACLED API is connected and returns event data, trajectory shifts to a data-driven calculation based on the percentage change in conflict event frequency over the prior 7-day rolling window, replacing the session-input-based calculation. The trajectory explanation beneath the label indicates which method is active. Time-to-threshold projects the number of days to reach the IHL Obligation Threshold (75%) assuming linear continuation of the current trajectory, using the formula: days = (75 − currentScore) / (meanChangePerCalc × 10), where 10 represents a reference assumption of 10 user updates per day. This projection is a planning tool only; actual trajectories are non-linear and should be read alongside ACLED trend data where available.
Each manually entered field carries a source credibility selector with the following multipliers applied to that field's weight: Unverified=0.7, Media Report=0.8, NGO Report=0.9 (default), UN/ICRC Verified=1.0, Government Official=0.85. Government data receives a slight discount (0.85) to reflect potential conflict-of-interest in government reporting on civilian casualties. UN/ICRC verified data receives full weight as the gold standard for humanitarian monitoring.
For annual-tier data sources (GCRI, static population figures), a staleness penalty is applied: weight is reduced by 10% per 3-month period since last update, with a minimum multiplier of 0.5 (50%). This reflects the reduced reliability of structural data as conflict situations evolve. Real-time API data (Open-Meteo weather) carries no staleness penalty for the first 30 minutes, then degrades to amber status.
The What-If tool identifies the five most impactful variables by testing each variable at ±25% of its range and measuring the resulting total score delta. The variable with the highest delta is listed first. For each selected variable, users can adjust the value via a slider and observe live changes in both endangerment and feasibility. This supports scenario planning and sensitivity analysis without modifying the main assessment.
Current weights were derived by the researcher from IHL doctrine review and face validity assessment against the 21 face-validity anchor cases. A documented next step is Analytic Hierarchy Process (AHP) validation with a panel of IHL practitioners and humanitarian operations experts. AHP would produce a quantified expert consensus weight matrix to replace the current researcher-assigned weights, strengthening the model's academic validity and supporting peer review.
GCRI (EU JRC Global Conflict Risk Index) and FEWS NET IPC phase are currently manual entry fields. The planned automation pathway is: GCRI via EU JRC API (annual cadence, amber freshness tier); FEWS NET via IPC API (monthly cadence, amber freshness tier). In the current prototype these serve as supplemental context inputs that inform the Decision Support Briefing text but are not weighted into the Dimension 1 endangerment score to avoid double-counting with direct conflict intensity inputs.
The 21 face-validity anchor cases used in CERAI's initial parameterisation function as a face-validity assessment, not as statistical calibration. Calibration in the statistical sense requires independent ground truth; for civilian evacuation decisions during armed conflict, no such ground truth exists. What the historical cases provide is a check on construct alignment: do the model's outputs for documented past cases align with the consensus retrospective expert judgment of how dangerous those situations were and how feasible evacuation was at the time? Where alignment is strong, the framework is at least not patently mis-specified. Where alignment is weak, either the framework requires revision or the case in question is one where the documented record itself is contested. This is the same evidentiary standard that operational humanitarian indices apply during development; INFORM Severity's initial weights similarly rest on expert judgment pending statistical refinement.
The CERAI case database currently contains 47 documented evacuation cases since 2000, spanning conflict-affected operations across three continents. Of these, four cases — Op Raahat Yemen 2015, East Aleppo 2016, Kabul 2021, and Mariupol 2022 — serve as initialisation anchors, representing the moderate through extreme risk bands used to set the scoring curve. All 21 of the original historical cases retain face-validity anchor status and appear in the Historical Case Dataset table. The remaining 26 cases serve the nearest-neighbour precedent comparator only and were added to increase coverage of under-represented conflict typologies, including protracted urban sieges, state-led displacement operations, and non-permissive NEO environments. Cases added beyond the original 21 carry researcher-review flags on their vector encodings pending confirmation of source values.
In the absence of live operational data, the scoring curve was initialised using four documented historical cases as reference anchors — Op Raahat Yemen 2015, East Aleppo 2016, Kabul 2021, and Mariupol 2022 — representing moderate through extreme risk bands. In a fully operational deployment, this initialisation would be replaced by live data inputs from sources including ACLED, INFORM Severity, and ECMWF, which would continuously update the model's inputs, improving assessment accuracy as live data replaces precautionary defaults.
The 47 cases in the comparator database are retained as contextual reference points rather than ground truth training data. The historical precedent comparator formalises this process: when the model identifies that a current scenario most resembles Mariupol 2022 or Mosul 2016-17, it is surfacing the kind of precedent-based reasoning that ICRC delegates and UN coordinators apply in the field.
The term 'ground truth risk' as used in this tool refers to scores assigned to historical cases through a composite rubric based on documented outcome severity, recorded casualty levels, and primary threat classifications. These scores are educated approximations, not statistically validated ground truth. Rigorous ground truth derivation would require expert panel validation using the Analytic Hierarchy Process (AHP) — documented as the next methodological step in this research.
The comparator encodes each historical case as a 5-dimensional vector [threat, population, ops, political, destination] derived from archival case analysis. Euclidean distance in normalised vector space is used to identify the three closest analogues. Euclidean distance was preferred over cosine similarity as it penalises magnitude differences (overall severity level) rather than only direction (relative profile shape).
Each case in the comparator database is encoded as a 5-dimensional vector on a 0–100 scale. The five dimensions correspond to: Threat profile — composite hostile-event intensity and targeting risk; Population characteristics — size, density, and vulnerability composition of the affected population; Operational conditions — corridor availability, transport sufficiency, and logistics capacity at the time of assessment; Political consent — degree of armed actor consent to evacuation, government authorisation, and international mediation status; Destination capacity — receiving area shelter, medical, and protection capacity.
Vector values were assigned by the researcher through analysis of primary source documentation including OHCHR situation reports, ICRC operational memoranda, OCHA humanitarian access monitoring, and peer-reviewed post-conflict studies. Where primary documentation was insufficient, values carry a researcher-review flag indicating that the vector requires confirmation before operational use. The five-variable encoding predates the three-dimension CERAI v3 restructure; recoding to the D1/D2/D3 framework is a documented next step. The five-variable encoding remains a valid approximation for precedent-surfacing purposes and does not affect the scoring calculation.
The current query vector is constructed from the user's live inputs as follows: threat = current Threat Environment score (D1); population = a composite of total population and vulnerability percentage inputs; ops = the mean of Zone Exit and Route Conditions sub-scores (D2); political = a composite of ceasefire, government consent, and armed group consent inputs; destination = current Protection at Destination sub-score (D2). Cases added to the database beyond the original 21 face-validity anchors carry researcher-review flags on individual vector dimensions where source documentation required estimation. The comparator's query vector uses the pre-multiplier Dimension 1 raw score as the threat component, rather than the final multiplied endangerment score. This is intentional: the comparator matches on objective threat conditions, not on the population-adjusted output, ensuring that cases with similar threat profiles are surfaced regardless of differences in population vulnerability characteristics.
The Assessment Summary panel, displayed persistently between the output tab section and the Assessment Inputs section, provides a consolidated three-value summary of the current assessment state. The three values are: Threat Environment Score — the final Dimension 1 endangerment score after the Dimension 3 vulnerability multiplier has been applied, identical to the value displayed on the Endangerment gauge; Feasibility Score — the composite Dimension 2 feasibility score from Zone Exit, Route Conditions, and Protection at Destination sub-scores, identical to the value displayed on the Feasibility gauge; Vulnerability Multiplier — the Dimension 3 multiplier value with its percentage adjustment from neutral (1.00×), reflecting the population vulnerability profile. The panel is visible regardless of which output tab is active, allowing the user to reference all three key values simultaneously without tab switching.
The weights below were derived by the researcher through IHL doctrine review — each variable is assigned a weight reflecting the primacy of its corresponding IHL obligation. Specific numerical weights were assigned through Claude-assisted analysis of the relative severity and frequency of each obligation in documented evacuation case literature. This is the same methodology INFORM Severity used for its initial weight set, which ACAPS describes as best estimates pending expert refinement. CERAI's weights are similarly pending Analytic Hierarchy Process (AHP) validation with a practitioner panel of IHL lawyers, ICRC delegates, and OCHA humanitarian coordinators — documented as the next methodological step.
| Dimension | Component | Weight | IHL Basis |
|---|---|---|---|
| Dimension 1 | Hostility Intensity | 20% | AP I Art. 51 — civilian protection from attacks |
| Dimension 1 | Armed Group Proximity | 14% | AP I Art. 58 — passive precautions, distance from military objectives |
| Dimension 1 | Imminent Attack | 10% | AP I Art. 57(2) — feasibility of warning |
| Dimension 1 | Retaliation Risk | 8% | AP I Art. 51(2) — prohibition on acts of violence against civilians |
| Dimension 1 | Chem/Bio Threat | 7% | CWC absolute prohibition; mass casualty multiplier |
| Dimension 1 | Escalation Risk | 12% | Trajectory component of precautionary planning |
| Dimension 1 | Non-Conflict Hazards | 8% | Guiding Principles on IDP Principle 5 — all hazards |
| Dimension 1 | Food/Water (IPC) | 10% | GCIV Art. 54 — starvation prohibition; AP I Art. 70 |
| Dimension 1 | Energy Availability | 6% | Medical dependency; GCIV Art. 18 hospital protection |
| Dimension 1 | Vulnerable Population % | 5% | GCIV Art. 16; AP I Arts. 77-78 |
| Dimension 2 (Zone) | Corridor Status | 30% | OCHA AMRF; corridor denial = key encirclement indicator |
| Dimension 2 (Zone) | AG Consent Zone | 28% | GCIV Art. 49; ICRC access negotiations |
| Dimension 2 (Zone) | Mine Load Zone | 24% | Ottawa Treaty; UNMAS clearance obligations |
| Dimension 2 (Zone) | Imminent Attack (Zone) | 18% | AP I Art. 57(2)(c) — effective warning |
| Dimension 2 (Route) | Weather | 26% | Operational modifier; environmental hazard |
| Dimension 2 (Route) | Daylight | 20% | Operational modifier; safety and targeting risk |
| Dimension 2 (Route) | Transport Sufficiency | 21% | IOM logistics cluster; AP I Art. 57(2)(c) |
| Dimension 2 (Route) | Mine Load Route | 19% | Ottawa Treaty; UNMAS clearance obligations |
| Dimension 2 (Route) | Distance to Safety | 14% | Exposure time and logistics constraint |
| Dimension 2 (Dest) | Shelter Availability | 26% | UNHCR shelter cluster; Guiding Principles P.18 |
| Dimension 2 (Dest) | Medical Capacity | 24% | WHO health cluster; GCIV Art. 18 |
| Dimension 2 (Dest) | Destination Security | 22% | Non-refoulement; GCIV Art. 147 |
| Dimension 2 (Dest) | Willingness to Receive | 16% | UNHCR durable solutions framework; non-refoulement principle |
| Dimension 2 (Dest) | AG Consent (Destination) | 12% | GCIV Art. 49; ICRC access negotiations doctrine |
| Dimension 3 | Children Under 12 (%) | amplifier | AP I Arts. 77-78; GCIV Art. 24 — special protection for children |
| Dimension 3 | Elderly Over 65 (%) | amplifier | GCIV Art. 16 — special protection for elderly |
| Dimension 3 | GBV Risk | amplifier | UN Women/UNFPA protection monitoring; SGBV as IHL violation |
| Dimension 3 | Prior Evacuation Experience | amplifier (inverse) | Operational capacity factor |
| Dimension 3 | Financial Resources / Mobility | amplifier (inverse) | Sen capability approach; entitlement framework |
Sub-weights within each Dimension 2 sub-score (Zone, Route, Destination) sum to 100% independently. The three sub-scores are then combined using a geometric mean to produce the composite Feasibility score.
Dimension 3 variables operate as a multiplier (0.7×–1.3×) applied to both Dimension 1 and Dimension 2 scores rather than as additive weighted inputs.
CERAI operates at the same evidentiary standard as the leading operational humanitarian severity instrument. INFORM Severity Index is a composite indicator built on 31 core indicators across three weighted dimensions, with weights described by ACAPS as best estimates pending refinement through expert analysis and statistical methods, and is used by UNOCHA and IFRC for operational funding and resource allocation. CERAI's researcher-assigned weights pending Analytic Hierarchy Process validation operate at the same evidentiary level as INFORM's expert-assigned weights pending statistical refinement. The difference is not in the rigour of weight derivation but in the construct measured: INFORM Severity measures the intensity of humanitarian need across a crisis; CERAI measures the conditions relevant to an evacuation decision under IHL. CERAI's geometric mean aggregation throughout reflects a stronger non-compensability assumption appropriate to the evacuation context — a closed corridor cannot be averaged out by good weather — whereas INFORM Severity uses a hybrid arithmetic-geometric scheme appropriate to its broader severity construct.
The same endangerment score represents different levels of individual risk depending on population characteristics. Dimension 3 surfaces this by producing a risk range rather than a single score. A population with high child and elderly proportions, limited resources, and no prior evacuation experience may face conditions significantly worse than the headline score suggests at the individual household level. This approach aligns with INFORM Severity's principle that not all people affected by a crisis are equally affected. The Dimension 3 multiplier makes this distribution explicit rather than assuming a homogeneous affected population.
The 75% marker on the Endangerment gauge and spectrum bar represents the threshold at which conditions are considered to demand evacuation under GC IV Article 49, which requires evacuation "where the security of the population or imperative military reasons so demand." This threshold was derived through face validity assessment against the historical case dataset — cases above 75% in the dataset consistently involved IHL evacuation obligations being triggered or having been triggered in retrospective legal analysis. Cases below 75% generally involved precautionary planning obligations rather than immediate evacuation demands. The threshold is indicative only and human legal judgment is required in all cases. Above this threshold, the model reflects conditions that historically correlate with IHL obligations; it does not itself create legal obligations.