This document builds on UNHCR’s ongoing transformation to cloud computing, which has delivered significant operational benefits while also introducing structural governance risks inherent to reliance on hyperscaler services — most notably exposure to extraterritorial US jurisdiction affecting data protection, and concentration risk affecting business continuity. Against this background, it proposes a practical strategy and action plan to mitigate these risks, structured for straightforward implementation within UNHCR’s operational and institutional context. It addresses both risks generic to cloud adoption and those specific to humanitarian operations, and translates the response into clear objectives, workstreams, and implementation steps.
Overview of the current situation
UNHCR manages data that informs every protection decision, assistance distribution, and resettlement referral. Its data environment spans two categories: highly sensitive beneficiary data (registration, biometrics, RSD case files, protection monitoring, cash assistance records) and operational and administrative data (finance, supply chain, HR, programme management).
Over the past eight years, both categories have migrated from on-premises legacy systems to commercial cloud platforms — principally US-headquartered hyperscalers (Oracle, Microsoft, Google) — delivering real-time visibility, elastic capacity, and security at scale. However, this dependency places the organisation’s most sensitive data within the jurisdictional reach of US law and the operational footprint of providers entangled in military and political dynamics. The current environment is therefore operationally modernised but structurally exposed — the circumstance this strategy is designed to address.
Rationale for the focus of this document
UNHCR’s Digital Transformation Strategy 2022–2026 commits the organisation to data-driven decision-making, while its Data Protection Policy and Policy on the Protection of Personal Data of Persons of Concern establish confidentiality, purpose limitation, and data subject rights as binding obligations. The current hyperscaler-dependent posture creates a measurable gap between these commitments and operational reality: UNHCR cannot guarantee that protection-critical data remains beyond the reach of foreign jurisdiction or politically motivated service disruption.
This document focuses on closing that gap. It proposes a tiered data-classification approach that uses UNHCR’s own data — sensitivity profiles, user access patterns, operational criticality, and incident history — to drive evidence-based hosting decisions aligned with existing policy. By treating the problem as data governance rather than vendor selection, the strategy operationalises commitments UNHCR has already made, rather than introducing new obligations.
Objectives & timescale
- Governance (2026) Establish a Cloud Data Classification and Hosting Standard, owned by the Data Governance Committee, mandating sensitivity tiers and permissible hosting per tier — operationalised through a data catalogue (Unity Catalog) that tags assets by sensitivity tier and maintains a live inventory, with data lineage (OpenLineage) tracing Tier-1 data flows to verify hosting compliance and flag non-compliant hosting.
- Interim mitigation (2026–2027) Migrate the top three protection-critical data assets to air-gapped or sovereign-hosted environments, modelled on NATO’s Google Distributed Cloud deployment.
- External advocacy (2027) Secure UNHCR endorsement of, and table at the Digital and Technology Network of the High-Level Committee on Management (HLCM), a formal proposal for UN sovereign cloud infrastructure.
- Long-term transformation (by 2030) Eliminate hyperscaler hosting of Tier-1 protection-critical data through open-source platforms, sovereign cloud, or UN-jurisdiction on-premises hosting.
| Phase | Period | Key milestone |
|---|---|---|
| Foundation | 2026 | Classification standard endorsed; Tier-1 inventory complete |
| Interim mitigation | 2026–2027 | Top three Tier-1 assets migrated to sovereign hosting |
| Advocacy | 2027 | UN sovereign cloud proposal tabled at HLCM |
| Long-term transformation | 2028–2030 | Tier-1 hyperscaler dependency eliminated |
Organisational roles involved
Implementation responsibility is distributed across four organisational bodies, aligned with the phases of the strategy.
- The High Commissioner Provides political endorsement, tables the UN sovereign cloud proposal at the Digital and Technology Network of the HLCM, and engages member state counterparts where required.
- The Senior Management Committee Sponsors the strategy at executive level, integrates it with the Digital Transformation Strategy, and secures the resource envelope required for implementation.
- The Data Governance Committee Owns the Cloud Data Classification and Hosting Standard, approves tier assignments, monitors compliance across operations, and reports progress to the Senior Management Committee.
- DIMA (Division of Information Management and Analytics) At headquarters and bureau level, leads technical implementation: the Tier-1 data inventory, migration to sovereign hosting, and architectural choices for long-term transformation.
| Action | High Commissioner | Senior Management Committee | Data Governance Committee | DIMA |
|---|---|---|---|---|
| Endorse strategy | C | A | C | R |
| Tier-1 inventory & migration | I | I | C | R / A |
| Classification standard | I | I | R / A | C |
| UN sovereign cloud advocacy | R / A | C | I | I |
R = Responsible · A = Accountable · C = Consulted · I = Informed
Key technologies & methodologies
This action plan applies tools and methodologies examined in the Modern Data Architecture and Data Governance sessions.
- Data governance, cataloguing and lineage (Unity Catalog, OpenLineage) The mechanism that makes the classification standard enforceable. Unity Catalog’s federation tags assets by sensitivity tier and unifies UNHCR’s dispersed sources into a single governed catalogue; OpenLineage traces Tier-1 data flows to verify hosting compliance end-to-end.
- Medallion architecture (Bronze / Silver / Gold) Applied as a complementary quality-classification layer — sensitivity tiering governs where data is hosted, while Medallion governs how data is refined, supporting transparency and reprocessing without re-ingesting raw data.
- Open-source, portable data stack (Trino, dbt, Spark, Delta Lake, Airflow) A vendor-neutral lakehouse stack adopted for Tier-1 workloads. By decoupling storage from compute, it eliminates proprietary hyperscaler lock-in and guarantees portability to sovereign infrastructure — the technical foundation of the long-term transformation objective.
How the course content shaped this document
This document draws directly on all eight modules of the programme.
Module 1 — Amazing AI. Data Pipeline Automation informed the choice of pipeline-orchestration tools (Airflow, dbt) in the open-source stack proposed for Tier-1 workloads.
Module 2 — Frameworks for Continuous Data Innovation. Shaped the document’s structure: Decision-Making Frameworks underpin the rule for sorting data — the repeatable criteria (sensitivity, criticality, jurisdiction exposure) that determine which tier each asset belongs to — while Design of Organizations informed the roles-and-responsibilities mapping across the High Commissioner, Senior Management Committee, Data Governance Committee, and DIMA.
Module 3 — Architecture and Querying Data. Particularly Architect Simplicity — shaped the strategy’s core design move: following Roger Sessions’ principle that complex architectures become manageable when elements are grouped into layers by scale, the strategy partitions UNHCR’s data into sensitivity tiers and applies proportionate hosting and governance at each tier, rather than imposing a single, uniform solution across all data.
Module 4 — The Importance of Data. Supplied the rationale’s central argument through the Artisan vs. Factory concept: that UNHCR must treat data governance as a systematic, repeatable discipline (“factory”) rather than an ad-hoc, operation-by-operation practice (“artisan”).
Module 5 — Data Platforms and Database Design. Informed the distinction between data categories (beneficiary vs. operational) and the sensitivity-tiering logic at the heart of the classification standard.
Module 6 — Data Science Acceleration and the Modern Data Platform. The Modern Data Stack and Modern Data Stack Patterns directly supplied the open-source, portable lakehouse stack (Trino, dbt, Spark, Delta Lake, Airflow) proposed for Tier-1 workloads.
Module 7 — The Cloud. This module is the document’s foundation: The Cloud and Data Leadership framed the central tension between the operational benefits of cloud adoption and the loss of control it entails — the precise problem this strategy addresses.
Module 8 — Ethics, Information Governance, and the Modern Data Organisation. Underpins the entire argument: Data Governance and Compliance, and the AI ethics material, connect directly to UNHCR’s obligation to protect persons of concern, framing data sovereignty as an ethical as well as a technical imperative.
Reference is made to all eight modules, though their influence varies by design: Modules 6, 7 and 8 are load-bearing — supplying the technology stack, the central cloud tension, and the governance-ethics foundation respectively — while others provide conceptual grounding. A focused strategy draws unevenly but deliberately on a broad curriculum.
Other research. The strategy also draws on UNHCR’s Digital Transformation Strategy 2022–2026 and Data Protection Policy, the US CLOUD Act, NATO’s Google Distributed Cloud deployment, and reporting on hyperscaler service suspensions affecting international institutions.
Note on using AI for this work
I utilised an LLM model (Claude, by Anthropic) to assist in developing this document. I began by setting out the problem — the governance and sovereignty risks arising from UNHCR’s dependence on commercial hyperscaler services — and iteratively developed the strategy and action plan through a guided conversation, supplying my own source materials at each stage.
The AI assisted in reviewing and refining my drafts of each section (purpose, current-situation overview, rationale, objectives, timescale, organisational roles, and key technologies); summarising and synthesising source documents I provided, including the Modern Data Architecture session materials and my earlier analyses of UNHCR’s cloud and enterprise-systems context; mapping the document’s content to the relevant course modules; and checking conceptual consistency — for example, distinguishing maturity-based (Medallion) from sensitivity-based (Tier-1 / Tier-2) classification, and identifying where a proposed tool would have contradicted the document’s sovereignty argument.
I directed the scope and structure of the document throughout, made all editorial decisions independently, and brought my own knowledge of UNHCR operations and the source materials to each iteration. I challenged outputs that were inaccurate, over-stated, or inconsistent with the humanitarian context — including correcting module misattributions and rejecting a proposed analytical platform that conflicted with the document’s open-source and data-sovereignty commitments. I acknowledge that AI language models can produce inaccurate, incomplete, or unsourced content, and I remained vigilant regarding these limitations throughout. I take full responsibility for the accuracy and integrity of this document.
Citation. Claude AI (Opus 4.6, 4.7 and 4.8, Anthropic). Used for source summarisation, iterative drafting and refinement, course-content mapping, and consistency checking. Example prompts used include: “Review the below text on the scope of the document”, and “How can I link and use the contents of the attached document in supporting my objectives?”