Context
You are part of a consortium that operates a multidimensional data warehouse on top of a lakehouse to analyse tourism, employment, and social vulnerability indicators.
Data comes from:
- Edge: mobile apps used by tourists, local sensors (mobility, energy), tablets used by community organisations.
- Fog: regional servers operated by local tourism offices / municipalities that pre-aggregate and anonymise data.
- Cloud: a lakehouse platform (Bronze / Silver / Gold) plus multidimensional marts (fact tables such as F_TOURISM_OBSERVATION and dimensions like Time, GeoArea, Gender, AgeGroup, Indicator, etc.).
Several countries are involved (EU and non-EU). Communities have raised concerns about data sovereignty, environmental footprint, and the social impact of concentrating data on big cloud providers.
Learning goals
By the end of this exercise, you should be able to:
- Decide where to run which part of a lakehouse + data warehouse (edge, fog, cloud).
- Reason about compute and storage resources, economic cost, and sustainability.
- Discuss responsibility, data sovereignty, and the social impact of infrastructure choices.
Material
- Lakehouse + Tourism Star Schema + Cloud–Fog–Edge Scenarios: https://gist.github.com/gevargas/d0b20cc26a5d66ba961c718b2f1f3080
Part 1 – Workload characterisation
1. Briefly describe the main components of the system (½–1 page):
- Lakehouse layers (Bronze, Silver, Gold).
- Data warehouse star schemas and fact/dimension tables.
- Types of workloads: ingestion / transformations, OLAP dashboards, ad-hoc SQL, data science notebooks.
2. For each workload type, classify:
- Latency sensitivity: low / medium / high.
- Latency sensitivity = how strongly a workload depends on quick response time.
- If a workload is high latency-sensitive, users or systems need answers in (near) real time; even small delays are problematic (e.g. interactive dashboards, transaction authorizations).
- If it is medium, a delay of seconds or minutes is acceptable (e.g. most ad-hoc analytics).
- If it is low, long delays are fine (e.g. nightly batch ETL, offline reports).
- Latency sensitivity = how strongly a workload depends on quick response time.
- Data volume: MB / GB / TB per day (order of magnitude).
- Privacy sensitivity: low / medium / high.
Part 2 – Two deployment scenarios
Design two possible deployments:
- Scenario A – Cloud-centric: most compute and storage in one major cloud region.
- Scenario B – Hybrid cloud–fog–edge: more pre-processing and partial storage in fog/edge, slimmer core cloud.
For each scenario:
1. Draw a deployment diagram showing:
- Edge devices.
- Fog nodes (regional servers / sovereign cloud regions).
- Cloud lakehouse (object storage + query engine + data warehouse marts).
2. For each lakehouse layer (Bronze, Silver, Gold), explain:
- Where data is stored and processed (edge/fog/cloud).
- Why did you choose this location (performance, privacy, operational reasons)?
3. Justify the deployment in terms of:
- Performance and latency.
- Operational complexity (few big nodes vs many smaller nodes).
- Ability to enforce local data sovereignty (e.g. data that cannot leave the country).
Part 3 – Resources and economic cost
For each scenario:
1. Propose approximate resource profiles:
- ETL/ELT jobs (vCPUs, hours per day). A vCPU (virtual CPU) is a virtual processing core that a cloud provider assigns to your virtual machine or container.
- Roughly: 1 vCPU ≈ 1 hardware thread (often one hyper-thread on a physical CPU core).
- Your VM’s “number of vCPUs” tells you how many tasks it can process in parallel and is the unit you pay for in most cloud pricing (e.g. “4 vCPUs, 16 GB RAM”).
- OLAP usage (concurrent users, hours per day).
- Storage volumes for Bronze / Silver / Gold (relative magnitude).
- Network egress (GB per month).
2. Based on public cloud price calculators (approximate values), estimate:
- Monthly compute cost.
- Monthly storage cost.
- Monthly network egress cost.
- Total monthly cost for each scenario.
3. Identify which components dominate the cost and explain why.
Part 4 – Sustainability and environmental footprint
For each scenario:
1. Identify the most energy-intensive parts of the system (frequent ETL, large clusters, heavy egress, etc.).
2. Discuss the environmental impact:
- Energy use and CO₂ emissions (centralised data centre vs many fog nodes).
- Hardware lifecycle and potential e-waste at edge/fog.
- Choice of regions/providers with greener energy.
3. Propose two concrete design decisions that could reduce the environmental footprint.
Part 5 – Responsibility, data sovereignty, social impact
For each scenario:
1. Analyse data sovereignty:
- Where is personal/community-sensitive data physically stored?
- Which jurisdictions apply?
- How can communities control access to their data?
2. Identify risks of digital extractivism:
- Who controls infrastructure and derived value?
- Are local actors able to run their own analyses?
3. Propose at least one governance and one technical mechanism to support:
- Local control over permissions.
- Auditable access logs.
- Community opt-out / deletion / aggregation of data.
4. Conclude with a ½–1 page reflection: Which scenario (or hybrid of both) better supports responsible, sovereignty-aware analytics? Why?
Deliverables
- A short report (4–6 pages) covering Parts 1–5.
- 5-minute in-class presentation defending your design decisions.