| Producing public health reports at the district level from urban data requires combining population cohort records held by municipalities with environmental measurements maintained by separate organisations. Raw administrative records cannot be shared across institutional boundaries without pseudonymisation, aggregation, or other privacy-protective transformations required under the General Data Protection Regulation (GDPR). Large language models offer a practical route to accessible, natural-language reporting for non-technical public health officials, but only if the underlying data is both privacy-safe and semantically coherent across sources. We propose a three-pipeline architecture in which provider-side pipelines, guided by the SemT-X framework, apply pseudonymisation, field suppression, and controlled generalisation before aggregating records into aggregate-only district-month products. A third consumer pipeline uses an LLM with SQL tool-use to query the aggregated datasets and generate structured reports, with all interactions logged for auditability. End-to-end feasibility is demonstrated through a case study combining municipal kindergarten enrolment data with multi-source air quality measurements, producing an illustrative report via a reproducible DuckDB shared data plane. |
*** Title, author list and abstract as submitted during Camera-Ready version delivery. Small changes that may have occurred during processing by Springer may not appear in this window.