Data enrichment is a crucial component of data preparation, aimed at improving the quality of input datasets by integrating external information from third-party sources. One approach to support this process is reconciliation-based data enrichment. In this approach, data providers organize their information around entities and expose reconciliation services for data consumers. Data consumers use these services to match tabular values in their data with entities in the source they want to access, enabling them to fetch additional information by querying the source with canonical identifiers. This approach reduces information asymmetry and facilitates access to external data, making enrichment more effective and user-friendly. The chapter introduces reconciliation-based enrichment, reviews current techniques for entity reconciliation in tabular data, and highlights the emerging role of large language models in supporting such tasks. Finally, it discusses the advantages of reconciliation-based enrichment with respect to sustainability principles, including reducing computational effort, limiting human intervention, and promoting resource efficiency in data preparation pipelines.

De Paoli, F., Palmonari, M., Belotti, F. (2027). Reconciliation-Based Data Enrichment. In B. Pernici, T. Catarci, M. Palmonari, G. Simonini (a cura di), Discount Quality for Responsible Data Science Human-in-the-Loop for Quality Data (pp. 39-59). Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-29481-4_4].

Reconciliation-Based Data Enrichment

De Paoli F.;Palmonari M.;
2027

Abstract

Data enrichment is a crucial component of data preparation, aimed at improving the quality of input datasets by integrating external information from third-party sources. One approach to support this process is reconciliation-based data enrichment. In this approach, data providers organize their information around entities and expose reconciliation services for data consumers. Data consumers use these services to match tabular values in their data with entities in the source they want to access, enabling them to fetch additional information by querying the source with canonical identifiers. This approach reduces information asymmetry and facilitates access to external data, making enrichment more effective and user-friendly. The chapter introduces reconciliation-based enrichment, reviews current techniques for entity reconciliation in tabular data, and highlights the emerging role of large language models in supporting such tasks. Finally, it discusses the advantages of reconciliation-based enrichment with respect to sustainability principles, including reducing computational effort, limiting human intervention, and promoting resource efficiency in data preparation pipelines.
Capitolo o saggio
data enrichment, data engineering, semantic web
English
Discount Quality for Responsible Data Science Human-in-the-Loop for Quality Data
Pernici, B; Catarci, T; Palmonari, M; Simonini, G
26-set-2026
2027
9783032294807
2658
Springer Science and Business Media Deutschland GmbH
39
59
De Paoli, F., Palmonari, M., Belotti, F. (2027). Reconciliation-Based Data Enrichment. In B. Pernici, T. Catarci, M. Palmonari, G. Simonini (a cura di), Discount Quality for Responsible Data Science Human-in-the-Loop for Quality Data (pp. 39-59). Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-29481-4_4].
open
File in questo prodotto:
File Dimensione Formato  
978-3-032-29481-4_4.pdf

accesso aperto

Tipologia di allegato: Publisher’s Version (Version of Record, VoR)
Licenza: Creative Commons
Dimensione 817.02 kB
Formato Adobe PDF
817.02 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10281/627621
Citazioni
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
Social impact