Data enrichment is a crucial component of data preparation, aimed at improving the quality of input datasets by integrating external information from third-party sources. One approach to support this process is reconciliation-based data enrichment. In this approach, data providers organize their information around entities and expose reconciliation services for data consumers. Data consumers use these services to match tabular values in their data with entities in the source they want to access, enabling them to fetch additional information by querying the source with canonical identifiers. This approach reduces information asymmetry and facilitates access to external data, making enrichment more effective and user-friendly. The chapter introduces reconciliation-based enrichment, reviews current techniques for entity reconciliation in tabular data, and highlights the emerging role of large language models in supporting such tasks. Finally, it discusses the advantages of reconciliation-based enrichment with respect to sustainability principles, including reducing computational effort, limiting human intervention, and promoting resource efficiency in data preparation pipelines.
De Paoli, F., Palmonari, M., Belotti, F. (2027). Reconciliation-Based Data Enrichment. In B. Pernici, T. Catarci, M. Palmonari, G. Simonini (a cura di), Discount Quality for Responsible Data Science Human-in-the-Loop for Quality Data (pp. 39-59). Springer Science and Business Media Deutschland GmbH [10.1007/978-3-032-29481-4_4].
Reconciliation-Based Data Enrichment
De Paoli F.;Palmonari M.;
2027
Abstract
Data enrichment is a crucial component of data preparation, aimed at improving the quality of input datasets by integrating external information from third-party sources. One approach to support this process is reconciliation-based data enrichment. In this approach, data providers organize their information around entities and expose reconciliation services for data consumers. Data consumers use these services to match tabular values in their data with entities in the source they want to access, enabling them to fetch additional information by querying the source with canonical identifiers. This approach reduces information asymmetry and facilitates access to external data, making enrichment more effective and user-friendly. The chapter introduces reconciliation-based enrichment, reviews current techniques for entity reconciliation in tabular data, and highlights the emerging role of large language models in supporting such tasks. Finally, it discusses the advantages of reconciliation-based enrichment with respect to sustainability principles, including reducing computational effort, limiting human intervention, and promoting resource efficiency in data preparation pipelines.| File | Dimensione | Formato | |
|---|---|---|---|
|
978-3-032-29481-4_4.pdf
accesso aperto
Tipologia di allegato:
Publisher’s Version (Version of Record, VoR)
Licenza:
Creative Commons
Dimensione
817.02 kB
Formato
Adobe PDF
|
817.02 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


