EconPapers    
Economics at your fingertips  
 

Data Lake Governance: Towards a Systemic and Natural Ecosystem Analogy

Marzieh Derakhshannia, Carmen Gervet, Hicham Hajj-Hassan, Anne Laurent and Arnaud Martin
Additional contact information
Marzieh Derakhshannia: LIRMM, Univ. Montpellier, CNRS, 34090 Montpellier, France
Carmen Gervet: Espace Dev, Univ. Montpellier, IRD, Univ Guyane, Univ. Réunion, 34293 Montpellier, France
Hicham Hajj-Hassan: CNRS-L, Beirut P.O. Box 11-8281, Lebanon
Anne Laurent: LIRMM, Univ. Montpellier, CNRS, 34090 Montpellier, France
Arnaud Martin: CEFE, Univ. Montpellier, CNRS, 34293 Montpellier, France

Future Internet, 2020, vol. 12, issue 8, 1-16

Abstract: The realm of big data has brought new venues for knowledge acquisition, but also major challenges including data interoperability and effective management. The great volume of miscellaneous data renders the generation of new knowledge a complex data analysis process. Presently, big data technologies provide multiple solutions and tools towards the semantic analysis of heterogeneous data, including their accessibility and reusability. However, in addition to learning from data, we are faced with the issue of data storage and management in a cost-effective and reliable manner. This is the core topic of this paper. A data lake, inspired by the natural lake, is a centralized data repository that stores all kinds of data in any format and structure. This allows any type of data to be ingested into the data lake without any restriction or normalization. This could lead to a critical problem known as data swamp, which can contain invalid or incoherent data that adds no values for further knowledge acquisition. To deal with the potential avalanche of data, some legislation is required to turn such heterogeneous datasets into manageable data. In this article, we address this problem and propose some solutions concerning innovative methods, derived from a multidisciplinary science perspective to manage data lake. The proposed methods imitate the supply chain management and natural lake principles with an emphasis on the importance of the data life cycle, to implement responsible data governance for the data lake.

Keywords: data lakes; data governance; sustainability; supply chain management; natural lake; ecosystem (search for similar items in EconPapers)
JEL-codes: O3 (search for similar items in EconPapers)
Date: 2020
References: View references in EconPapers View complete reference list from CitEc
Citations:

Downloads: (external link)
https://www.mdpi.com/1999-5903/12/8/126/pdf (application/pdf)
https://www.mdpi.com/1999-5903/12/8/126/ (text/html)

Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.

Export reference: BibTeX RIS (EndNote, ProCite, RefMan) HTML/Text

Persistent link: https://EconPapers.repec.org/RePEc:gam:jftint:v:12:y:2020:i:8:p:126-:d:390384

Access Statistics for this article

Future Internet is currently edited by Ms. Grace You

More articles in Future Internet from MDPI
Bibliographic data for series maintained by MDPI Indexing Manager ().

 
Page updated 2025-03-19
Handle: RePEc:gam:jftint:v:12:y:2020:i:8:p:126-:d:390384