TY - GEN
T1 - Towards a Semantic Extract-Transform-Load (ETL) framework for big data integration
AU - Bansal, Srividya
PY - 2014/9/22
Y1 - 2014/9/22
N2 - Big Data has become the new ubiquitous term used to describe massive collection of datasets that are difficult to process using traditional database and software techniques. Most of this data is inaccessible to users, as we need technology and tools to find, transform, analyze, and visualize data in order to make it consumable for decision-making. One aspect of Big Data research is dealing with the Variety of data that includes various formats such as structured, numeric, unstructured text data, email, video, audio, stock ticker, etc. Managing, merging, and governing a variety of data is the focus of this paper. This paper proposes a semantic Extract-Transform-Load (ETL) framework that uses semantic technologies to integrate and publish data from multiple sources as open linked data. This includes - creation of a semantic data model to provide a basis for integration and understanding of knowledge from multiple sources, creation of a distributed Web of data using Resource Description Framework (RDF) as the graph data model, extraction of useful knowledge and information from the combined data using SPARQL as the semantic query language.
AB - Big Data has become the new ubiquitous term used to describe massive collection of datasets that are difficult to process using traditional database and software techniques. Most of this data is inaccessible to users, as we need technology and tools to find, transform, analyze, and visualize data in order to make it consumable for decision-making. One aspect of Big Data research is dealing with the Variety of data that includes various formats such as structured, numeric, unstructured text data, email, video, audio, stock ticker, etc. Managing, merging, and governing a variety of data is the focus of this paper. This paper proposes a semantic Extract-Transform-Load (ETL) framework that uses semantic technologies to integrate and publish data from multiple sources as open linked data. This includes - creation of a semantic data model to provide a basis for integration and understanding of knowledge from multiple sources, creation of a distributed Web of data using Resource Description Framework (RDF) as the graph data model, extraction of useful knowledge and information from the combined data using SPARQL as the semantic query language.
KW - Big data
KW - Data integration
KW - Ontology
KW - Semantic technolgies
UR - http://www.scopus.com/inward/record.url?scp=84923917788&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=84923917788&partnerID=8YFLogxK
U2 - 10.1109/BigData.Congress.2014.82
DO - 10.1109/BigData.Congress.2014.82
M3 - Conference contribution
AN - SCOPUS:84923917788
T3 - Proceedings - 2014 IEEE International Congress on Big Data, BigData Congress 2014
SP - 522
EP - 529
BT - Proceedings - 2014 IEEE International Congress on Big Data, BigData Congress 2014
A2 - Chen, Peter
A2 - Chen, Peter
A2 - Jain, Hemant
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 3rd IEEE International Congress on Big Data, BigData Congress 2014
Y2 - 27 June 2014 through 2 July 2014
ER -