Graphic-tablet-with-icons

How much data lake is in a scalable data hub?

To begin with, it can be said that the terms Data Lake and Scalable Data Hub are not synonymous with each other. But they are not completely different either. The Scalable Data Hub has already been described in the blog post Business intelligence (BI) and analytics with the Scalable Data Hub very well described and from the classic Enterprise Datawarehouse (EDW) demarcated. However, the EDW is not the only system whose positive characteristics are incorporated into the Scalable Data Hub. The properties of a data lake are very well suited to compensating for and supplementing the weaknesses of an EDW.

Scalable data hub meets data lake

A scalable data hub should be characterized by the Use of scalable technologieselastic at the changing demand for computing and storage resources can be adapted. This means that it can be operated on a minimal scale when data volumes are low. When data volumes increase, the Scalable Data Hub can "grow" accordingly and "shrink" again when demand decreases. This feature enables a Efficient use of resourceswhich in the Big data environment is an advantage.
 
Advances in the areas of Industry 4.0 and IoT are also increasing the Amount of data from different data sourcesthat need to be processed. What's more, systems must not only be able to process data, cope with large amounts of datathey must also monitor and in the event of anomalies with respond with low latency. The pursuit of Cost efficiency in this context is an additional trigger for the search for a system that can meet these challenges.
 
This is where the properties of a data lake come into their own in the Scalable Data Hub. The term "data lake" was originally coined in the Hadoop Community and describes a Storage for raw data. This storage system solves Previous data silos and bundles the data collected in it centrally.
 
The loss of information already mentioned by Sandra Rueßthat can occur during an Extract, Transform, Load (ETL) process is minimized in the Data Lake by the Storage of raw data bypassed. This means that large volumes of data can be collected from a wide variety of source systems.
 
This is done under the Use of horizontally highly scalable technologiesas used, for example, by the Hadoop ecosystem and its components can be offered. In the Hadoop Distributed File System (HDFS), you can Raw data stored to be used at a later point in time in a technical context. With horizontal scaling, additional nodes (computing units) can be added to a computer cluster and released again. This technically enables the previously mentioned Elasticity achieve[2].

Added value of the Data Lake for the Scalable Data Hub

 
As Central data hub the Scalable Data Hub offers a Central, complete and consistent databasewhich serves as the basis for business intelligence (BI) and data analyses, among other things. This database also includes Raw data that has been processed without can be loaded into the Scalable Data Hub. It is precisely this point of raw data storage that is covered by the Properties of the data lake supported.
 
To ensure that the raw data can still be used after a long period of time, the Collection and management of metadata an essential component of a data lake. The metadata can, for example Information about the origin, quality, quantity and structure of the data give. For Explorative data analyses this metadata provides a valuable basis. For example Information about the data sources The quality of the data can be determined by comparing how good the data quality is in comparison to other sources. For Use cases from the areas of Industry 4.0 and IoT this would mean, for example, that sensors could be monitored preventively. A deterioration in the data quality of a sensor could therefore indicate an imminent defect in the component.

Raw data and its influence on Enterprise Ware House and Scalable Data Hub
The use of typical EDW functions is of course not excluded by the influence of the data lake in a scalable data hub. On the contrary! One Raw data basis provides the optimal data basis for ETL processes in the sense of classic EDW. Using the conventional approach Schema-On-Write[1], the user had to think carefully about which data he needed and which he did not need before loading the data. A later Changing the data models or data structureschanging technical requirements, is extremely easy to implement using the Schema-On-Write. elaborate.
The user of the Scalabale Data Hub, on the other hand, can download the so-called Schema-On-Read[1], whereby the technical requirements for the data structures and models can be oriented to the observation of the raw data. This means that previously undiscovered Connections between data can be identified from various data sources. This expands the possibilities for Obtaining new key performance indicators (KPIs).
 
As you might expect, the Requirements of a data lake for data storagethrough a "simple" relational database not fulfilled become. Like the Data Lake, the Scalable Data Hub therefore relies on Highly scalable technologieswho are committed to the Storage of raw data suitable. It is important that the Data stored independently of its structure can be used. For this reason, Hadoop can be used for the operation of a data lake with its Hadoop Distributed File System (HDFS) can be used as storage. Many Hadoop-related technologies have already proven themselves in practice in this context.
 
As the Scalable Data Hub with its centralized data storage one Single-Point-Of-Truth the company, it also bears responsibility for the Security of the data it contains. These include, for example how long the data is stored in the Data Hub should be, who has access to the data and which actions are performed on the data may be used. The Data Lake offers concepts for the Dealing with data security, governance and lifecycle managementwhich are essential for a central data hub such as the Scalable Data Hub. This includes the support of Procedure for authentication and authorization of users and systems that want to interact with the Scalable Data Hub. A data lake also contains components that Monitor data life cycles and the Guidelines for uniform data quality guarantee.
 
Conclusion
The Scalable Data Hub therefore utilizes the strengths of a data lake in conjunction with big data technologies. Because for a Central data storage it needs Resilient, high-performance and scalable technologieswhich is a High availability of data guarantee. The resulting Possibilities for data analysis from all disciplines are enormous. The Data Lake therefore offers very good Supplementary properties which also Impact on the security, lifecycles and quality of data in a scalable data hub have.


Sources:

[1] https://blogs.oracle.com/datawarehousing/entry/big_data_sql_quick_start10
[2] https://de.wikipedia.org/wiki/Rechnerverbund

Find out more about the Scalable Data Hub here

 

Patrick Treiber

About ME

All contributions from Patrick Treiber

Learn more

Further information on our website and in our newsletter

Arrow up