Interface documentation with Swagger

Business intelligence (BI) and analytics with the Scalable Data Hub

The concept of classic data warehousing can no longer meet today's requirements for efficient data management in terms of availability, data diversity and flexibility at an acceptable cost-benefit ratio.

iStock_000027365243_bloggroesse

Good scalability is not a given in traditional solutions, neither from a technical nor an economic point of view. High costs per terabyte make it uneconomical to store both aggregated data and raw data over long periods of time. Business intelligence (BI) and analytics with existing, proprietary technologies therefore often exceed the available budgets.

The holistic approach of a scalable data hub takes up the concept of classic data warehousing to some extent, but uses Hadoop However, the company is focusing on a framework for highly scalable, massively distributed data processing as a new technological basis.
This article highlights the problems of traditional data warehouses (DWH) and describes how the Scalable Data Hub with Hadoop opens up new possibilities in terms of Performance, granularity and the integration of unstructured data opened.

Problems and weaknesses of classic data warehousing

The idea behind business intelligence and data warehouses, which was developed in the 1990s, has proven its worth to this day. The data warehouse integrates all relevant data from heterogeneous and distributed data sources and provides comprehensive views that can be used by Evaluation tools easily processed can be used. However, the classic data warehouse approach in this form brings with it some fundamental problems:

- In order to be able to react agilely to new business situations, specialist users such as controllers, managers or analysts require granular data. In traditional DWHs, however, this data is sometimes not stored at all or is moved to offline archives too soon after loading. It is therefore no longer possible to process and evaluate this data. Business users therefore quickly reach their limits with current systems.
- Often, only the data for which the department has already defined analytical requirements is loaded into the DWH. It is therefore not possible for business users to determine which new Use cases could be realized with new data (not available in the DWH) - important potential is being wasted.
- Unstructured data are hardly present in classic DWH. Big data analyses are therefore not fully feasible. Historical data are only kept to a limited extent for cost reasons, which in turn makes many approaches to predictive analytics impossible.
- The scalability with regard to increasing data volumes and users is basically given with classic DWH. However, the resulting costs for hardware, software and operation increase almost linearly and are therefore no longer in proportion to the benefits.
- To achieve an acceptable Performance many DWH databases require a complex management of Aggregates, Partitioning and Indices. This approach is suitable for reporting and, to a limited extent, for drill-down analyses. For ad-hoc formulated queries, such as Data Exploration However, it cannot provide the required response times.

Apache Hadoop: Advanced architecture for big data and analytics

Many ideas and concepts have been developed in recent years to find solutions to the problems listed. Since then, many companies have carried out initial proof of concepts (PoCs) with new big data analytics technologies. Various approaches have been investigated, such as Hadoop, predictive analytics, in-memory databases (IMDB) or a combination from these. However, the architectures of these PoCs are usually very specifically tailored to the underlying use cases. A transfer to other use cases is only possible with a high degree of customization.
But how could an architecture for all BI and big data analytics use cases in a company be structured? For this future holistic architecture contributes above all to the Hadoop technology with.

Apache Hadoop - highly scalable with distributed data processing

Apache Hadoop is a free framework, written in Java, for highly scalable, massively distributed Data processing. The central elements are the MapReduce alogrithm and the Hadoop file system (HDFS) is the result. MapReduce parallelizes the data processing and distributes it to all participating nodes of the computer cluster, whereby considerable speed advantages can be achieved. In HDFS, extremely large amounts of data (petabyte range) can be stored cost-effectively and accessed in parallel.

Hadoop has matured significantly in recent years and has proven itself for cross-industry use in companies. Today, other Apache open source projects are grouped around the core of distributed data storage in HDFS and processing in MapReduce in areas such as Data access, integration, security and operation. Essential further developments in the Hadoop ecosystem are, for example:

- Ad-hoc data access via interactive SQL interfaces.
- Tight integration with leading providers of predictive analytics solutions.
- SQL-based tools now also enable specialist users to access the HDFS.
- The topic Data protection was previously a weak point of Hadoop. Today, Hadoop management tools are available that Audit accesses and access rights at the level of Hadoop files, schemas, tables and views.

Scalable data hub based on Hadoop: Powerful architecture to solve the problems of traditional data warehousing

With the Scalable Data Hub, doubleSlash has developed a new and holistic concept for BI and analytics. Similar to the classic DWH, the Scalable Data Hub covers the Connection, integration and aggregation of data from heterogeneous sources from. With Hadoop, however, the scalable hub relies on a Powerful architecturewhich the Processing complex data and providing advanced functional services is made possible. Using this technical development, the Scalable Data Hub offers solutions to most of the problems of the classic DWH:

ProblemSolution in the Scalable Data Hub
Cost-benefit ratioThe underlying software for Hadoop is largely open source. Hadoop itself does not place any special requirements on the hardware. Existing systems can be used or replaced by inexpensive available hardware. This makes Hadoop the most cost-effective form of data storage for large data volumes today and up to a factor of 1,000 cheaper than classic DWH RDBMS.
Unstructured dataThe Scalable Data Hub enables data storage and provision for both structured and unstructured data (e.g. sensor and RFID data, emails, log files or data from the connected vehicle).
Granular dataIn contrast to traditional DWH, the data in the Scalable Data Hub is kept granular (often also raw data). This prevents premature aggregation of the data and the associated exclusion of potential new analyses. This is particularly relevant for predictive analytics.
Historical dataDue to the favorable cost-benefit ratio, the Scalable Data Hub also enables the storage and provision of data that is rarely accessed, such as data that must be kept for compliance reasons.
Data without a technical requirementThe Scalable Data Hub serves as a central platform for the provision of all resulting data for business users. New and meaningful analytical use cases can be identified and implemented on this comprehensive basis.

The main arguments in favor of Hadoop are the Very favorable ratio of costs per data volume and the Scalability with growing data volumes. With the Scalable Data Hub based on Hadoop, companies can build their data management in a sustainable and future-proof way, restructure their IT costs and gain the necessary flexibility for further technical development and analytical use cases.

Conclusion

Hadoop is one of the core elements of a new and holistic architecture for BI and big data analytics. The concept of the Hadoop-based scalable data hub takes up the idea of the classic DWH, but is technically superior to it and meets today's requirements - such as cost-effective data volumes, scalability and a variety of analyzable data types. A data hub based on this powerful technology represents a future-proof platform for BI and analytics.

 

Find out more about the Scalable Data Hub here

 

Sandra Rueß

About ME

Sandra Rueß studied Business Informatics with a focus on Business Engineering (Bachelor of Science). She has been working at doubleSlash since 2015 and, in her role as Business Consultant, specializes in the following areas Requirements management, conception and IT design specialized.

All contributions from Sandra Rueß

Learn more

Further information on our website and in our newsletter

Arrow up