Data diagram

Data Mesh: Data Lake Evolution

,

Motivated by the exciting presentation "Data Mesh in Practice - How Europe's Leading Online Platform for Fashion Goes Beyond the Data Lake"1 by Max Schultze (Lead Data Engineer at Zalando) and Arif Wider (Lead Technology Consultant at ThoughtWorks),

we would like to describe in this article why a Data Lake is not the solution for all data problems of modern and networked systems and which of these problems can be better solved with the data mesh approach.

What exactly is a data lake? A data lake is a very large data store that holds data from a wide variety of sources in its raw format. It can contain both unstructured and structured data and can be used for big data analysis. (https://www.bigdata-insider.de/was-ist-ein-data-lake-a-686778/). Examples of data sources can be ERP or CRM systems, but also systems for product usage, sales, website visit data or other systems.

The goals of big data analyses can include, for example, finding relevant customer groups, customer behavior patterns or deriving product improvements.

In the meantime, data lakes have gone from being a hype topic to the "Valley of tears" relegated. The great expectations placed on the Data lake approach have only been partially fulfilled.
While the approach - bringing together data from multiple sources and formats to gain new insights - is very promising at the outset, it tends to turn into more of a Swamp of datawhich can no longer be kept track of, but which nobody wants to switch off.2 Once this situation has occurred, the data lake costs money without generating any benefit.

A major problem in most cases is that there is no clearly defined responsibility for maintaining good data quality gives: Correctness, relevance, reliability and consistency of the data or availability on different systems.3

Challenges with data lakes

The organization often Existing data sources connected to the data lake. The Source systems have no direct advantage from the use of their data by others (e.g. in the context of analyses). They therefore have little or no motivation to adapt their systems so that they generate data of good or better quality. On the other hand, the users of the data lake are faced with a non-transparent accumulation of datawhich they cannot penetrate. In case of doubt, they turn to the operators of the data lake, as they do not know the source systems of the data in the lake (these are supposed to be abstracted to a certain extent by the data lake). However, the operators of the data lake often have no knowledge of the content of the data. They are focused on providing a technical platform for data storage and use. When it comes to content-related issues, they can only act as an intermediary between data producers and data users. This makes them a bottleneck for mediation, which can sooner or later lead to dissatisfaction on all sides.

The Centralized model can work for organizations that have a small number of different consumption cases. For a company with a large number of sources and a diverse consumer base, failure is inevitable.4

Monolithic Data Platform
View of the monolithic data platform - Source: https://martinfowler.com/articles/data-monolith-to-mesh.html

Data mesh approach breaks down barriers

To achieve this Front formation to break upthe data mesh approach brings the idea of so-called "Data products". A key influencer of this new trend is an old acquaintance: Martin Fowler. The idea is similar to what we know from the agile approach and other approaches such as DevOps, Microservices etc. or from the Functional Owners (FOs): Instead of a separation of responsibilities into layers (analogy: frontend, middleware, database), there is a Focus on a "product"which is then developed and maintained holistically across all sub-areas. These products can of course also be networked and combined as required.

Data mesh is an architectural and organizational paradigm that challenges the age-old assumption that we must centralize big analytical data to use it, have data all in one place or be managed by a centralized data team to deliver value. Data mesh claims that for big data to fuel innovation, its ownership must be federated among domain data owners who are accountable for providing their data as products (with the support of a self-serve data platform to abstract the technical complexity involved in serving data products); it must also adopt a new form of federated governance through automation to enable interoperability of domain-oriented data products. Decentralization, along with interoperability and focus on the experience of data consumers, are key to the democratization of innovation using data. Source: https://www.thoughtworks.com/de/radar/techniques/data-mesh

The Data Product Owner

Data Production
Data Production - Source: own presentation

For each of the data products offered within an organization, there is then a Data PO (Data Product Owner)which has the following tasks:

  • He is responsible for ensuring that his product "works". To this end, he is in contact with the users of the data product as well as with the "manufacturers" (development team, data engineers, data scientists) and the "suppliers" (source systems).
  • He maintains his data product like a real product: he sells and explains his product, collects suggestions for improvement from his customers, develops new product ideas, checks whether his product is still "profitable" or is being used and supports it throughout its life cycle.

There can be a separate team behind each data product. This makes it easier to scale the number of data products with this approach than with the central data lake, where the operating team is in principle responsible for all products created from the data lake data. There can be simple data products (data products that are very close to the source data) and complex data products (data products that are composed of source data from several sources or even other data products).

Positive user experience is crucial

A data product should consist not only of the user data itself, but also of a description of the data (what exactly does the data say) and other metadata (e.g. how many data records are there, how up-to-date are they, how often are they used). If only the user data is available, a user requires instructions on how to use it. They cannot use the data independently - which ultimately makes the "User Experience" of the data has a negative impact. It has a positive effect on the user experience if the data is presented in a common standard are available and the Standardized access is possible. In order to maintain a good data user experience, a Central committee which defines standards across all data products and enables their implementation.

These thoughts are clearly summarized in the following graphic:

Ecosystem of data products
Ecosystem of data products - Source: https://www.slideshare.net/ArifWider/data-mesh-in-practice-how-europes-leading-online-platform-for-fashion-goes-beyond-the-data-lake

This recognizes both the individual data products (and the requirements for these), which can be related to each other and are each managed by a data product team. The provision of the data products can be based on a common "Data Infra as a Platform", which defines standards and ensures the simple production and use of data products without having to know the data in the platform in detail (that's what the data product teams are there for).

Conclusion

Even a data mesh does not solve all the problems of modern networked IT systems. However, it expands the data lake to include the idea of data products and the approach of overarching processing of this product.

In manageable environments, a data lake itself can be a good starting point for the Bringing data sources togetheror data offerings and data users or data requirements. A data lake is successful as long as the Manageable number of data sources and data users is and from a central team can be managed. However, if the number on one or even both sides gets out of hand, the administration of the data lake alone is time-consuming. Content-related tasks, such as ensuring data quality, suffer as a result of the pure administrative effort.

If there is then a very dynamic environment If this is added to the mix of data sources and data users (high-frequency changes to interfaces, data models, etc.), as is increasingly the case in modern microservice architectures, dissatisfaction on all sides is inevitable. The data lake and the data lake team become a bottleneck for all parties. Declining data quality turns the data lake into a data swamp.

This is where the Data Mesh comes in. Through the Separation of the management of data lake infrastructure from the content-related consolidation of data into data products there is a clear focus and thus a division of tasks.

The focus of cross-divisional teams on concrete data products From the data source(s) to the data users, these teams offer the opportunity not only to manage data, but also to focus on data content. Among other things, this ensures data quality and a good user experience of the data.

But even this clear focus on data products is only successful if all parties are in harmony. All technology and all processes are useless if people (data users and data providers) do not talk to each other, or if they are not supported by data product teams in the data centers. Dialog to coordinate their needs and plans with each other.

 

Do you need support in identifying data offerings and data requirements in your company? Then take a look at here over.

You may also be interested in this blog post on the topic of Data Lake: How much data lake is in a scalable data hub?


Sources:

1 https://www.meetup.com/de-DE/ThoughtWorks-Muenchen/events/270948960/

2 https://www.computerwoche.de/a/wie-ihr-data-lake-sauber-bleibt,3331146

3 https://www.alexanderthamm.com/de/blog/die-5-wichtigsten-massnahmen-fuer-eine-optimale-datenqualitaet/

4 https://martinfowler.com/articles/data-monolith-to-mesh.html

 

 

 

Marc Mai

About ME

Marc Mai studied Business Informatics (M.Sc.) and has been supporting companies in their IT development at doubleSlash since 2013. Data-Driven Journey. As a data architect, he develops cross-industry end-to-end solutions for data enablement - from the design of modern data lakehouses and intelligent data integration to the development of data products that generate real business value. Marc Mai combines technical expertise in backendArchitectures with a strategic understanding of data culture. His mission: to enable organizations to use data as a strategic competitive advantage and drive AI-supported innovation.

All contributions from Marc Mai

Learn more

Further information on our website and in our newsletter

Arrow up