Data quality

How can data quality be optimized in practice?

,

As part of my bachelor's thesis, I dealt extensively with the quality of data and created two artefacts to optimize it: a catalog of measures and an associated process model.

What is data quality?

Measures are decided and decisions are made on the basis of data. The quality of the data therefore plays a major role. Data quality indicates the "degree of suitability of data to fulfill the purpose for which it was collected or generated" [1]. Data quality therefore indicates the extent to which the available data is suitable as a basis for further planning. Data quality is subjective and depends a lot on feelings and individual assessments, especially in companies [2]. This is mainly because every company, every system and every employee needs different data and has different perceptions of its relevance [3].

Differentiation between technical and functional data quality

At doubleSlash and also in the context of the thesis, a distinction is made between functional and technical data quality. We have defined both as follows:

Technical data quality

  • technical factors of the data and the system of their storage
  • Base layer (elementary basic components of a data set)
  • Zero values in columns or rows
  • Format of the data (e.g. XML, JSON, etc.)
  • Formatting of individual values (e.g. addresses, dates)
  • Adjusting the units (e.g. liters, milliliters)
  • Table structures

Technical data quality

  • Data content and links
  • Conceptual data models
  • Definition of parameters, responsibilities, terms and boundaries
  • Domain knowledge of employees (distinctive knowledge base in a respective specialist area/project)

Technical data quality is about an attribute being defined in accordance with the rules, i.e. correct syntax. In contrast, functional data quality is about the logical context of the data, i.e. correct semantics [4]. For example, as shown in the following figure, an attribute, in this case "date of birth", can be of good technical quality, while the functional quality is poor:

Syntax and semantics
Figure 1: Example of technical and functional data quality; source: own illustration

Data quality criteria

There are various criteria for assessing data and information quality. The following figure shows the 15 most important criteria according to the German Society for Information and Data Quality (DGIQ):

DGIQ Dimensions
Figure 2: The 15 most important criteria for data quality; source: own illustration based on [3]

The criteria are divided into four categories, in each of which different aspects must be considered. Overall, data and information quality is determined by all dimensions together, as each dimension is a critical success factor for a functioning overall system. It is therefore not enough for only certain dimensions to be of high quality; each individual dimension must have a high level of quality.

Creation of artifacts for the assessment of data quality

Since data quality is highly relevant, but its assessment is often not so trivial, I have created two artifacts to better assess data quality: a catalog of measures and an associated process model.

To create these artifacts, I conducted interviews with colleagues at doubleSlash who are involved in projects in the E-mobility sector are active. From the interviews, the Problems identified in practice and from this Requirements for the catalog and the process model was defined. It was also found that many companies are aware of the importance of data quality for the success of a project, but often fail to implement it due to its complexity. The concepts and models exist in theory, but are rarely used in practice due to the excessive effort involved. There is no compact procedure for optimizing data quality that can actually be used in projects.

The two artifacts created are intended to offer precisely this compact and clear approach to gradually improve data quality.

Catalog of measures to optimize data quality

The catalog of measures developed contains a total of 21 measures that can be used to optimize data quality. Here is a small excerpt from the catalog:

Excerpt from the catalog of measures
Figure 3: Excerpt from the catalog of measures

There is a detailed version of the catalog that contains specific implementation instructions for each measure and a corresponding DIN A4 page that lists all measures in a table and divides them into categories. In addition, there is a note column in this table for each measure, in which it is briefly mentioned which aspects are explained in more detail in the more detailed version. Finally, there is a column indicating which data quality criterion is mainly improved by the implementation of the measure.

Structure of the process model

The process model can be used as a guideline for the catalog of measures. It specifies a sequence for the implementation of the measures and uses questions to check whether the implementation of certain measures in a category appears to make sense in the respective application context.

And this is what the process model looks like:

Procedure model
Figure 4: Process model for optimizing data quality, Source: Own illustration

Application in practice

I applied the created artifacts to a sample data set from the automotive industry. To do this, I first used the process model and followed the corresponding path based on the questions. In doing so, I discovered that the organizational factors in this project fit so far, but that data cleansing had not yet taken place. For this reason, measures 13-18 were applied to the data set in order to be able to analyze it better.

Procedure model practice
Figure 5: Process model in the application; Source: Own illustration

In measure 13, for example, the duplicates were removed and in measure 16 the time stamps were brought into a standardized format.

Conclusion: Data quality can be achieved more effectively with the catalog of measures and the process model

Many companies are aware of the high relevance of good data quality. However, the requirements vary depending on the industry and data set. Optimizing data quality is a multi-complex topic and therefore very time-consuming. Many companies do not have the time to deal with this topic in detail. For this reason, the catalog and the process model are intended to provide added value by offering a time-saving way of finding the right solution for your own company. The use of the process model and the associated implementation of the measures is not a one-off process. To ensure a permanently high level of data quality, all steps must be run through and adapted regularly.

 

Learn more about software development

Sources

[1] Würthele, Volker (2003): Data quality metrics for information processes. Data quality management by means of holistic measurement of data quality. Dissertation. Swiss Federal Institute of Technology Zurich, Zurich.
[2] Rohweder, Jan P.; Kasten, Gerhard; Malzahn, Dirk; Piro, Andrea; Schmid, Joachim (2021): Information quality - definitions, dimensions and terms. In: Knut Hildebrand, Marcus Gebauer and Michael Mielke (eds.): Data and information quality. Wiesbaden: Springer Fachmedien.
[3] Harrach, Hakim (2010): Risk assessments for data quality. Concept and realization. Wiesbaden: Vieweg + Teubner Verlag | Springer Fachmedien.
[4] Weber, Kristin; Klingenberg, Christiana (2020): Data Governance: The guide for practice. Munich: Carl Hanser Verlag.

Janina Stifel

About ME

Janina Stifel completed her Bachelor of Science in Business Informatics at the HTWG in Constance and is currently doing her Master's degree at the Technical University of Munich. She has been working at doubleSlash since 2020, where she supports the Design & Development department. She is particularly interested in business process modeling and the Data processing and visualization.

All contributions from Janina Stifel

Learn more

Further information on our website and in our newsletter

Arrow up