Everyone in a packed courtroom holds their breath. A lawyer has just quoted several court rulings to support his client's claim. But now the judge looks sternly over the edge of his glasses: "These cases... don't exist." The lawyer is stunned - the AIwhich he trusted blindly, deceived him. ChatGPT had provided him with the verdicts with completely invented names and details.1
What sounds like a tech thriller is reality. The incident made headlines in 2023 and is emblematic of a problem that has become even more pressing since then: even the most modern AI systems make serious mistakes.
At random: unpredictable AI outputs
Why do such mistakes happen?
Generative AI works probabilistically. This means that it selects words or solutions on the basis of probabilities. The model rolls the dice for each new query. This means that identical inputs can generate different answers. On the one hand, this randomness leads to creative diversity. On the other hand, it leads to unpredictability, which undermines the reliability of AI models. A recent example of this "capriciousness" of AI models was provided by Google's Gemini language model. Normally it gives harmless answers - but sometimes it gets completely out of hand. For example, Gemini insulted a user, saying that he was a "blight on the universe" and wishing him to "please die".2
Such disturbing lapses illustrate how unpredictable the responses of modern AI models can be when certain conditions come together unfavorably.
Hallucinations instead of facts: When AI invents content
Closely related to the random principle is the phenomenon of AI hallucination. AI outputs content that is false or fictitious. AI models generate answers based on their training data. If this data is incomplete or misleading, the AI "guesses" and presents made-up information in the guise of apparent facts. However, this is often not recognizable to the user, as the AI always presents these false facts in an absolutely convincing way.
The New York courtroom from the introduction is a perfect example of this, as ChatGPT completely hallucinated several court rulings. The lawyer had no reason to doubt the seriousness of the case and thus ended up in an embarrassing situation in court. He is not alone in this, as incidents like this continue to cause a stir around the world. In the past, AI chatbots have repeatedly quoted imaginary scientific articles, invented laws and accused people of false deeds.
Outdated knowledge: A questionable database
Generative AI models know a lot, but by no means everything. Even if a model has devoured millions of documents, it is still unaware of world events after the date of its training. Many well-known AI systems therefore have a level of knowledge that is several months or years behind the present. The half-life of AI knowledge is short because reality is constantly creating new facts. Models such as Bing Chat or newer GPT versions try to solve this by integrating live web search. However, this only works well if the AI classifies the information it finds correctly, accesses reliable sources and does not generate complete nonsense on this basis.
Model collapse: when AI learns from AI
Model collapse is a phenomenon in which generations of AI models are increasingly derailed by the training of AI-generated data. Put simply, if AI A generates content and AI B learns it again, the quality of the output continuously decreases. This phenomenon has already been observed in a number of experiments. In one study, for example, researchers first trained a language model on purely human data and then used its own generated texts again as a training basis for the next model. They repeated this several times, using the responses of new models to train the next generation. The result was alarming. Even the 10th "generation model" only spat out meaningless gibberish when asked a test question. One of the scientists involved described: "At some point, the model was practically completely meaningless." This is precisely the phenomenon of model collapse.3
It is therefore important for users and companies to understand this: The quality of AI output depends fundamentally on the quality of the training data. If the data is outdated, distorted or of dubious origin, this is inevitably transferred to the answers of the models. Vigilance with regard to the database therefore becomes a basic requirement if AI systems are to be used in a trustworthy manner.
Conclusion: Stay vigilant when dealing with GenAI
Artificial intelligence is a powerful tool and, as with any new instrument, we must learn to use it responsibly. An autopilot in an airplane relieves the pilot of work, but it does not relieve him or her of the need to monitor the instruments and intervene in an emergency. The same applies to generative AI. If we treat its outputs with healthy mistrust, check the results and make corrections, then GenAI can enrich our everyday lives without becoming a danger.



