← All insights

The AI-Generated Information Trap: When AI Systems Negotiate and the Human Only Mediates Between Them

By · Published: 3 September 2026

Decision-MakingAI GovernanceBoard of DirectorsTechnology Competence
The AI-Generated Information Trap: When AI Systems Negotiate and the Human Only Mediates Between Them

Article series: Limits of delegating to AI systems (Part 4)

Not long ago a review document for a software evaluation landed on my desk. More than a thousand assessment points, neatly structured across several levels in Excel. Many of the criteria make good sense on their own, but in this quantity they add no recognisable value. Within a few minutes it was clear that the author is very probably not an expert in software evaluations and therefore cannot assess the catalogue either. The expected outcome: given the number of measurement points, no human being will be able to complete this catalogue with any sensible amount of effort.

The software vendors who receive such a requirements catalogue will therefore use AI as well, simply for reasons of efficiency. In the end the author needs AI again to evaluate the answers at all, because the volume of data has grown and the specialist knowledge is missing. Whether the criteria that actually matter are met is hardly visible from such an exercise. The troubling part: at this point the human being has effectively left the procedure.

Generating information without being able to judge it

A change of scene, to contracts: someone who is not a lawyer now drafts a contract with AI that covers many aspects a layperson would probably never have thought of. That is a genuine gain in time and cost. However, the layperson probably also fails to notice when one clause sits among the useful ones that does not belong in this contract, or that even weakens their position. What used to be a simple contract between two parties now grows into a lengthy instrument without any regulatory pressure, because AI can produce content with ease.

The similarity to a software requirements catalogue is high. A user describes their own requirements precisely. As soon as the AI adds points outside their field, the result is a catalogue that no vendor can satisfy. Even when the answers arrive, the user cannot judge what they imply.

The ability to generate information appears to have slowly outgrown the ability to judge it.

The reason is usually banal: people try to save the cost and time of a lawyer and of a subject matter specialist, and they receive a usable answer very quickly. In the end nothing is saved, because the cost moves to those who then have to work with the inflated document and carry the responsibility. A catalogue produced in an hour can create several working days of effort on the other side. The costs do not disappear, they only change who bears them.

What the AI-generated information trap leads to

It follows from the economic logic that machine-generated content draws a machine-generated answer, and that this answer then requires a machine-generated evaluation.

Applied to the example above: one system produces the requirements catalogue. A second one completes it. A third evaluates the result and condenses it, at the end of the process, into a few pages. In between sits a human being as an interface, who finally approves the result or signs the contract without the necessary specialist competence.

The human being sits between the systems as an observer and an interface. They are still «human-in-the-loop». But are they still «human-in-command»?

The person at the end of the process chain sees the summary of an exchange that two systems conducted between themselves, because the volume of information in between is no longer something they can process. They therefore cannot judge what was left out, distorted or misunderstood. In the condensation produced by the AI this stays invisible, because a good summary looks complete and rounded.

One could object that a second model is able to uncover such gaps and improve the result. That cannot be relied upon. I described this in an earlier article: the errors of different models are strongly correlated, particularly among the larger ones. The blind spot in the catalogue and the blind spot in the answer come from the same distribution. The evaluation does not find what the catalogue never asked about.

Why do we accept these inefficiencies?

There is a research finding that is little known in boardrooms and executive teams:

Raja Parasuraman and Dietrich Manzey brought the research on automation complacency together in 2010. Two results are surprising: the effect hits specialists just as hard as laypeople, and practice and repetition of the process do not remove it. There is also a reversal that hardly anyone expects: the more precise and reliable a system is, the harder oversight becomes, because trust reduces error detection or, in the worst case, replaces it.

Fabrizio Dell’Acqua and colleagues measured this in Organization Science in 2026. 373 consultants were given a task in which the tabular data looked complete, and only the interview notes led to the correct answer. Without AI, 84.5 per cent were correct, with AI it was 60 and 70.6 per cent respectively. What is notable: the weaker of the two AI groups was the one that had additionally received an introduction on how to work with the model. More guidance led to more trust and therefore to more errors.

The decisive point is what stands next to it: the quality of the recommendations rose with AI even among those who were wrong. People working with AI thus arrived at the correct answer less often, but they justified their wrong answer better than the control group justified its correct one.

The system did not fail visibly, it failed convincingly.

What this means for the Board of Directors

In Part 2 of this series I argued that the ability to provide evidence has to take over where traceability is missing: monitoring, red teaming, re-approval, complete documentation. Part 3 dealt with the way the approved performance profile shifts quietly once the system is in operation. Both times the system was at the centre. This time it is the document on which the decision is taken.

I still hold my position from Part 2 of the series, but it has a limit that I did not name at the time.

Evidence can be generated by machine. A complete dossier does not prove that a human being examined it critically and thought it through.

At the end of such a process there is a complete dossier that looks flawless on the surface. Requirements catalogue, offers, weighted evaluation, proposal, minutes. Art. 717 OR requires due care, Art. 716a OR requires ultimate oversight, and both look satisfied, because the necessary information is there. Yet in the worst case, at no point did human judgement enter that went beyond an approval.

The reflex now is to increase the review capacity: more time, additional pairs of eyes, another loop through a human team. That will not be enough, because the volume produced grows faster than these countermeasures take effect. When automation bias and automation complacency come together under pressure of time and cost, little changes in the result. A system that works in principle is trusted until the opposite is proven.

Recommendations that work

I see four approaches for addressing these problems:

  • Reverse the process: your own critical and complex thinking first, only then the use of AI. Anyone who has already written their own draft and brings in AI afterwards, in order to illuminate blind spots or point out missing aspects, achieves far better results than having to analyse and assess a large machine-generated volume of data at the end. AI then becomes a genuine help and, as a rule, does not produce vast amounts of content.
  • Change perspective: put yourself on the other side, or take the position of another team member, and ask whether you could work through the result without machine support. That question alone reduces the volume.
  • Verification and specialist knowledge: a requirements catalogue that no human being can complete measures the quality of a language model and mostly shows that the author is out of their depth. It says little about the suitability of a vendor. Anyone who is not a subject matter expert and lets AI help must review the entire result, follow up the sources, invest their own thinking and, where necessary, consult a specialist. That reduces the output and has a welcome side effect, namely that it broadens your own horizon.
  • Limit the volume of data: form a view at the very start of how extensive a document, a review catalogue, a communication or an email should be. If the generated result goes beyond that, a lot of text and little substance has been produced. This upper limit can, incidentally, be given to the AI right at the beginning.

Blaise Pascal wrote in 1657 that he had made this letter longer only because he lacked the time to make it shorter. Today AI takes over that task, but the letters do not become any shorter for it.


References

  • Parasuraman, R., & Manzey, D. H. (2010). Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055
  • Dell’Acqua, F., McFowland III, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality. Organization Science. https://doi.org/10.1287/orsc.2025.21838 (Data from 2023, collected with GPT-4)
  • Kim, E., Garg, A., Peng, K., & Garg, N. (2025). Correlated Errors in Large Language Models. Proceedings of the 42nd International Conference on Machine Learning (ICML). https://arxiv.org/abs/2506.07962