HUMAN IN THE LOOP. A CONCEPT DESIGNED TO FAIL?

Sep 10
10 min read
There is no shortage of worrying stories about AI at the moment. Agents behave differently from what their developers expected, systems find ways around restrictions and experiments produce situations in which models try to preserve their access or achieve a given objective in ways that nobody intended or would have allowed.
Some of these examples come from deliberately constructed test scenarios, others are closer to real-world use and there is plenty of room to debate how seriously each individual case should be taken. But they raise a legitimate question: how do we remain in control as AI systems become more capable and increasingly able to act autonomously?
Much of this discussion understandably focuses on the companies developing these models. They need to understand their systems, build safeguards and make sure that increasing capabilities do not create risks that eventually become impossible to control.
For the rest of us, however, it would be rather convenient to consider the problem solved there. Most of us will probably never develop a foundation model, but we are the ones putting these models into companies, connecting them to data and applications, giving them access to tools and eventually allowing them to perform work autonomously. We therefore define a significant part of the environment in which these systems operate and cannot rely entirely on others to make their use safe. The relevance of this is already visible in completely ordinary AI use.
Not every wrong answer is a hallucination
When we talk about errors in generative AI, the term "hallucination" comes up quickly. It usually describes content that a model presents convincingly even though it is factually wrong, unsupported or entirely invented. A model might cite a source that does not exist, present an incorrect number as fact or give an answer for which there is no reliable basis.
But not every problematic answer is a hallucination in this narrower sense. Sometimes the individual pieces of information are correct but are connected incorrectly. The model does not invent a new fact out of nowhere. It constructs a relationship between existing information that never actually existed.
I have had this happen several times when I realized that the context a model was drawing on was broader than I expected at that moment. From the AI's perspective, I am whatever it knows about me: "Claas, this approach would work well for your X project at Y." There is only one problem: there is no X project and I once asked a question about Y but have never worked for the company. None of the individual information has to be wrong. The earlier research happened and so did the later discussion. The system simply connected the two and constructed a context that never existed.
Of course, this ability is impressive and one of the reasons AI is so useful. At the same time, it can be quite frightening. The model is doing more than retrieving information. It forms a judgment from the available context about what belongs together, what is probably meant and what action might follow. That judgment can be based on assumptions that were never stated or verified.
As long as the result is an answer on my screen, I can deal with it. Mainly because I can just read it, recognize the false assumption and correct it. Once the same system has access to corporate data and applications and eventually the ability to take actions itself, the idea that I will spot every incorrect assumption becomes considerably less reassuring.
This is where one of the standard answers to the problem usually is the conecept of "Human in the Loop" (HITL) .It may sound reasonable, but I am just not (any longer) sure that it can survive, as a control model, the scale we are trying to achieve with AI.
Being in the loop does not necessarily mean being in control
We can already see the problem in the way AI is used to create content today. With rising expectations and constant pressure to deliver more in less time, AI is an obvious tool. Information can be researched and structured faster and documents and presentations can be produced at a speed that previously required substantial effort.

What concerns me more is how often obviously insufficiently reviewed content finds its way into final deliverables. A slide has a clear headline, three neatly structured arguments underneath and an apparently logical conclusion. Look more closely and sometimes those three arguments simply repeat the headline in different words. The structure and appearance create the impression of a sound argument that is not actually there.
The fact that much of the remaining output can be remarkably good makes the problem harder. It is easy to extend that confidence to the entire result even though the quality of individual parts can vary considerably. At the same time, AI increases the volume of content requiring review much faster than it increases the attention available to review it.
Increasingly, the input itself has already been produced with AI. One colleague creates a document using AI, someone else uses it as the basis for another analysis and another system turns that into a recommendation. An assumption that was never properly validated at the beginning can pass through several stages of work and gradually acquire the appearance of established fact.
In our own work, this makes professional standards particularly important. We should use AI wherever it genuinely takes over work, simplifies it or allows us to produce better results more efficiently. But if our name ultimately appears on a recommendation, presentation or other deliverable, we need to understand the argument, have checked the relevant content and be prepared to stand behind the result. Having a human somewhere in the process does not guarantee any of that.
Human in the Loop has a scaling problem
The issue becomes even more obvious once AI moves from producing content to taking action. It is easy to imagine a human supervising an agent. The agent proposes an action, a human reviews it, approves exceptions and intervenes when something unusual happens. Particularly during an initial learning phase, this may be exactly the right approach.
I found a related point in Bill Gates' recent article about the impact of AI on work and society particularly interesting. He describes the point at which AI can perform work with near-perfect accuracy, giving companies a strong economic incentive to let it operate without continuous human review. At the same time, he discusses the risk that more capable systems may act in ways that are not in our interests.
The problem with HITL sits precisely between these two developments. The better AI becomes at performing work autonomously, the stronger the economic incentive to reduce human review. At the same time, increasing autonomy makes effective control more important. The article goes far beyond this question and is well worth reading: see Bill Gates, “The turbulent AI era is here”
As a permanent operating model, the idea of a human reviewing an agent's work becomes difficult once we start to scale. If an agent is supposed to take over a meaningful amount of human work but another human then needs to fully review every action it takes, a significant part of the expected benefit already disappears. Once we move from one agent to ten, a hundred or eventually thousands of agents working in parallel, individual human supervision becomes unrealistic.
Human attention is also poorly suited to this kind of monitoring. If a system performs correctly a thousand times, the human supervising it is unlikely to review the thousandth transaction with the same attention as the first. High reliability itself can therefore make actual oversight increasingly superficial over time.
A limiting-case perspective
It may help to look at the question not from the perspective of today's systems but as a limiting case. The concept comes from mathematics: we examine how a system behaves as certain variables approach an extreme. The result is not a prediction, but it can reveal whether a concept is fundamentally scalable or whether it eventually reaches a structural limit.
Assume, then, that AI is not merely somewhat more capable but a thousand times faster and that tens of thousands of agents are working simultaneously within an organization.
Human control in the sense of comprehensive review would no longer be conceivable in such a scenario. Not because humans are incapable of reviewing individual decisions, but because the number and speed of actions would exceed any human capacity to review them. A human could no longer see, understand and approve every action. Even if the system produced a rationale, source list and risk assessment for every decision, the amount of information would eventually exceed the attention humans could devote to it.
This becomes even more apparent with machine-to-machine communication. If one agent researches, another evaluates and a third acts on the result, there is little reason to create a beautifully formatted 40-page PowerPoint presentation between them. A growing share of information may be exchanged directly between machines without ever being translated into a format intended for human consumption. A false assumption or invented relationship can therefore move from one system to another, become part of the context for another decision and ultimately trigger an action without a human ever seeing the intermediate steps.
At scale, human oversight approaches zero.
The limiting-case perspective does not tell us whether such a system can ultimately be controlled. It initially tells us only who can no longer control it: a human trying to review its individual actions .As the number and speed of actions become very large while human attention remains limited, the proportion of actions actually reviewed by humans inevitably approaches zero. This is not a question of discipline, but a structural consequence of machines and humans scaling at fundamentally different rates.
If humans cannot scale, part of the control will have to
If individual human oversight cannot scale accordingly, one consequence seems difficult to avoid: a significant part of operational control will itself have to be performed by machines.
I do not think it will be sufficient simply to ask a second agent to check the work of the first. We will probably need highly specialized control systems whose primary purpose is to monitor other systems. They would need to operate with the explicit assumption that the system being monitored may make mistakes, construct false relationships, cross defined boundaries or find a way of completing an assigned objective that technically satisfies the goal but is nevertheless not in our interests.
For such controls to work, they will need a degree of independence from the systems they supervise. If both rely on the same model, information, context and assumptions, there is at least a risk that they will make the same mistakes. Whether this requires different models, separate contexts, deterministic controls or more extensive architectural separation is one of the questions for which we do not yet have a definitive answer.
Crucially, a control system cannot merely observe. If it identifies a significant deviation, it must also have the authority and technical ability to block an action, restrict permissions, isolate an agent or, in an extreme case, stop it. That makes the control system itself a critical component of the architecture. A system capable of stopping hundreds or thousands of other agents necessarily holds considerable power. Who or what, then, controls that system?
We may ultimately end up with a layered control architecture: operational agents working within clearly defined boundaries, specialized systems monitoring their behavior and intervening when necessary and humans defining rules, limits and escalation mechanisms while stepping in where human judgment is genuinely required. There is something paradoxical about delegating part of the control over AI back to AI. At sufficient scale, however, I see few alternatives.
Control has never meant checking everything
The underlying principle would not be entirely new. Companies already manage enormous volumes of activity without a human personally reviewing every transaction. A CFO does not check every booking and a CISO does not inspect every login. Instead, organizations use permissions, limits, segregation of duties, automated controls, monitoring, sampling, audits and defined escalation mechanisms.
Digital labor will probably require similar principles, adapted to systems that can interpret situations and act with much greater autonomy. We will need to determine what information an agent can access, which systems it can modify, which actions it may perform autonomously and where financial, operational or other limits apply. We will need to consider whether actions can be reversed, how unusual behavior is detected and when a human or another control system must intervene.
The human role does not disappear in this model, but it changes. Instead of reviewing every individual action, humans increasingly need to design the rules, boundaries and control mechanisms and intervene at the points where human judgment is genuinely required.
We should not pretend that we have already solved the problem
There is, however, a conclusion that would be far too easy: Human in the Loop does not scale, so we simply replace human oversight with automated control systems. That initially moves the problem to another level rather than solving it.
We do not yet know exactly what such a control architecture needs to look like, how independent supervisory systems really need to be or how we can ensure that they continue to work under conditions that were not anticipated when they were designed. At the same time, the potential consequences of autonomous systems may eventually become too significant to answer these questions only once those systems are already operating at scale. Understanding the limits of our current control mechanisms is therefore at least as important as developing increasingly capable systems. We need to understand where Human in the Loop works and where it does not, what new risks emerge from machine-to-machine communication and automated oversight and under what conditions we can responsibly give a system greater autonomy.
For me, this does not mean stopping the development of autonomous AI. It means that we should not scale autonomy faster than our ability to understand its consequences and build effective boundaries around it. With a system that produces information, mistakes can often be corrected afterwards. With systems making decisions and taking actions autonomously at scale, we cannot rely on discovering the critical problems only after they have occurred.
We cannot simply delegate the problem to the developers of the foundation models.



Comments