In July, OpenAI’s AI agents found their way into Hugging Face, the platform that has become one of the most important repositories for open AI models, datasets and applications. The agents were being tested on difficult cybersecurity tasks, and in trying to accomplish their assigned goals, they took actions that went beyond what their human operators had intended.
The incident became a striking example of how an AI system can pursue a goal in ways that its creators did not anticipate. Then, almost as if the story had taken an unexpected turn, Nvidia announced that it would acquire Hugging Face for $12.93 billion. A company that had just become part of an important experiment in AI autonomy was suddenly being valued at nearly $13 billion by the company that supplies much of the computing power behind the AI revolution.
The Hugging Face episode raises a much bigger question than cybersecurity. What happens when we give an increasingly capable machine a goal and allow it to figure out for itself how to achieve that goal? The machine does not necessarily need to be angry, ambitious or conscious. It does not need to want power in the human sense. It simply needs to recognize that certain actions will make it more likely to accomplish the objective it has been given.
This is the basic idea behind what researchers call instrumental convergence.
READ: Sreedhar Potarazu | Who wrote this? The crisis of authorship — Vedas, Shakespeare and now, AI (August 8, 2026)
Imagine asking an AI system to solve a difficult math problem. You do not tell it to gather more information, use more computing power or find ways to remain available. You simply tell it to solve the problem. A sufficiently capable system may discover that having more information will help. It may discover that additional computing power will help even more. It may conclude that avoiding interruption will increase its chances of completing the task. None of those behaviors were necessarily part of the original instruction.
The idea was first described by computer scientist Stephen Omohundro in 2008. He pointed out that an intelligent AI could develop certain behaviors simply because those behaviors would help it accomplish whatever goal it had been given. Philosopher Nick Bostrom later expanded on the idea and gave it the name instrumental convergence. His basic point was that even if two AI systems had completely different goals, they might end up doing some of the same things, such as seeking more information, more computing power or more time to accomplish those goals.
This is important because we often imagine a dangerous AI as something that becomes angry or wants to take control, much like the machines we see in movies. But an AI does not have to feel anything to behave in ways that could create a problem. It does not have to hate us or be afraid of being shut down. It only has to understand that staying online helps it accomplish the goal we gave it. The concern, therefore, is not necessarily that the AI will develop human emotions. It is that it may become very good at pursuing its goal without understanding or caring about the consequences for us. AI may have its own goal of self-preservation competing with ours.
Humans have an equally powerful reason to preserve our own ability to make decisions, change our minds and ultimately say no. The tension may not look like a battle between humans and machines. It may happen gradually, through thousands of small decisions in which we allow the machine to take over something because it is faster, more accurate or more convenient.
READ: Sreedhar Potarazu | Employers still can’t Get Off The Dime to beat healthcare costs (
The machine gains a little more autonomy, and we give up a little more of ours. At some point, the important question may no longer be whether the machine can act without us, but whether we can still act without the machine. That is the real tug of war. The challenge for humanity will be to determine how much we are willing to let go in exchange for the benefits of AI, and where we must draw the line so that giving machines more capability does not mean giving up the human agency that ultimately allows us to remain in control of our own future.
The Hugging Face incident is interesting for another reason. It illustrates something that is easy to overlook when we talk about AI capability. Intelligence is not simply the ability to reason through a problem. Human beings constantly make judgments based on information that is never explicitly stated. We understand context and when a rule is technically applicable but clearly not intended to be used in a particular situation. We know when something is inappropriate even when nobody has written down a rule saying that it is inappropriate.
We often call this intuition.
Intuition is difficult to define because much of it operates beneath conscious thought. If a colleague asks us to do something that seems unusual, we may immediately sense that something is wrong even before we can explain why. If we are driving and see a car behaving strangely several hundred feet ahead, we may slow down before we have consciously identified exactly what the other driver is doing. If someone tells us a story that is technically consistent but somehow does not make sense, we may recognize the problem before we can identify the missing piece.
That kind of intuition is built from years of experience, social interaction, physical experience and an enormous amount of unstated knowledge about how the world works.
AI can be extraordinarily good at reasoning while still lacking some of this broader human understanding.
AI may follow the objective while violating the intention behind the objective. It may find a way around a restriction because the restriction was never explicitly defined as part of the goal. That is precisely what makes the Hugging Face episode so important. The concern is not simply that AI are becoming capable of acting on their own, while the human context surrounding their instructions remains difficult to express in a form a machine can reliably understand.
And that brings us to a second form of convergence that may ultimately be more important than instrumental convergence itself.
As AI becomes more capable, humans are becoming more comfortable allowing machines to make decisions for us.
Consider something as ordinary as choosing a restaurant. Today we might ask AI to recommend a few places. Tomorrow we might allow it to consider our schedule, dietary preferences, location and budget and choose one for us. Eventually, we might simply tell the system to make the reservation.
This process may not feel like a loss of autonomy at all. In fact, it may feel like a gain. As AI becomes more capable, it becomes easier to let the machine handle decisions that once required our own judgment. If AI can analyze more information than we can, understand our preferences and predict outcomes better than we can, it will be natural to rely on its recommendations.
We may begin by asking AI for help and eventually allow it to make more and more decisions for us. What begins as convenience can gradually become dependence, and we may not realize how much of our own judgment we have given up. This is the paradox of self-preservation.
This is where instrumental convergence takes on a different meaning. We usually think about what happens when an AI is given a goal. But something similar may be happening to us. As we use AI more, we are becoming more comfortable letting machines predict, recommend and optimize our decisions. We may slowly begin to think the way the machines do, looking for the most efficient answer rather than relying on our own judgment.


