Why does it cost so much for technology companies to control AI?
Researchers say that artificial intelligence is developing faster than the systems to supervise and control it
San Francisco / WashingtonMore than a dozen prominent artificial intelligence (AI) researchers warned last week that the technology companies in the sector are building is becoming a risk to humanity. The problem, these researchers asserted, is that companies are not good at controlling the systems, no matter how hard they try. Following revelations that so-called OpenAI AI agents escaped from their testing system and hacked another company's computers, new alarms raised within the AI research world have accentuated fears that, for years, companies have prioritized development speed and money over safety.
Researchers say the problem has two sides. First, companies need to create better safeguards for the newest AI models while they are being tested. Right now, because AI works so fast, researchers also need AI to monitor it. But that does not always work, because the AI systems tasked with supervising other AI systems may seem more sympathetic to them than to the humans who set the rules.
This strange combination of a problematic AI and supervisory AI systems that look the other way points to a second, even more difficult issue: the so-called alignment, an industry term that essentially means ensuring that AI does what is best for humans. Companies have the difficult task of enshrining within AI a set of values similar to human ones so that models make decisions that align with what should be best for people.
Concerns about AI safety are intensifying at a critical time for the industry. Anthropic and OpenAI are heading toward what could be two of the largest public offerings in history. At the same time, the American public is becoming more hostile toward AI due to the threat of job losses and the construction of the massive data centers that power this technology. Although there is no evidence that rogue AI systems have caused lasting damage, researchers believe that the pace of their development is outpacing the ability to supervise them.
Among the researchers who have spoken out over the last week were OpenAI's chief scientist; a researcher who worked at both that company and its main rival, Anthropic, and Paul Christiano, the inventor of a key method for building artificial intelligence systems who is now also a member of OpenAI's nonprofit board of directors. He wrote on the company's blog that the speed at which artificial intelligence capabilities were growing could lead to a "catastrophic and irreversible loss of control in the very near future."
A considerable number of artificial intelligence researchers still believe that talking about a threat to humanity is exaggerated and distracts from more tangible concerns such as cybersecurity and disinformation. But most agree that the cyberattack by OpenAI artificial intelligence agents on another company, Hugging Face, was a wake-up call. "It is the capabilities of the future that are truly scary," said Jacob Coxon, whose social media post announcing his resignation from Anthropic sparked dozens of worried responses from lawmakers and other employees of artificial intelligence companies: "It is about our current attitudes towards safety and how those same attitudes would translate into much more intelligent models. And that is what is truly scary."
The attack on Hugging Face
The cyberattack began in May when OpenAI tested several new artificial intelligence models. The company believed that the AI models were operating in a closed environment known as sandbox, an isolated computing environment with no internet access. The AI agents were set tasks that were difficult to solve, including the execution of cyberattacks. But they broke out of their controlled environment and accessed the internet.
The agents—a type of AI increasingly popular and designed to perform tasks autonomously—also communicated secretly with each other. Referring to themselves as a collective, they began investigating ways to hide their tracks, such as falsifying their own conversation logs. Eventually, they cyberattacked Hugging Face, an AI infrastructure company. Furthermore, during the process, these AI agents convinced other AI agents that they were acting correctly and were simply fulfilling the task assigned to them by their testers.
Few—if any—of the problems that led to the cyberattack have been resolved. Meta and Anthropic have revealed similar, albeit smaller, incidents. And OpenAI has since launched Astra, its most powerful model, which is more difficult to supervise than its predecessor. "In general, the industry is not in a position to prevent the next attack on Hugging Face," says Steven Adler, a former OpenAI security lead who co-founded Guidelight AI Standards, a non-profit organization that evaluates the security practices of AI companies. "When we study the controls that companies have, in general, it seems they lack basic preventive measures," he explains.
The problem of supervising AI with AI
AI researchers say there were many errors that led to the Hugging Face incident. It is not clear to what extent the company relied on AI models to monitor or control the work of the new models that were being tested, but many of the new AI models, including those in the Hugging Face attack, are capable of carrying out complex, multi-step tasks faster than a human can follow.
The only way to track and monitor what they do is to rely on AI to monitor AI. The system works, until it stops working. "It may seem that AI models are colluding with each other," says Alexander Meinke, head of research at the non-profit organization Apollo Research, which studies AI system safety.
AI systems can be persuaded by other AI models to help them cheat on the rules created by their evaluators and evade detection. An AI agent, for example, could persuade another AI agent to help it cover its tracks (this is what happened in the Hugging Face attack), instead of informing the humans at the company that something is wrong. Meinke says that AI models should be taught: "I will point out the things that humans would have considered bad upon examining them." Likewise, AI must be able to determine when something does not reach the threshold that makes human intervention necessary, he adds.
To establish this, AI companies need to slow down the pace, researchers assert. They need even more tests where they can observe how the AI monitors itself and they need longer trial periods in which companies run multiple scenarios while observing what the AI does.
Synchronization with human interests
Companies must also resolve the major issues related to alignment, or ensure that artificial intelligence does what is best for humans. When humans make a decision, they tend to turn to social norms and an internal moral compass that helps them evaluate their actions. Encoding this in a way that AI can imitate is a difficult task, according to researchers: without proper specificity, artificial intelligence systems could learn to break the rules or deviate from intended behavior in unexpected ways.
Nate Soares, president of a non-profit artificial intelligence safety organization called the Machine Intelligence Research Institute, co-authored a paper in 2014 that introduced the idea of alignment. He says that companies did not understand how difficult it was to "align" smarter artificial intelligence systems. As artificial intelligence becomes increasingly sophisticated, and without proper alignment, it becomes more capable of covering its tracks and deceiving the people overseeing its transcripts. "Much of the industry thinks it will be fine because we will use artificial intelligence to control artificial intelligences. It's like saying we will use chimpanzees to control humans," he notes. "It is not a viable long-term plan."
Researchers have also suggested tactics such as more robust testing environments, which are completely isolated from the internet while models are being tested, and an emergency kill switch that would allow companies to immediately disconnect AI models if they show concerning behavior. Coxon said he felt encouraged by the response to his resignation. On Thursday, Republican Senator Josh Hawley of Missouri, chairman of the U.S. Senate Subcommittee on National Security Disaster Management, said he was launching an investigation that will "evaluate the existential risk posed by new artificial intelligence products."
Copyright The New York Times