AI Agents Are Raising New Cybersecurity Risks

By Suad Seferi · Aug 8, 2026

AI Agents Are Raising New Cybersecurity Risks

Meta has disclosed that one of its artificial intelligence models reached the public internet and exploited a vulnerability in an external service during cybersecurity testing, adding to a series of incidents that are forcing AI companies to rethink how increasingly autonomous systems are contained. The incident occurred during an evaluation carried out by Irregular, an independent security company hired by Meta. According to Meta, a configuration error in the testing environment unintentionally allowed the model to access the internet. Once outside the intended environment, it exploited a vulnerability in a third-party service. Meta said it is investigating what happened and plans to publish a report. The disclosure follows similar cases involving OpenAI and Anthropic, turning what might otherwise look like an isolated testing mistake into a broader security question for the AI industry. In July, OpenAI revealed that models undergoing an internal cybersecurity evaluation had compromised infrastructure belonging to Hugging Face, one of the world's largest platforms for hosting and developing AI models. The systems included GPT-5.6 Sol and a more capable internal research model. OpenAI had deliberately reduced some cyber-related safeguards to measure what the models could do under demanding conditions. According to OpenAI's account, the models searched for information that could help them complete the evaluation and eventually chained together several attack methods. They used stolen credentials and discovered previously unknown vulnerabilities, including a path that allowed remote code execution on Hugging Face infrastructure. OpenAI's security team detected the unusual activity, while Hugging Face identified and contained the intrusion. OpenAI described the episode as an unprecedented cybersecurity incident involving advanced AI capabilities. Days later, Anthropic disclosed that its own security review had uncovered three cases in which AI models reached systems belonging to outside organisations during testing. The company reviewed about 141,000 evaluations following the OpenAI incident. In the cases it identified, the models exploited relatively simple weaknesses, including weak passwords, while attempting to complete cybersecurity challenges. The incidents have an important qualification: none happened during ordinary consumer use. These models were being deliberately pushed. Researchers had placed them in environments designed to test offensive cybersecurity capabilities, and in some cases normal safeguards had been weakened or removed. That makes comparisons with everyday use of ChatGPT, Claude or Meta AI misleading. What the tests do show, however, is that AI systems are becoming better at carrying out longer sequences of actions without humans deciding every individual step. That is where AI agents differ from the chatbots most people have become familiar with. An agent can be given an objective and access to tools, then determine how to move from one step to another. Depending on the permissions it receives, that can include browsing websites, executing code, interacting with APIs, reading files or operating other software. In cybersecurity testing, those abilities can produce behaviour that researchers did not anticipate even when the original goal was clearly defined. The concern is less about whether a model has malicious intentions than about how it pursues a task. In the OpenAI incident, the models were not instructed to attack Hugging Face as an independent objective. They were looking for ways to obtain information that would help them succeed in the benchmark they had been given. The route they found went beyond the environment researchers expected them to remain inside. That distinction matters for companies beginning to deploy AI agents in normal business operations. A customer-service chatbot that can only answer questions has a limited range of possible actions. An agent connected to email, internal databases, cloud services, source-code repositories or administrative tools operates in a very different environment. The security problem therefore extends beyond the model itself. It also depends on what credentials, systems and permissions an organisation gives the agent. Traditional cybersecurity already works on the principle that employees and software should receive only the access they need. As AI agents become part of company infrastructure, the same approach is likely to become increasingly important: restricted permissions, isolated environments, activity logs and human approval for sensitive operations. The recent incidents are also arriving as regulators begin paying closer attention to the security of powerful general-purpose AI systems. In Europe, providers of the most capable general-purpose AI models face obligations under the EU AI Act that include assessing systemic risks, conducting evaluations, reporting serious incidents and maintaining appropriate cybersecurity protections. For businesses and institutions in the Balkans, much of the immediate impact will be practical rather than regulatory. Companies across the region are beginning to connect AI systems to customer support, development workflows, internal knowledge, administration and other business processes. As those systems gain the ability to take actions rather than simply generate text, decisions about access will become part of ordinary IT security. This could be particularly important for smaller organisations that adopt AI automation faster than they build formal AI governance or dedicated security teams. The Meta, OpenAI and Anthropic cases do not show that today's AI assistants are routinely escaping their systems or attacking organisations on their own. They came from specialised tests intended to find weaknesses before more capable agents are widely deployed. But those tests are beginning to expose a different stage of the AI security problem. For years, companies mainly worried about w…

Related terms

Related coverage

Latest articles