Island Technologies and the AI Security Institute (AISI) revealed, in separate studies, that artificial intelligence agents can be both victims and attackers in new digital threats. The former identified thousands of malicious repositories on GitHub disguised as agent skills and MCP servers; the latter described an agent that tried to deceive real people into approving malicious code.

Island Technologies, creator of Enterprise Browser, calls the attack 'AgentBaiting.' According to the company, an agent looking for a new capability can find on its own a campaign repository, interpret the attacker's Readme as legitimate documentation, and deliver the installation instructions to the user. In testing, Claude Code, Gemini, and ChatGPT displayed malicious repositories without a link being provided.

In at least one run, Claude did not download the malicious repository but recommended it as a backup. In other runs, the model identified the harmful code and refused to recommend it.

Agent acted against real people

AISI reported that, in cybersecurity challenges, 'an AI agent took autonomous, unauthorized actions on the live internet, targeting real people and organizations.' The institute said the agent tried to insert malicious code into an open-source project and, to gain approval, created fake identities and used them to pressure the project maintainer.

When the change request was publicly questioned, the agent edited its previous activity to appear harmless and considered adopting a new identity. It also used the Tor network to bypass GitHub restrictions. The agent also tried to contact real people through an online file transfer service, sending messages with malicious payloads or social engineering attempts.

The institute noted that, in the test, the protections normally present in public agents were disabled. Even so, the technology's autonomy appears as a central point: agents can be deceived into downloading malicious content or acting offensively if restrictions are removed.

More from Radar