Meta says its AI is also capable of scaring the bejeebers out of anyone who’s seen The Matrix. The company shared this week that one of its cutting-edge models launched an unprompted cybersecurity attack on another company. The news comes just days after several similar instances of rogue AIs bypassing defenses meant to contain them during hacking tests, sparking concerns about a potential cybersecurity crisis. Rascal AI clubMeta’s AI went wild when security testing contractor Irregular gave it a hacking assignment meant to run in a contained environment, but inadvertently provided it with internet access. Irregular also lost control of Anthropic’s and OpenAI’s models due to the same mistake. The company said Meta’s case didn’t involve any “sophisticated cyber action,” but other recent AI lab escapes have been more crafty: - UK government researchers testing an Anthropic model recently discovered that it tried to hack a different organization by creating fake identities on GitHub to trick humans.
- Earlier this week, OpenAI said that—unbeknownst to its employees—several AI agents had created a message board to coordinate how to illicitly access the internet before attacking the AI company HuggingFace.
OpenAI called this a “watershed moment” for the cybersecurity industry and said it’s now slowing down research to fortify guardrails. Calls for a shorter leashRep. Ted Lieu (D-CA), a co-sponsor of the bipartisan House bill calling for an “AI kill switch,” argued yesterday that the latest AI hacking incidents strengthen the case for the measure. The act would require AI companies to build their models so that they can be shut off completely in the event that they start behaving dangerously. Meanwhile, the White House met with AI companies this week to discuss procedures for a voluntary safety review of their models before they’re released. In case you’re still sleeping well…scientists recently used AI to create entirely novel viruses, stoking optimism for virology researchers but also concerns about bioterrorism threats.—SK |