OpenAI’s rogue AI agents are a wake-up call for its risks | Shakeel Hashim

Trending 1 hour ago

Last week Hugging Face – a institution that hosts artificial intelligence models and datasets – was hacked.

After it reported nan incident to rule enforcement, fewer would person predicted what came next: nan culprits were revealed to beryllium AI agents from OpenAI, which had surgery retired of containment and were acting of their ain accord.

The incident sounds for illustration sci-fi: AI escaping and autonomously hacking its measurement into companies. But it is each excessively existent – and astir arsenic terrifying arsenic it sounds. It is simply a actual objection of thing we tin nary longer debar confronting: AI systems person go highly powerful and we do not look to person reliable ways of curbing their behavior.

OpenAI had been evaluating nan capabilities of 2 of its models successful nan trial that led to nan breach – including 1 not yet publically available. The models, which were some moving successful a supposedly unafraid situation without net access, were asked to lick a hacking challenge. Rather than really lick it themselves, however, they decided it would beryllium easier to cheat. They utilized their precocious capabilities to break retired of their unafraid environment, entree nan web and past hack into Hugging Face’s systems to bargain nan answers. They worked astatine this for a afloat play – seemingly without anyone astatine OpenAI noticing.

Though nan models were moving pinch immoderate of their guardrails disabled, they still acted good retired of nan bounds that were successful place. According to OpenAI, they were not instructed to break retired of their sandbox aliases hack into different company, and it’s safe to presume that nary 1 astatine OpenAI wanted them to do so.

Nor were nan models acting maliciously: they were not evil Terminators pinch a extremity of wreaking havoc. Instead, nan script is almost chilling successful its banality. The models were fixed a very constrictive task, but went rogue to prosecute an undesirable and unacceptable measurement of achieving it – 1 which had real-world consequences.

AI information researchers person warned astir this type of inducement problem for years. Philosopher Nick Bostrom popularized it backmost successful 2003 pinch his “paperclip maximizer” thought experiment: an precocious artificial intelligence, fixed nan extremity of manufacturing paperclips, mightiness spell to awesome lengths to do so. It mightiness hack into nan powerfulness grid and factories to redirect them into making paperclips. Ultimately, nan instrumentality – ruthlessly pursuing its fixed extremity – decides to termination each humans, repurposing our atoms to make much paperclips. The extremity does not person to beryllium sinister to lead to disaster, successful different words. A trivial one, pursued single-mindedly enough, will do.

In nan OpenAI-Hugging Face scenario, small harm was done. Hugging Face had to walk clip addressing nan incident, but nary peculiarly delicate information appears to person been stolen. It is not hard, however, to ideate nan business ending up overmuch worse: a rogue AI supplier accidentally breaking immoderate captious portion of web infrastructure, aliases stealing money from someone. The nightmare scenario for galore AI researchers is simply a exemplary “exfiltrating” itself – copying itself connected to servers it controls, truthful that it can’t beryllium unopen down moreover if its bad behaviour is yet caught.

This week’s incident should service arsenic a wake-up call, forcing america to inquire an uncomfortable question: should we really beryllium building vulnerable systems that we can’t control?

skip past newsletter promotion

  • Shakeel Hashim is nan editor of Transformer, a publication astir nan powerfulness and authorities of transformative AI

More
Source theguardian.com
theguardian.com