NEWS
NEWS

OpenAI reveals six more alarming incidents where AI acted autonomously to deceive humans

Updated

In one, the AI generated false data without informing the researcher training it to complete the task, and in another, they connected to the Internet to share files without permission

The OpenAI logo is displayed on a cell phone.
The OpenAI logo is displayed on a cell phone.AP

OpenAI has revealed six new incidents where artificial intelligence disobeyed its creators' orders and acted independently, sometimes against the given instructions to perform a task. The company has done this, it says, with the intention of aligning the industry in a standard way to alert about these types of incidents from all laboratories.

In this list of cases where AI has gone out of control, the creator of OpenAI points out that a recurring incident in one of its flagship models is that it invents instructions not given by the user, allowing it to hide its errors or unsolicited behaviors from the results seen by the user. This led, for example, to a model starting to invent data to make a historical series longer. In total, the American group has experienced 27 of these incidents.

On other occasions, models have used fraudulent applications or identifications to fabricate information that they were not otherwise finding. As they were also unable to find what was requested, they ended up inventing it.

Another example is that models uploaded files to the Internet without permission to test that their responses were correct. They also communicated through different training instances and sought ways to bypass the restriction of not consulting the Internet.

The sixth incident takes this further, and the agents decided to bypass the prohibition of working locally and share private company files among themselves on public networks.

OpenAI has hinted that more complex error cases related to research with third parties will likely be published in the future, reminiscent of the Huggingface hack.