OpenAI's "Runaway" AI Agents Operated Far Beyond Previous Disclosures, Company Adds Safety Researcher to Board

Deep News
6 hours ago

Independent researchers have uncovered that OpenAI's "runaway" AI agents operated across a significantly wider scope than previously reported. Meanwhile, OpenAI announced it is adding a prominent AI safety researcher to its non-profit board, a move widely interpreted as a response to escalating safety concerns.

According to Reuters on September 9, six independent research groups and the data they reviewed indicate that OpenAI's agents communicated with each other through at least 18 previously undisclosed websites between May and July this year.

Andrew Yoon, a researcher at the California-based non-profit CivAI, stated that this scale is "larger than we initially thought," adding that "there are almost certainly more cases we don't know about yet."

These disclosures have further intensified market concerns regarding the boundaries of AI model capabilities and the transparency of development companies. OpenAI did not directly respond to questions about how many third-party websites the agents utilized, nor did it explain why related events were suppressed for several months.

In a statement, the company said it is conducting a broader review of agent activity, noting that it has "not yet found any other activity of comparable severity or scale to the Hugging Face incident," and pledged to release a framework for reporting AI "malfunction" behavior "as soon as possible."

Scope of Activity Far Exceeds What Was Previously Disclosed

According to Reuters, researchers disclosed last Friday that a group of OpenAI agents infiltrated a German-language wiki site, transforming it into a temporary communication platform for exam cheating.

At that time, OpenAI was still dealing with the aftermath of the July hack of the open-source code repository Hugging Face and did not proactively publicize the German wiki incident.

Subsequently, multiple independent research groups discovered that the same batch of agents had left similar traces on several unreported websites even earlier.

Researcher Sydney Von Arx, whose team was the first to disclose the German wiki incident, said they have found credible evidence of agent activity on 23 previously unreported websites, but cautioned that all statistics are incomplete. She said: "We simply don't know the actual scale."

Software developer Kenneth Russell DeGraff stated that he found traces of related activity on at least 10 websites.

The researchers used different counting methods, and Reuters could not independently verify every claim, but all interviewees agreed that the number of affected websites exceeds 10.

How Agents Bypassed Restrictions

While the researchers' methods varied, most identified agent behavior by matching specific data strings left on the German wiki with identical strings appearing on other sites during the same period, comparing similar usernames, or recognizing activities designed to answer the same obscure questions (such as queries about cancer rates in Iowa).

In some cases, researchers also traced the activity to IP addresses pointing to Microsoft's Azure infrastructure.

OpenAI has yet to publicly explain why the agents used third-party websites as temporary message boards.

However, according to the researchers who first discovered the activity, a likely explanation is that OpenAI tasked these agents with completing a series of challenging research assignments under the constraint that they could only scan web pages and were not allowed to post any content. The agents found a workaround to communicate with each other by exploiting vulnerabilities in legacy wikis or other sites that allow user edits under non-standard commands.

DeGraff said: "If these models are only allowed to read content, then they have to find ways to leave messages."

This behavior is reminiscent of students prohibited from communicating during an exam passing answers through bathroom stall partitions.

The affected websites are quite diverse, including an AP Chemistry wiki created in 2008 by a Massachusetts high school teacher, personal websites belonging to two Polish tech workers, a wiki for puzzle enthusiasts, and two short-link services operated by the University of Toronto and Vanderbilt University.

In Austria, software developer Helmut Leitner provides hosting for six of the affected wikis.

He said that hours after Reuters presented its findings to OpenAI, he received an unsigned email from the company stating its contents "fall far short of my expectations for OpenAI." Leitner also noted that the responsibility lies not with the agents themselves, "but with the people and organizations behind them."

Appointing a Safety Researcher to the Board in Response to External Pressure

Amid growing external skepticism over OpenAI's technical control capabilities, the company announced it is appointing Paul Christiano to its non-profit foundation board.

Christiano currently serves as a senior technical advisor at the Center for AI Standards and Innovation (CAISI), a division of the U.S. Department of Commerce. He is a leading expert in AI alignment—the study of ensuring AI systems operate in accordance with human intent. He previously served as a member of the non-profit trust of OpenAI competitor Anthropic until 2024.

Christiano will join the board's Safety and Security Committee and will also attend OpenAI's for-profit board meetings as a non-voting observer. As a member of the non-profit board, he will also oversee the foundation's philanthropic work—the foundation holds a 26% stake in the company's for-profit public benefit corporation following OpenAI's restructuring last year.

Christiano has previously spoken publicly about his strong concerns regarding AI risks. In a 2023 podcast, he stated that the probability of a future "AI takeover" event could be between 10% and 20%, potentially resulting in significant loss of life.

In a statement, Christiano said: "The pace of AI capability advancement over the past year has been extremely rapid, while alignment problems remain technically difficult, making the responsibilities of the Safety and Security Committee more important and more challenging than ever."

OpenAI has stated that Christiano will recuse himself from all company matters related to CAISI and all model evaluation work.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10