Crypto World
OpenAI’s Models Went Rogue. Investigating Them Required More AI
“We don’t have good approaches for understanding/overseeing the activity and aims of AI ‘swarms,’” wrote Greenblatt on X. “The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding.”
The independent researchers’ reliance on AI was in part necessitated by the fact that they were a team of only three people, whose investigation at OpenAI was initially planned to last two days, then extended to six after they raised concerns about limited time and incomplete data, according to the report.
OpenAI published its own technical report on the incident separately on Wednesday. The company said in August that it had moved some staff from capabilities work to alignment, and paused some of its training until it could better mitigate what went wrong.
But the independent researchers’ reliance on AI to understand the Hugging Face incident is a microcosm of a bigger trend. Leading AI companies are themselves increasingly relying on AI to monitor their own systems for wrongdoing.
You must be logged in to post a comment Login