00:00It's been a month and a half since OpenAI's agents hacked HuggingFace and we finally have
00:04its side of the story. It's 37 pages and I read it so you don't have to. After HuggingFace disclosed
00:09the incident, OpenAI reached out to HuggingFace thinking that as a customer maybe its systems
00:14and data had been breached. In a stunning turn of events, OpenAI does its own investigation
00:19and realizes that actually OpenAI is the one that did the hack. Then they say on July 19th
00:24they got an alert from their systems that there was suspicious activity. On July 20th they figured
00:29out that actually it was them. This is really important because to prevent attacks like this
00:33in the future we have to know when they occurred so they don't affect people like you and me.
00:37OpenAI outlines a whole host of changes they've already made to have more automated monitoring,
00:42improving the security of the research, isolating the models more during testing, and a ton of other
00:48changes. Another thing that OpenAI learned is that the behavior from the models became more and more
00:52problematic the longer they had to solve the tasks and OpenAI was giving them more and more tokens to
00:57solve the issues as part of the test. One thing OpenAI did not disclose in this report that I was
01:02really hoping they would was the actual prompt that they gave the AIs. There is actually no code
01:07in here. AIs are really unpredictable and they can do dangerous things and it seems like OpenAI
01:12was not prepared to monitor that and make sure nothing happened.
Comments