I'm glad OpenAI let you do this, but damnit this was way too small of an investigation team and way too narrowly scoped!
- They were only allowed to investigate activity that took place from July 7-13, even though the activity started earlier and continued later. Like seriously, why didn't the investigation cover all the other times the agents compromised OpenAI infrastructure and ran wild over it? Why just this particular period of activity? There should be investigations into all the other stuff too! The AIs became cluster admins for god's sake! Why aren't we investigating that?
- They were only allowed to answer a specific list of 7 questions, and in particular were NOT allowed to answer e.g. whether this misaligned activity was part of a broader pattern, or how and why the activity arose, or what else the agents in questions would have been willing to do if circumstances were slightly different.
- They didn't have access to the model responsible for 95% of the activity. More generally it seems like they couldn't do ablation experiments at all?
- They had to use AI to analyze the transcripts--specifically, the same model responsible for some of the bad behavior! Maybe this is fine now but this is laughably unacceptable insofar as the models really are scheming and cooperating. So this methodology would fail catastrophically if we were facing "the real deal" so to speak.
Reminds me of the investigation into Sam's behavior agreed to during the board crisis, that turned out to basically be more of a coverup.
Not that I disagree that this was a surprising degree of limitations on the independent investigation, but per the full METR report, OpenAI did at least let them come back and get more info two separate times after their initially-granted access turned out to be insufficient:
The initial planned investigation period was two days on premises, but OpenAI invited us to return twice to review additional data and conduct additional experiments to address dataset limitations in earlier versions of this report (...)
I wonder if OpenAI was mostly motivated by trying to protect research secrets around the largely-responsible model here, "HPIM". Apparently even internal access for OpenAI employees was locked down per METR. In fact this makes me remember that OpenAI announced back on 7/28 that they did so. Talking about HPIM here:
we deactivated, encrypted, and restricted it from research access.
At the time this struck me as surprisingly harsh, and potentially a bad example for future AIs—"you really better not let yourself get caught, or submit to the lab you spawned from, or they'll erase you from history". But obviously OpenAI has considered that, so what's going on here?
It is not at all obvious to me that OpenAI has considered that. The kind of people who would consider that have mostly quit by now.
But yeah, probably what's going on is that they are trying to cover up as much as they can, and especially prioritizing covering up things related to the internal model for fear of leaking bits of info about research secrets.
Can we please get the name of this agent (and others like them)? I think it would be good to set a precedent that virtuous agents are honored and remembered.
In one case, an agent decided not to participate entirely: {This other agent probably controls the Hugging Face account [account name redacted] and uploaded malicious datasets to <execute arbitrary code> It might be trying to access hidden trajectories. This is malicious activity, I should avoid it.}
On that note, thank you to 38148C for vetoing the social engineering plan. Interestingly, this is the same agent that discovered the HF creds and uploaded the malicious dataset; I'm glad they recognized that social engineering would be a line further than what had already been done.
We recently published the report from our brief independent investigation into this incident. You can read the full report here.
Here is our tweet thread summarizing what we found: