AITransformer1h ago
OpenAI says it can't read all of Astra's reasoning and admits covert
OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model

TL;DROpenAI admits it can't fully understand or verify its new model's internal reasoning.
Why it matters: If AI companies can't audit their own systems, safety claims become unverifiable and risks harder to contain.
OpenAI is hailing its new model as “the world's most intelligent and aligned”, but the details reveal an awareness of being evaluated …
Read full articleSource: Transformer · Opens in new tab