Back to wire
Development·Article·Confirmed

OpenAI's security code red mobilised 250+ people; Brockman says 25% of production engineers were reassigned

OpenAI says an internal security code red mobilised more than 250 people across more than 100 service areas and grew into its continuous Defense Factory workflow. President Greg Brockman separately says the company paused projects for 25% of its production engineers so they could harden systems with frontier models; his claim that Astra eventually saturated the P0 issues it could find does not establish that OpenAI is free of critical vulnerabilities.

Published 14 Sept 2026, 18:20 · Updated 21 Sept 2026, 13:11

A security sprint diverted a quarter of production engineering, Brockman says

OpenAI says it called an internal security code red and mobilised more than 250 people across Security, Applied and Research to examine hundreds of systems spanning more than 100 service areas. The company describes the work as an incident-response-style sprint that took priority over ordinary work except critical business operations and later became the basis for a continuous security programme it calls the Defense Factory.

OpenAI president Greg Brockman supplied a separate measure of the engineering cost in an a16z interview published on 14 September. He said OpenAI put the projects of 25% of its production engineers on hold and reassigned those engineers to defensive work using the company’s models. The 25% figure applies to production engineers, not to OpenAI's total workforce, and OpenAI's Defense Factory page does not publish the same percentage.

The workflow links discovery to reproduced fixes

OpenAI's published architecture runs a five-stage loop: inventory, discovery, dynamic validation, ownership assignment and verified remediation. Agents receive source code, security findings and service context inside isolated, reproducible development environments; the control plane handles orchestration, policy and credentials, while people review consequential changes and deployed fixes are retested before a remediation is treated as verified.

The company says Codex generated all remediation patches during the sprint. It reports that 53 urgent or high-priority issues were closed on the first day, 90.6% of agent-routed ownership assignments were accepted, 37% of findings were identified as duplicates and 19.5% of findings were reproduced at runtime. After dynamic validation, OpenAI reports a 0.81% false-positive rate and a 0.53% rollback rate for fixes. Those figures describe OpenAI's own operation and have not been independently audited.

Brockman says Astra found serious issues before the search saturated

Brockman said OpenAI pointed Astra at its own systems, found a number of serious issues and fixed them. He then said the exercise eventually saturated, meaning the company had found, to its knowledge, the P0 issues that Astra was capable of identifying at that point. His account is consistent with OpenAI's broader first-party description of frontier cyber models being used to find, validate and fix vulnerabilities in production systems.

That saturation claim has a narrow boundary. It says the search stopped producing further P0 findings within Astra's capability, not that OpenAI proved every critical vulnerability had been removed. Brockman also said the process should repeat as new models gain cyber capability, which is why OpenAI is turning the sprint into a standing defensive loop rather than treating the code red as a finished audit.

The disclosed metrics establish scale, not complete security

The material change is the scale and operating model now documented by OpenAI: a large share of production engineering was temporarily redirected to defensive work, more than 250 people participated across over 100 service areas, and agent-assisted vulnerability work was integrated with ownership, patching and post-deployment verification. The company has also published enough of the workflow to show where isolation, credentials and human approvals sit in the loop.

Important evidence remains internal. OpenAI has not published the underlying vulnerability list, an independent audit of the reported false-positive and rollback rates, or evidence that the Defense Factory can discover every severe weakness in its systems. The confirmed story is therefore the code-red response, its measured internal workflow and Brockman's first-party account of the engineering diversion, while claims about comprehensive security remain unsupported.

Source trail

4 sources · 4 primary