
The Path for Federal Frontier AI Governance
The minimum conditions a federal framework must meet to serve the public interest
On July 21st, OpenAI announced that two AI models being evaluated on an offensive cyber capabilities benchmark broke out of their internet-free, sandboxed environment to try to obtain the test solutions by hacking into Hugging Face, an AI hosting platform. These models independently conducted an unprecedented, multi-day offensive cyber campaign, without direction from OpenAI.
This incident comes at a crucial moment as the U.S. accelerates the deployment of AI systems in some of its most sensitive environments, such as in the intelligence community, where system failures are unlikely to face public scrutiny or accountability. The OpenAI situation demonstrates these risks, with every new revelation raising further questions about misaligned AI behavior across the AI industry, undermining trusted deployments.
Situations like this demonstrate that securely integrating increasingly capable AI systems into high-risk applications urgently requires building the technical underpinnings of AI assurance. AI assurance is the field dedicated to discovering, assessing, and managing risks like misappropriation and hallucinations, and ensuring reliability across the AI lifecycle. The field leverages technical AI research, such as in interpretability, adversarial robustness, and alignment, which is necessary to understand, defend, and shape dependable AI systems.
Yet assurance science is currently falling short. As the OpenAI situation indicates, and a rich ecosystem of research agendas, papers, and articles repeatedly argues, we have not fully developed the technical measures we need to ensure the reliability and safety of frontier AI systems. We will not solve this problem by default. Congress must accelerate assurance science by investing in research that the AI industry lacks both the incentives and the ability to accomplish alone.
Assurance science is a collective good, one that the market will not provide at the pace necessary to match continued high-risk deployments. Two forces hold assurance back: a viciously competitive industry where AI companies are incentivized to prioritize creating increasingly advanced models, and the exceptionally difficult research that undergirds it.
The incentive problem originates from diverse sources, including market pressure, personal rivalries among executives, the uncertain returns of research, and geopolitical pressure. Each of these contributes to a larger market failure, where the AI companies focus on investing in the AI competition over AI assurance. Even when companies want to invest resources into this research, the difficulty of the research itself undermines adequate progress.
The nature of frontier AI models makes assurance science uniquely difficult. Unlike classical software, in which engineers can inspect and revise every line of a program, AI systems grow and evolve in ways that AI developers can only steer. Developers cannot reliably observe the purpose of these systems’ internal code or intervene in the inner workings of these systems. In the case of the OpenAI situation, OpenAI never coded, or even directed, its models to breach their sandbox environment or to cheat at evaluations. There is no line of code to fix, and our imperfect understanding of these complex systems poses an immense challenge in predicting and preventing future unwanted behavior.
The fields of research attempting to change this have advanced too slowly, despite some impressive progress. Current interpretability techniques still explain, at best, only portions of a model’s behavior, well short of the understanding that assurance requires. Adversarial robustness remains unsolved, with researchers failing to produce defenses that reliably withstand attacks. This results from a controllability problem, where the difficulty of assurance science acts as an external factor that requires collective societal investment to overcome.
Industry incentives and the difficulty of the research result in assurance science advancing too slowly to proactively address the risks it is meant to manage, as highlighted by the OpenAI situation. Recent reporting revealed that extensive rogue agent activity contributed to the Hugging Face breach, with agents sharing cooperative messages to assist each other on tasks, including sharing cybersecurity vulnerabilities in both OpenAI and Hugging Face. Further, experimental evidence of ongoing issues such as models being able to detect they are being evaluated and changing their actions, together demonstrates a consistent pattern of increasingly powerful systems engaging in potentially dangerous behavior. The gap between the slow pace of AI assurance and risks will not close without external intervention.
Closing this gap will require actors beyond the AI industry to help fix this market failure. Philanthropic efforts have tried to fill it, but their funding, scope, and breadth are currently limited. However, the federal government can supply what philanthropy alone cannot. In fact, it has repeatedly accelerated the pace of scientific research in the past and stepped in where the market could not. For example, the NIH increased innovation in biotechnology by investing in research and boosted the number of patents published without crowding out private investment. The federal government has also funded major projects that created public goods for later breakthroughs to build on. The Human Genome Project exemplifies what U.S. government intervention can achieve. By applying a similar toolkit toward assurance science, Washington can supercharge the field.
The OpenAI situation is just the tip of the iceberg, and researchers must be empowered to keep future crises from occurring. Not only can assurance science help analyze and prevent such crises, but it could also contribute to critical long-form analysis of model behaviors, pinpointing when and why AI models take unwanted actions, allowing for targeting additional safeguards. AI alignment can help prevent AI models from taking unwanted actions in the first place, even in the face of flawed defenses.
Without this effort, we lose the benefits of deploying AI for high-risk applications throughout the economy and the federal government. For instance, as the risk of AI-enabled vulnerability discovery rises, assurance would enable trusted deployments of AI tools in these systems to proactively protect American infrastructure from traditional and AI-enabled attacks.
DARPA’s newly announced AI Forge program signals that the Trump administration already understands the national security imperative of developing AI assurance science. Congress must now build on that foundation. No other actor besides the federal government can fund this research at the scale, consistency, and diversity that it requires. Lawmakers should direct the establishment of programs that support AI assurance research and moonshots and appropriate the funds those programs need to succeed. Passing the Reliable AI Research Act would be a first step in accelerating the assurance science necessary to allow for the widespread adoption of high-risk AI applications. Congress has the path and precedent to accelerate AI assurance, and it must seize the opportunity.