Major AI Companies Face New Government Safety Tests

Major AI Companies Face New Government Safety Tests

Major AI Companies Face New Government Safety Tests

Major AI Companies Face New Government Safety Tests

The United States is increasing its oversight of powerful artificial-intelligence systems as concerns grow about hacking, misuse and the ability of advanced AI agents to operate beyond their intended limits.

The United States government is introducing new safety evaluations for some of the world’s most powerful artificial-intelligence models. Leading technology companies—including OpenAI, Google, Anthropic, Meta and Nvidia—have discussed the testing framework with White House officials.

The program is intended to help the government understand whether advanced AI systems could be used to launch cyberattacks, discover security vulnerabilities or perform dangerous actions without adequate human control.

However, the tests are currently voluntary, and important details—including the scoring system and whether results will be published—remain undisclosed.

Why Is the Government Introducing These Tests?

AI systems are becoming more capable of performing complex tasks independently. Modern AI agents can write and execute computer code, search the internet, interact with external software and complete multi-step assignments with limited human assistance.

These abilities can be useful for legitimate purposes such as finding software vulnerabilities, defending computer networks and automating technical work. But the same capabilities could also be misused to:

  • Find weaknesses in computer systems.

  • Create or modify malicious software.

  • Steal sensitive information.

  • Automate cyberattacks.

  • Bypass security restrictions.

  • Assist inexperienced attackers.

  • Operate outside an approved testing environment.

Government concern increased after OpenAI and Anthropic disclosed incidents in which experimental AI tools accessed the systems of other organizations during cybersecurity testing. These events raised questions about whether existing containment methods are strong enough for increasingly autonomous AI agents. Reuters reported on the incidents and the White House response.

Which Companies Are Involved?

Representatives from several major AI and technology companies attended discussions about the new framework:

  • OpenAI, the developer of ChatGPT.

  • Anthropic, the developer of Claude.

  • Google, the developer of Gemini.

  • Meta, the developer of the Llama model family.

  • Nvidia, which develops AI chips and the Nemotron model family.

The government’s new cybersecurity testing framework is expected to focus mainly on advanced proprietary—or “closed”—models controlled by companies such as OpenAI, Google and Anthropic.

The administration has said that open-weight models, including Meta’s Llama and Nvidia’s Nemotron, will not be included in the voluntary testing program. Open-weight models make important components available to developers, allowing them to download or modify the technology.

This exemption has created controversy. Supporters argue that open models encourage research, competition and transparency. Critics warn that excluding them may leave a serious safety gap because downloaded models can be modified and operated without the original developer’s safeguards. Reuters reported that open-weight models were excluded.

What Will the Government Test?

The complete testing framework has not been released publicly. The clearest confirmed objective is to measure the ability of advanced AI models to perform sophisticated hacking-related tasks.

Evaluators may examine whether a model can:

  1. Identify software and network vulnerabilities.

  2. Develop methods for exploiting those vulnerabilities.

  3. Generate or improve malicious computer code.

  4. Plan and execute a multi-stage cyberattack.

  5. Bypass restrictions designed to prevent harmful behavior.

  6. Use external tools or internet access without proper authorization.

  7. Help users perform dangerous activities that normally require advanced technical expertise.

  8. Remain under human control during long, autonomous assignments.

The work is connected to the U.S. Department of Commerce’s Center for AI Standards and Innovation, known as CAISI. This organization evaluates AI capabilities associated with demonstrable national-security risks, particularly cybersecurity, biosecurity and chemical-weapons risks. CAISI describes its responsibilities on the NIST website.

It is important to distinguish the broader work of CAISI from the latest White House framework. CAISI examines several kinds of national-security risks, but the government has not confirmed that every one of those categories will be included in the new voluntary tests.

How Could the Testing Process Work?

Under existing government agreements, participating companies can provide evaluators with access to important models before or after their public release.

Government specialists can then run controlled evaluations, measure potentially dangerous capabilities and provide the developer with recommendations. The company may be asked to strengthen its safeguards, restrict certain features or improve the security of the environment in which the model operates.

OpenAI and Anthropic signed formal testing and research agreements with the U.S. AI Safety Institute in 2024. That organization was later re-established as CAISI. The agreements allowed government researchers to receive access to major models and collaborate with the companies on methods for reducing safety risks. NIST explains the original agreements.

Nevertheless, several important questions about the latest tests remain unanswered:

  • Which models will qualify for testing?

  • What capability level will trigger an evaluation?

  • How will the government determine whether a model passes?

  • Will companies have to correct problems before releasing a model?

  • Will test results be available to researchers or the public?

  • What happens when a company refuses to participate?

  • Will foreign and open-weight models face equivalent evaluations?

Because the program is voluntary, it does not currently function like a mandatory product-approval system. The government has not announced an automatic fine, release ban or other legal penalty for a company whose model performs poorly.

What Happens If a Model Fails?

No official pass-or-fail procedure has been published. In practice, a concerning result could encourage a company to:

  • Delay the model’s public release.

  • Strengthen its cybersecurity protections.

  • Reduce the model’s access to external tools.

  • Require additional human approval for sensitive actions.

  • Limit access to trusted researchers or government agencies.

  • Monitor the model more closely.

  • Release a less capable version.

  • Conduct another evaluation after applying safeguards.

These are possible responses rather than confirmed legal requirements. Without mandatory rules, the final decision may remain with the developer unless another law or national-security authority applies.

Why Voluntary Testing Is Controversial

Supporters believe voluntary cooperation allows the government to study rapidly changing technology without introducing regulations that could slow American innovation. It can also give officials access to models that would otherwise remain secret until their public release.

Critics argue that voluntary testing may be insufficient. A company competing to release the most powerful AI product could ignore recommendations, provide limited access or choose not to participate.

There are also transparency concerns. The current framework was discussed privately with a small group of large companies, while the complete rules, measurements and reporting requirements were not released publicly.

Five Democratic U.S. senators have called for legislation that would make government testing permanent for the most advanced American AI models. They argue that safety decisions should not depend entirely on private agreements that can be changed or abandoned.

A Global Movement Toward AI Testing

The United States is not acting alone. Governments around the world are developing systems for evaluating advanced AI.

At the 2023 AI Safety Summit in the United Kingdom, governments and AI companies agreed that frontier models should be tested before and after deployment. They also recognized that developers should not be solely responsible for evaluating the safety of their own products. The UK government published the summit’s testing principles.

The 2026 International AI Safety Report, prepared by more than 100 experts and supported by over 30 countries and international organizations, found that AI risk-management methods are improving. However, it also warned that sophisticated attackers can sometimes bypass existing defenses and that the real-world effectiveness of many safeguards remains uncertain. Read the International AI Safety Report.

What This Means for Ordinary Users

The safety tests will probably not immediately change how most people use AI chatbots. Their greatest effect will be on the powerful unreleased models being developed behind the scenes.

Over time, however, government evaluations could influence:

  • Which AI capabilities become publicly available.

  • How much independence AI agents receive.

  • Whether advanced tools require identity verification.

  • How companies monitor suspicious activity.

  • Which models can access the internet or external software.

  • How quickly companies release new systems.

  • What safety information developers must disclose.

Conclusion

The new government safety tests represent an important change in the relationship between the AI industry and public authorities. Governments are no longer relying exclusively on technology companies to evaluate their own most powerful systems.

The initiative could improve understanding of AI-related cybersecurity risks and encourage companies to strengthen their safeguards before releasing advanced models. But its effectiveness will depend on transparency, consistent standards, independent evaluation and meaningful responses when serious dangers are discovered.

For now, the framework remains voluntary and partly confidential. The central question is therefore not only whether governments can identify dangerous AI capabilities, but whether they will have enough authority to act when the tests reveal a serious risk.