This week, leaders of Google, OpenAI, Anthropic, Meta, xAI and Nvidia signed the White House Accord on Super Intelligence. The companies committed to four layers of oversight for frontier models: internal controls, an internal team that verifies them, an independent external auditor or evaluator, and an independent board committee. The commitments are voluntary for now, and the accord says they could eventually be written into law.
Of the four layers, the outside assessment carries the most weight. It’s the only layer that puts someone outside the company in a position to say whether safeguards work. It also has the least definition. The accord doesn’t say who qualifies as an independent assessor, what standard they assess against, how much access they get, or how independence holds up when an evaluator also sells services to the company it reviews. Each company picks its own assessor.
Under those terms, two companies can both announce that they passed an independent assessment and mean very different things. Closing that gap will take answers to three questions.
What Should Decide Who Qualifies
A credible definition of an independent AI assessor needs to cover at least four things.
Independence. An assessor shouldn’t review safeguards it helped design, and it shouldn’t sell remediation or consulting work to the same company it assesses. Mature audit fields also watch fee dependence. When one client accounts for a large share of an assessor’s revenue, the assessor has a reason to go easy.
Competence matched to the job. “Audit” covers several different activities in AI. Testing a model for dangerous capabilities, such as helping with a cyberattack or a biological weapon, requires machine learning researchers and domain experts. Checking whether a company’s safeguards are designed well and working in practice requires experienced controls auditors. An assessor qualified for one isn’t automatically qualified for the other, and an assessment report should state which one was performed.
Access and scope. An assessor who sees only documentation can confirm that a safeguard exists on paper. Confirming that it worked takes access to systems, logs, approval records and people over a period of time. Frontier models change between training runs and deployments, so the assessor also needs a clear rule for when a change requires a fresh look.
Oversight of the assessor. Someone has to check the checkers. In established audit markets, accreditation bodies review assessors’ work and can take away their standing when it falls short. That backstop is what gives an assessment’s conclusions any weight.
The Infrastructure That Already Exists
None of this has to be invented from scratch. ISO 42001, the international standard for AI management systems, gives organizations a framework for governing AI, with documented controls, clear owners and continuous improvement. Its companion standard, ISO 42006, sets requirements for the bodies that audit and certify against ISO 42001, including the competence auditors must demonstrate and the impartiality rules they must follow. Because certification bodies operate under accreditation, a third party oversees the auditors themselves.
On the technical testing side, newer frameworks are filling in. AIUC-1 certifies specific AI agents in specific deployments, with recurring third-party testing for problems such as hallucinations, jailbreaks and unsafe tool use, and builds on ISO 42001’s controls.
States are also testing models for supervising assessors directly. Connecticut passed a law this year creating a pilot program, starting July 2027, in which the state’s Department of Consumer Protection will approve up to five independent verification organizations. Each must enter a state-supervised agreement defining its scope, methods, reporting obligations and governance. That structure answers several of the accord’s open questions: who approves assessors, what they’re held to, and who supervises them.
These pieces weren’t built for frontier model evaluation, and they shouldn’t be presented as a complete answer. The work ahead is connecting them, so that management system assurance, technical model testing and assessor oversight fit together under one credible definition of independence.
Why the Rules Need to Come First
The accord’s commitments are voluntary today. If they become law, the rules for who can serve as an assessor need to be in place before the mandate takes effect, for three reasons.
First, early practice becomes precedent. Once signatories start naming assessors and publishing results, the market will settle on working definitions of “independent” and “qualified.” Definitions set under no outside pressure tend to favor convenience. Lawmakers who arrive later will find those definitions already embedded in contracts and expectations.
Second, qualified capacity takes time to build. Accrediting assessment bodies, training evaluators who understand both frontier models and controls testing, and developing shared methods all take years. A mandate that arrives before that capacity exists will be met by whoever is available.
Third, a requirement without a qualified assessor market behind it produces box-checking. Companies will satisfy the letter of the rule, regulators will have little basis to challenge weak assessments, and the public will get the appearance of oversight without much substance.
Some argue it’s too early to set these rules because the science of evaluating frontier models is still young. That’s a fair concern about testing methods, which should keep evolving as models do. The questions raised here are about the assessor: whether they’re free of conflicts, qualified for the job, given enough access and subject to oversight. Those questions have stable answers, and other assurance fields have been answering them for decades.
What Organizations Can Do Now
Companies building or deploying advanced AI don’t need to wait for a mandate. When choosing an assessor, they can ask what other services the firm provides to them, what qualifications its team holds for the specific type of assessment, how much access the engagement includes, and who oversees the firm’s work. Building an AI management system now also gives any future assessor something concrete and documented to evaluate.
The accord sets up a sound structure for AI accountability. Its value will depend on who fills the assessor role and what they’re held to, and the time to define that is before anyone is required to use it.
Credit: Source link


























