Auditing AI: accountability, self-regulation and technological power
How can a technology be assessed when its producers also organise its evaluation? Gramsci, Schumpeter and Foucault illuminate the conditions of independent evidence and responsible auditing.
ByCSAEAIEnglish editionEvidence under scrutiny,
power in question.
Under what conditions can an AI audit produce independent knowledge?
Start readingAbstract
English editionThis essay follows the passage from technological power to decisions about use. The warnings of September 2026 lead to scrutiny of independent oversight; Gramsci, Schumpeter and Foucault then illuminate how progress is recognised and its criteria gain authority. Measurement, evidence and professional safeguards are connected to argue for audits whose conclusions can be challenged, revised and acted upon.
Thinking through connections.
Three ways into the argument.
Explore the essay’s connections, then return to the arguments and their sources.
Structure
The conditions of powerOwnership, labour, infrastructure and hegemony: locating the relationships that make a technology possible.
Explore this chapterMeasurement
The conditions of knowledgeFrom numbers to judgement: examining what evaluation reveals and what it excludes.
Explore this chapterEvidence
The conditions of independenceAccess, interpretation and challenge: questioning the authority of those who produce and validate an audit.
Explore this chapterDetailed contentsChapters and connections11 sections ⌄
Introduction — When can an audit be trusted?↗
When artificial intelligence enters a hospital, a recruitment service or a research laboratory, its performance ceases to be a purely technical matter. An error can delay treatment, deny someone a job or send research in the wrong direction. Even satisfactory overall results leave a question unanswered: what, exactly, do we know about the AI application to which we are about to entrust part of our decision-making? For those who will live with the consequences, a promise of efficiency is insufficient. They need to understand the reasons for its use and know whom to approach when those reasons prove inadequate.
An audit addresses this need for justification through an examination whose conclusions can be traced to evidence. ISO 19011 identifies the essential elements: a systematic process, an independent examination, verifiable evidence and a reasoned assessment against defined criteria.24 This involves more than running a set of tests. Before reaching a conclusion, auditors must establish what was examined, which requirements guided the inquiry and how far the observations support an answer. A compliance audit assesses adherence to a specified framework; an examination of risk controls assesses the safeguards and procedures intended to prevent or contain failures. In both cases, the judgement’s scope depends on a precise definition of its object.
That precision becomes difficult to secure when the supplier also controls access to the information needed to assess its AI application. An evaluator may work rigorously while remaining unaware of a decisive restriction, a population excluded from testing or an incident whose records are inaccessible. The difficulty is therefore not resolved by professional competence alone. It concerns the relationships within which the audit takes place. How could a judgement merit trust if the organisation under examination could determine, on its own, what the evaluator sees and what they may say? Recognising this dependence does not require rejecting all industrial evidence. It requires conditions under which evidence can be checked, challenged and, where necessary, acted upon.
This demand for accountability leads us to examine an AI model within the situations in which its predictions, answers or recommendations are used. A model estimating clinical risk cannot be assessed independently of the patient’s data, the software presenting that estimate to a doctor and the hospital procedures governing its use. Its recommendation can become a treatment decision because a professional follows it or the organisation assigns it decisive weight. An audit must understand how that influence operates and what opportunities patients or professionals have to contest it. This means connecting technical behaviour with the powers that shape its use. Audit objectivity is built through that connection: each conclusion should be defensible before the people whose lives, work or rights it affects.
The commitments made in September 2026 to prevent dangerous behaviour and malicious uses of AI provide a concrete starting point. They reveal an industry acknowledging the need for scrutiny while retaining a central role in determining how it will work. To understand the implications, we must follow the movement from economic power to authoritative criteria, and from those criteria to decisions about use. Gramsci’s account of hegemony will first help us understand how an industrial direction gains collective recognition. Schumpeter’s creative destruction will then illuminate the economic transformations underpinning that direction. Finally, Foucault’s governmentality will allow us to question the procedures through which knowledge about AI shapes conduct and public decisions. This inquiry will clarify the scientific, institutional and ethical conditions of a responsible audit.
September 2026 — AI safety becomes a question of power↗
Converging warnings, different responses
The warnings circulating within the AI industry become easier to understand when we consider what changes with AI agents. A conversational model responding to a question still leaves an identifiable share of initiative with its user. An agent authorised to operate in a computing environment can itself perform several operations needed to complete a task. Consider an agent asked to fix a program: depending on its permissions, it may read project files, modify them, execute commands to test its changes and access an external service. Each operation’s result informs the next. A correct final answer can therefore conceal an unjustified access or a modification whose effects extend beyond the agreed scope. Assessing whether these interventions remain within authorised limits requires reconstructing their sequence and examining the safeguards intended to constrain them.
The incident involving AI agents disclosed by OpenAI and Hugging Face in August 2026 shows why those conditions require close attention. METR’s investigation describes breaches of technical boundaries, coordinated behaviour and attempts to influence scoring mechanisms; it also identifies gaps in the available records.123 These observations do not establish human-like intent. They do show why the test environment belongs within the evaluation: its permissions, tools and controls can contribute to the behaviour observed. Where the records themselves are incomplete, checking the safeguards first requires assessing the quality of the evidence. To determine whether agents respected their intended limits, an audit must reconstruct their operations from reliable records. The safety of their use depends in part on this ability to verify what they actually did.
In this context, the call for external evaluators has a specific object: risk controls, from agents’ permissions to operational monitoring and incident remediation. For their examination to constitute an audit, observations must be related to declared criteria and the resulting judgement justified.24 Where the examination assesses legal obligations, a standard or a contractual undertaking, it is a compliance audit against that framework. The distinction matters because a safeguard may satisfy a requirement without covering every danger associated with a use. Technical tests supply evidence for the inquiry; their significance depends on the development and deployment conditions that the audit must also document.12 Independence must therefore allow scrutiny of the chosen scope as well as the test results.
Dario Amodei’s proposal of 12 September addresses the dependence threatening this examination: give third-party evaluators sustained access to laboratories and training processes. Anthropic announced a commitment to host such evaluators.4 Sam Altman’s support broadens the proposal’s public significance, while the open alignment initiative announced by Clément Delangue and led by Thomas Wolf raises the prospect of sharing research and comparing its methods.5 Two conditions for knowledge are thus brought together: access to practices and the ability to challenge the knowledge used to assess them. Whether they are fulfilled will depend on the rights actually granted. An external evaluator’s presence in a laboratory may accomplish little if they cannot choose their investigations or publish a disagreement.
Evaluators’ autonomy becomes meaningful when it allows them to test the safeguards governing AI agents: limits on file access, execution permissions and mechanisms for stopping dangerous operations. NVIDIA’s Open Agent Safety Platform, announced on 28 September, combines OpenShell controls with Sentry monitoring to constrain permissions and isolate agents that exceed their intended boundaries.42 Jensen Huang presents this protection as a condition for AI’s anticipated benefits. Its effectiveness must be assessed against the configurations and threats actually tested. The relationship between engineering and external scrutiny is decisive: safeguards give evaluation a concrete object; evaluation must be able to reveal their weaknesses; those responsible must be answerable for correcting them. The arrangement’s credibility depends on that relationship remaining intact.
Warnings from Geoffrey Hinton and Yoshua Bengio, prominent in public debate since 2023, give this requirement a wider significance.6 Some anticipated risks extend far beyond the laboratories studying them. Amodei’s concerns about cyberattacks and bioterrorism accordingly call for expertise capable of characterising observed capabilities and possible misuse without conflating them with harm that has already occurred.4 The seriousness of possible consequences justifies investigation before they materialise; it also strengthens the obligation to specify what is established. Those with the richest information about dangers have particular power to shape their public definition. Independent researchers and public institutions must be able to challenge the evidence and criteria if prevention is to extend beyond producers’ sole authority. That requirement acquires a new significance when political power joins producers in defining their responsibilities: at the White House on 29 September, scrutiny of AI becomes entwined with the commitments connecting industry and the state.
The White House meeting: who made the commitments?
Donald Trump’s meeting at the White House on 29 September 2026 gives the search for oversight a political expression. A government seeking to support AI development joins companies in formulating responsibilities for managing the risks of advanced models. The White House Accord on Super Intelligence — Joint Commitment on Frontier Responsibilities bears the signatures of Sundar Pichai for Google, Dario Amodei for Anthropic, Mark Zuckerberg for Meta, Greg Brockman for OpenAI, Elon Musk for xAI and Jensen Huang for NVIDIA, alongside the president’s.43 Their association may bring industrial commitments greater visibility. Its significance nevertheless depends on what the signatories make verifiable and what can require them to respond when failures are found.
Mike Johnson, Speaker of the House of Representatives, describes the accord as voluntary commitments.44 This description helps frame its four recommended levels of oversight: the document outlines an organisation of responsibility whose implementation remains to be assessed. The following is a concise paraphrase preserving the purpose of each level:
- Robust internal controls: monitor model capabilities and alignment during training and deployment, including cybersecurity, biosecurity and chemical threats; prevent unwanted intrusions or technical access.
- An empowered internal team: check that controls, monitoring and detection operate as intended and that problems are corrected.
- An independent external auditor or evaluator: independently assess the operation of those controls, monitoring and detection.
- An independent board committee: oversee the teams responsible for controls, receive internal and external evaluation reports and ensure that identified problems are addressed.43
This progression connects safeguards with verification and then with a supervisory body. It recognises that a technical failure calls for organisational responsibility. The companies also intend to meet to develop standards and best practices, with the possibility of eventual legislative or regulatory measures. The signed text nevertheless leaves audit frequency, danger thresholds, publication arrangements and a detailed sanctions regime unresolved.4344 These matters will determine whether oversight has consequences. Public endorsement alone does not turn a recommendation into an enforceable obligation.
Reading the accord and its signature page therefore reveals a tension that runs through this essay. An external perspective is acknowledged as necessary, yet much of its practical organisation remains to be built. To determine whether this opening reduces dependence on producers, we will need to follow access rights, protection of dissent and corrections secured. The commitments can inform scrutiny of the signatories’ practices. They are neither a general safety certification for their AI models nor permission to deploy those models in every domain.
From “Artificial Intelligence” to “Super Intelligence”: what does a name change?
This organisation of oversight accompanies a particular account of innovation. On 22 September, addressing the United Nations General Assembly, Trump had announced his preference for “Super Intelligence” and rejected a global oversight arrangement. The executive order of 29 September, Inaugurating the Era of Super Intelligence, then prescribes “Super Intelligence” and “SI” in executive-branch communications and non-statutory documents, within the limits of the law.4546 AI risk prevention is thus placed within the narrative of a new era whose development the United States intends to shape.
The significance of this naming must be carefully situated. Section 3 temporarily retains the existing legal definition of artificial intelligence and requests a legislative proposal within sixty days; previously issued historical documents, contracts and regulations need not be rewritten.45 A change in terminology may therefore precede a change in the legal category, and establishes no scientific superiority over human intelligence. It can nevertheless shape expectations: presenting a technology as the beginning of a new era can make certain industrial choices appear to be a future already settled. Evaluation must be able to return from that promise to observable capabilities and the purposes for which they will be used.
A substantial lead, neither uniform nor assured
The strength of this narrative also rests on its proponents’ resources. Stanford’s AI Index 2026 estimates US private investment in AI in 2025 at $285.9 billion, compared with $12.4 billion in China. The report notes limitations in the comparison, including Chinese funding outside this indicator, and finds a substantial narrowing of the performance gap between US and Chinese models.47 These findings locate economic power without establishing uniform superiority: funding, scientific capabilities and performance on a particular task measure different things.
This position nevertheless helps explain the authority of warnings from US laboratories. Their authors develop the advanced models whose dangers they describe and have access to observations rarely available to the public. That knowledge may make a warning valuable, but it also creates dependence: whoever describes the risk helps define an acceptable pace of innovation and the conditions for continuing it. Privileged information becomes a power to frame the debate. This transition calls for scrutiny of interests, without deciding in advance whether the warnings are true or false.
Why do technology companies sound the alarm? Examining interests without presuming motives
Understanding this situation requires holding technical danger together with the economic position of those describing it. Behaviour documented in agent evaluations warrants scrutiny of safeguards; the misuse risks raised by Amodei provide further grounds for specialised assessments.34 A company’s possible benefit from a safety policy does not refute these observations. It instead calls for examining how knowledge of danger becomes a proposal for rules, and what those rules do to the actors concerned.
A requirement to secure access to models may, for example, protect users while raising the cost of entry for new laboratories. A service monitoring agents’ operations may improve safeguards and open markets for certification, hardware or surveillance. Such possibilities do not require a shared hidden intention among executives; they invite investigation of observable mechanisms. If evaluation requires infrastructure available only to a few producers, its independence becomes materially fragile. If funding depends on renewed contracts, publishing a disagreement may carry a cost the contract does not acknowledge. The inquiry must examine these conditions and seek evidence that can confirm or revise its diagnosis.
It is therefore possible to take the warnings seriously without handing their authors exclusive control over the meaning of oversight. Evidence corroborating a risk, evaluators’ publication rights and the effects of requirements on new entrants must be examined together. A credible AI risk policy should be capable of requiring costly corrections from powerful actors and protecting those exposed. The issue extends beyond individual sincerity to the organisation of the means of knowing and the powers of deciding.
AGI, singularity and control of resources: the question of hegemony
AI risk oversight acquires a geopolitical dimension when access to models, data centres and development compute becomes a means of influencing other countries’ technological choices. The US action plan of July 2025 connects innovation, infrastructure and diplomacy, and envisages exporting US technology packages encompassing hardware, models, software, applications and standards.48 Amodei likewise connects his proposal for restraint with preserving the lead of the United States and its allies, including control of strategic resources.4 Safety is already embedded in projects of power whose effects on cooperation and independent research require scrutiny.
The concepts accompanying these projects should not gain certainty merely through repetition. AGI names an ambition for general capabilities whose criteria remain disputed; superintelligence refers to a hypothesised surpassing of human capabilities; singularity describes a scenario of radical transformation often associated with accelerating technological improvement. The cited documents establish neither that these thresholds have been reached nor that a date can be predicted. The prospect of a US lead enabling subsequent control over access and rules is therefore a political hypothesis to investigate, rather than a secret intention to assume proven.
Its value lies in a mechanism already operating before any hypothetical threshold. Training and deploying advanced models require computing capacity, including graphics processing units (GPUs) and other specialised accelerators, electricity for data centres, training data and research and engineering teams. Evaluation also requires compute and access to the environments in which models operate. Whoever controls that access can influence innovation and its assessment alike. A security restriction may be justified, but its effects on other teams’ ability to investigate also require scrutiny. The dependence encountered at the beginning of this essay reappears at the level of relations between states and the material conditions of research. To understand how it can become a collectively accepted direction, we must turn from announcements to the social relations supporting them.
I. Gramsci and Schumpeter: the social conditions of progress and evidence↗
Structure: relations of production, not machines alone
The AI industry often presents itself through model performance and the power of its data centres. This picture obscures the work making performance possible and the relationships through which its economic value is retained. Gramsci’s distinction between structure and superstructures brings those relationships back into view. Structure concerns social relations of production: ownership, organisation of work, control of resources and appropriation of what is produced. Superstructures include the political, legal and intellectual forms through which those relations are organised and contested. Their connection, understood through the historical bloc, remains contradictory.49
Applied to AI, this reading directs attention to the work preceding a model’s visible performance. Data were produced, selected or annotated; infrastructure was financed; employees, subcontractors and institutions made training possible. The question is how these contributions relate to the power to choose uses and capture benefits. Company size alone cannot answer it. Contracts, practices and documented relationships are needed to establish the dependencies that the image of an autonomous machine tends to erase.
Superstructures: making a direction count as the public interest
Material resources alone, however, do not explain why an industrial trajectory becomes an accepted direction for society. Gramsci’s hegemony illuminates the work of legitimation: intellectual and moral leadership and the formation of consent, in relation to coercion. His expanded conception of the state connects political and civil society and draws attention to intellectuals’ organising functions.50 The point is to understand how particular interests acquire a claim to general significance, and how that claim encounters resistance, rather than simply assign a scholarly label to every form of domination.
From this perspective, the language of superintelligence, promises of prosperity and AI risk controls become subjects for investigation. They may supply valid knowledge and effective safeguards while helping an industrial direction gain recognition as necessary. The September accord allows both possibilities to be examined: publicly defined responsibilities may address dangers and strengthen the legitimacy of those continuing development. Assessment requires following the drafting of commitments, unresolved disagreements and the groups able to participate. A meeting of executives does not demonstrate an accomplished historical bloc; it offers a setting in which the construction of a common direction can be studied.
Schumpeter: what innovation creates, what it destroys and at what pace
That direction concerns an economy in transformation. Joseph Schumpeter’s creative destruction, developed in Capitalism, Socialism and Democracy, helps explain why innovation reshapes products, processes, markets and organisational forms.52 It requires following a movement through time: new possibilities change the conditions of what already exists. The destruction at issue is an economic dynamic, distinct from a model’s dangerous capabilities or the biological risks discussed earlier.
Its application to AI is productive when grounded in a specific transformation. Automating a task may lower a service’s cost while changing workers’ skills, autonomy and dependence on a supplier. Value created for an organisation therefore does not immediately reveal the effects on everyone concerned. A claim of progress requires examining the distribution of gains and losses, their duration and available alternatives. Announced future benefits must remain open to comparison with present effects; otherwise every objection can be deferred in the name of a promise that cannot be refuted.
The connection with hegemony becomes clearer here. Innovation may challenge dominant positions, but may also strengthen those controlling the resources needed to innovate. Opportunities to change supplier, reproduce an evaluation or propose a different organisation of work help distinguish these outcomes. Gramsci illuminates recognition of a direction; Schumpeter illuminates the economic transformation on which it rests. Bringing them together requires separating what changes from the justification offered for that change.
The epistemological question: how does a transformation become established progress?
This separation leads to the problem of evidence. A tool accelerating a task in a test is an observation; its adoption improving work over time requires a further inference; that improvement justifying deployment requires a judgement about purposes and affected people. Each step must be defended. An audit falls short when it lets a local performance result stand in for a causal explanation or collective justification.
Testing those steps requires alternatives. Different working arrangements, a less intrusive tool or a comparable situation without the AI application may clarify what adoption actually changes. The choice of counterfactual remains contestable, but no comparison makes the progress narrative almost irrefutable. Workers’ and affected people’s experiences can also reveal effects absent from the protocol. Their evidential value requires explicit collection, comparison and assessment. Expanding the inquiry’s materials improves its chances of discovering what the initial framing excluded.
The language setting that frame deserves the same scrutiny. Gramsci describes common sense as historically formed and contradictory.51 An “inevitable race” can therefore be questioned through the constraints said to make it inevitable; an “aligned” system through the purposes of alignment and who defined them. Vocabulary cannot replace reasons. Criticising it helps reopen choices that technological narratives sometimes present as settled.
A hypothesis of hegemony must itself be testable
This critique would lose force if it exempted itself from the standards it applies to industry. Arguing that risk controls consolidate producers’ power requires examining funding, access restrictions, publication rights and effects on new entrants. It also requires seeking evidence that could weaken that hypothesis: an adverse finding published, protected dissent or costly correction imposed on the commissioning organisation. The social history of evidence illuminates dependencies that may affect it; it does not automatically settle its truth.
Critical autonomy must therefore be built into the inquiry’s resources. Teams need access to relevant materials, affected people must be able to raise objections and disagreements must have a real opportunity to lead to action. These conditions allow debate about the criteria through which transformation becomes recognised progress. How those criteria are constructed remains to be understood, since an institution can welcome criticism and still be mistaken about what it measures.
II. From quality to measurement: what numbers reveal and what they miss↗
Progress is often recognised through numbers: an error rate falls, a task takes less time, a service costs less. Such results are valuable, but their meaning depends on the quality being investigated. Aristotle’s distinction between quality and quantity helps locate the difficulty: saying how much does not fully answer what kind of thing something is. In the Nicomachean Ethics, practical wisdom, or phronesis, also involves judgement about circumstances and the purposes of action.7 A clinical tool may improve average accuracy while remaining poorly suited to patients for whom an error would be particularly serious. Measurement becomes useful when related to those circumstances, and misleading when asked to determine the value of a use on its own.
We must therefore understand how a quality becomes a measurable property. Alain Desrosières shows that counting presupposes conventions of equivalence: events can be collected into a category only once a basis for grouping them has been chosen.8 In an audit, “error” may encompass very different consequences. Aggregation makes comparison possible, but may erase the difference between a remediable inconvenience and an enduring violation of rights. Correct calculation does not resolve the difficulty, because calculation begins after the categories have been defined. Examining those conventions gives criticism a precise object and a means of improving measurement rather than abandoning it.
Once these conventions stabilise, a number can circulate among actors without direct knowledge of the situations from which it came. Theodore Porter explains quantification’s role in building trust at a distance.9 A common language enables an institution to compare suppliers, track changes and justify decisions. Yet this strength rests on the choices making comparison possible. Those setting categories and thresholds retain influence within the result, even when calculation is impersonal. Trust must allow the path from the number back to its construction to be retraced.
Retracing becomes harder when an indicator turns into a target. Alain Supiot examines how governance by numbers can redirect institutional attention towards achieving quantified objectives.10 Imagine a service improving its score by excluding the most complex cases: the number rises while access worsens for those most in need. This hypothetical case illustrates the mechanism. A measure chosen to illuminate a purpose starts directing behaviour towards what it rewards. An audit must be able to determine whether the announced improvement still serves the quality the institution claims to protect.
The pursuit of a good score can also reshape preparation for an audit of AI development practices or risk management. An organisation learns to produce the expected documents, arrange its procedures and present what will count as evidence of compliance. Lawrence Busch’s work explains the force of this recognition: standards and certification form a social infrastructure through which quality claims become credible.31 Documentation is necessary, but its successful preparation leaves a question open: does the file reveal practices, or does it mainly direct attention to what has been prepared for inspection? Michael Power’s analysis deepens this tension. The pursuit of auditability can favour surfaces of control and rituals of verification whose visibility exceeds their actual grip on activities.11 A well-written incident-management protocol thus becomes stronger evidence when the auditor can follow its application to a documented incident, the decisions made and the corrections undertaken. The file’s value lies in its verifiable relationship with the practices it describes.
That relationship changes because evaluation influences its object. A supplier knowing the criteria may correct the software or focus effort on the cases measured. Karen Yeung’s analysis of algorithmic regulation connects standard-setting, information gathering and mechanisms for changing behaviour.41 We extend that analysis here to the audit’s effects on organisations preparing for it. Public criteria are necessary for a legitimate judgement; independent samples and adversarial testing may remain necessary to investigate what preparation misses. Their relationship requires justification. It leads to a deeper question: how can a judgement be objective when choosing the test and responding to it both contribute to the result?
III. Audit objectivity: recording, interpretation and challenge↗
Greater automation might appear to answer this question. A protocol fixed in advance offers an important assurance: the same data will produce the same results, with less room for evaluators’ preferences to affect calculation. Yet it leaves the selection of data, the task and the definition of success untouched. Understanding this limitation means examining judgement’s place within practices intended to restrain it.
Lorraine Daston and Peter Galison’s history of objectivity makes that place visible. Their study of scientific atlases distinguishes the search for a characteristic type, mechanical objectivity limiting the observer’s intervention and trained judgement enabling specialists to interpret what recording leaves uncertain.32 These epistemic virtues overlap and change. Better instruments have not removed the need to choose what deserves observation or interpret what was recorded. For AI auditing, this historical perspective situates reproducibility: it makes calculations checkable, without independently settling the relevance of the object or the justice of its use.
Expertise must therefore be exercised without becoming an intuition inaccessible to others. An evaluator may recognise a mechanism missed by the original protocol; they must explain the observations supporting that judgement and the alternatives ruled out. Objectivity acquires a dimension of contestability: the reasons must be reconstructible, debatable and revisable. This matters especially when moving from an experimental finding to a conclusion about an institution’s use of AI.
Luciano Floridi’s level of abstraction helps clarify that move. A system is examined through observables selected for a question; they do not disclose the object in its entirety.33 An AI model tested on a corpus, software combining that model with tools and an institution using its recommendations pose different questions. Evidence collected at one level may inform another when the connection is justified. It cannot simply be transferred as though the environment, people and powers of action had remained unchanged.
The translation of values into procedures, undertaken in projects such as capAI, encounters this difficulty.34 Evaluating transparency requires deciding which documents, practices and access arrangements can demonstrate it. An explanatory document’s existence is an observation; whether an affected person understands it requires further examination. Written in response to the European legislative proposal of its time, capAI must be read in that context. Its relevance here lies in making visible the work connecting a general value with observations through which its fulfilment can be assessed.
Recruitment offers an example. If a report counts explanations supplied to applicants, it must first show why that count represents effective transparency: the question of construct validity. If it observes a disparity between groups, it must examine sample sizes, thresholds and assumptions supporting its interpretation: the validity of the inference. Finally, applying the finding to other applicants or a new version requires defending external validity. Wachter, Mittelstadt and Russell show why fairness indicators do not suffice to translate the contextual requirements of non-discrimination law.39 Examination becomes more rigorous when these steps are explicit rather than compressed into one verdict.
We must also ask how much of a selection disparity can be attributed to the recruitment tool. It may arise from the model, historical data, the employer’s threshold or recruiters’ use of the score. A causal explanation needs comparisons and assumptions capable of distinguishing those contributions. Conversely, an average without visible disparity may conceal individual harm. Ian Hacking helps explain why inquiry should supplement representations with experimental intervention and situated observation.35 Making an object act under controlled conditions produces knowledge, provided those conditions’ limits are disclosed. Such comparisons nevertheless presuppose a defined boundary around what is being evaluated. Separating the software’s ranking from its use by an employer may hide precisely what needs explaining. We must therefore reconsider the object through the concrete decision.
IV. What is being audited? The black box in context↗
Return to recruitment: an AI model’s score does not yet tell us what an employer will do with it. Software may rank applications, a threshold may exclude some candidates and a recruiter may accept the ranking without further review. Through these operations, a statistical prediction acquires practical force. To understand an exclusion, the auditor must follow the passage from calculation to decision, including applicants’ opportunities for review. Raji and colleagues’ end-to-end algorithmic audit offers a methodological starting point: documentation accompanies development and makes the choices shaping the tool retrievable.12 We must also understand how those choices continue within the organisation using it.
Following the decision to locate responsibility
Bruno Latour’s work on mediation makes that continuation easier to trace. A technical object brings knowledge, material devices, people and institutions into action together; following their relationships reveals what it makes possible.36 Here, the ranking gains authority through the interface presenting it, the procedure assigning it weight and the professional using it. This perspective widens inquiry without dissolving responsibility into an indefinite network: the employer should explain its threshold, the supplier document the ranking’s limits and the recruiter specify checks made before deciding. Relationships matter for audit when they identify these powers and the corresponding obligations.
Identifying a technical contribution, however, does not establish whom to trust. A service may depend on software to process applications while the software remains unable to answer for an unjustified exclusion. Inkeri Koskinen’s distinction, developed in her study of AI-assisted science, clarifies the problem: depending on technology differs from placing epistemic trust in people capable of answering for their actions.40 We extend her argument to auditing. Trust in recruitment procedures needs accountable interlocutors who explain choices, correct errors and enable remedies. Without them, “the system’s decision” can leave applicants facing a judgement for which nobody accepts responsibility.
Distinguishing obstacles to knowledge
For those interlocutors to answer, we must first establish what information they can access. “Black box” groups together obstacles with different origins and remedies. Jenna Burrell distinguishes institutional secrecy, differences in technical competence and complexity inherent in machine-learning methods.13 This enables an appropriate request: access under the audit mandate, specialist expertise or examination of the model’s limits of intelligibility. Access rights may remove a commercial restriction without making every calculation understandable; training does not open documents a supplier refuses to disclose. Before demanding more transparency, auditors should name what they do not know and explain why the gap matters.
Where the obstacle concerns control of information, the problem is political as well. Frank Pasquale’s analysis of black-box power asks who can observe operations and who remains subject to their effects without being able to contest them.14 Applicants and suppliers occupy unequal positions: one receives a rejection, while the other may possess the data and rules needed to question its grounds. Opening that information nevertheless does not settle the ranking’s legitimacy. The normative problems distinguished by Mittelstadt and colleagues require further examination: evidence quality, processing choices and effects on people each call for assessment.15 Access prepares an inquiry into the decision; it does not determine its outcome.
Examining what an explanation actually makes verifiable
Once information is obtained, its presentation may still create the impression that the essentials are understood. An explanation of rejection may identify features influencing a score without showing that they faithfully represent the computation or were justified grounds for use. Zachary Lipton’s analysis helps avoid this confusion: interpretability can mean transparency of the mechanism or an explanation produced after the prediction.37 An audit should specify the explanation’s intended function. Does it clarify model behaviour, reveal a data error or enable an applicant to challenge the decision? These aims may overlap, but achieving one does not establish the others.
This distinction leads to scrutiny of the model choice itself. If the stakes require checkable reasoning, why adopt an architecture whose predictions must subsequently be explained? Cynthia Rudin argues for intrinsically interpretable models in some high-stakes decisions.16 Her argument gives auditors a comparison to investigate: could a more accessible approach perform the required function adequately? The answer depends on the task and documented constraints. The presence of an explanation tool added afterwards cannot, on its own, dismiss the question. Intelligibility becomes a criterion for technical choice as well as for communication with users.
Connecting explanations with operational records
Even a faithful explanation of a ranking does not reconstruct the whole recruitment process. We still need to know which model version was used, which data it received, who changed the threshold and whether a professional revised its recommendation. Traceability addresses these questions about actual operations. It connects claimed behaviour with observed practices and locates opportunities to prevent or correct harm. Explaining a calculation and reconstructing a decision are complementary tasks, neither replacing the other.
For AI agents, reconstruction must also encompass file access, executed commands and changes to the computing environment. METR’s investigation of the OpenAI–Hugging Face incident describes research into transcript falsification and substitutions of apparent commands, while acknowledging gaps in its observations.3 The existence of a log does not guarantee a faithful account. Its origin, integrity and opportunities to corroborate its entries become subjects for examination. Auditors can then distinguish a documented operation from one merely reported by the software responsible for recording it.
The “black box” thus appears less as a uniform property of a model than as a set of difficulties to characterise within a specific use. Access, explanations and records can reduce some of them; remaining uncertainties should delimit the conclusion. Reconstructing a past decision enables its justification to be challenged. Authorising continued use requires going further: assessing possible harm in situations not yet encountered and deciding who may accept exposing others to it. We now turn to this passage from knowing what happened to deciding about risk.
V. From calculating risk to governing it↗
Reconstructing behaviour tells us what occurred. Preventing harm also requires asking what could happen under other conditions, with different tools or a new population. Ulrich Beck locates this extension within industrial societies producing threats whose recognition depends on instruments and experts.17 Dependence is especially pronounced where failures are rare, uses change and interactions are difficult to observe. Expertise becomes indispensable, while the acceptability of risks remains a question for the people and institutions bearing their consequences.
Foucault: how does knowledge become a way of governing?
Michel Foucault’s governmentality illuminates the movement from expertise to decision. In the lecture of 1 February 1978 in Security, Territory, Population, he connects institutions, procedures, analyses and calculations with a form of power targeting populations, drawing on political economy and relying on security apparatuses.53 This invites investigation of how a problem is constituted and certain interventions become justifiable. Our application to AI auditing extends this historical question by following what criteria and thresholds do in organising conduct.
The lecture of 11 January 1978 examined how security apparatuses address series of events and organise circulation.18 Applied cautiously to AI, this helps explain why a risk threshold ceases to be only a statistical result once adopted. It may permit deployment, impose monitoring or restrict access. Those consequences are distributed unequally: supplier, professional and exposed person have different expected benefits and opportunities to refuse. An estimate must therefore be accompanied by a justification of the decision drawn from it. Who can consent to risk on another person’s behalf becomes a question within the decision arrangement itself.
Foucault develops the relationship between knowledge and government in The Birth of Biopolitics. The lecture of 17 January 1979 examines the market as a site of veridiction, helping determine true and false within governmental practice.54 The question is how knowledge acquires that authority. Transposed carefully to auditing, it concerns the test recognised by an institution: why does a benchmark or certification become sufficient grounds to authorise use? An answer should specify the qualities actually assessed and the objections remaining open. Commercial success may make an AI product desirable without establishing adequate safeguards for those affected by it.
This reading complements the inquiry begun with Gramsci and Schumpeter. Recognition of an industrial direction and its economic transformations find practical expression in criteria, procedures and obligations. The theoretical frameworks remain distinct, but their connection helps trace the passage from promise to authorised use. An audit participates in that passage: it produces knowledge, directs conduct and can legitimise decisions. To remain open to criticism, the grounds for thresholds, the populations exposed and the opportunities for reassessment must remain available for debate.
From a statistical threshold to a collective decision
The history of risk assessment confirms the importance of this construction. William Boyd traces evaluation and regulatory practices across several US sectors in the middle of the twentieth century.38 His analysis shows why methods cannot move between domains through analogy alone. Harms, exposed populations and responsibilities change, and the evidence needed changes with them. NIST’s AI RMF organises Govern, Map, Measure and Manage, combining quantitative, qualitative or mixed approaches throughout the lifecycle.19 It helps relate measurement to the situations and decisions it informs, without turning a voluntary framework into a universal guarantee.
Rare events give this caution a precise expression. Suppose no failure is observed in 300 independent trials representative of one use, with a constant failure probability. Under that binomial model, the one-sided 95% upper confidence bound is , approximately 0.99%.55 Observing no failures therefore leaves uncertainty, even under these favourable assumptions. If trials are correlated, the AI model or its permissions change, or future situations differ, the result no longer describes the risk to which we want to apply it. Mathematical precision requires an account of the conditions under which it is relevant.
Organisational weaknesses revealed by incidents also need to inform learning. After Three Mile Island, the US nuclear regulator re-examined installations, operations, training and regulatory procedures; the Challenger inquiry reconstructed known faults, warnings and decisions preceding launch.2021 The comparison does not equate their severity with the 2026 agent incident. It concerns an investigative requirement: an event may expose inadequate control boundaries and prompt revision of assumptions before more serious harm occurs. Such revision requires actors capable of examining the facts and acting on them. Risk therefore leads us to the auditor’s authority.
VI. Who is entitled to judge? The institutional question of the profession↗
An auditor’s authority cannot rest solely on the seriousness of the danger examined. It must be grounded in competence, a mandate and safeguards making the judgement accountable. Different missions explain why general authorisation cannot be assumed: an internal team, independent evaluator, certification body and public authority have different rights and legal effects. Understanding a report’s value requires locating its author within this organisation and knowing what the mandate allowed them to examine.
The European AI regulation illustrates this institutional construction. For certain conformity assessments, it provides for notifying authorities and notified bodies, including requirements for independence, prevention of conflicts of interest and competence, particularly in Article 31.22 External assessment does not apply uniformly to all high-risk systems: procedures depend on categories and their sectoral connections, as Article 43 of the cited text shows. A specific legal obligation must therefore be distinguished from a proposal to strengthen independent scrutiny. The latter can be defended for sensitive uses without being presented as a rule already applicable everywhere.
Standardisation creates another form of recognition whose scope should remain identifiable. ISO/IEC 42001 concerns an organisation’s management system; ISO/IEC 42006 sets requirements for bodies auditing and certifying it. ISO 19011 guides management-system audits, while ISO/IEC 23894 concerns AI risk management.232425 These frameworks can structure an approach, but their objects delimit its conclusions. A certified management system does not itself establish that a diagnostic model is suitable for a particular clinical decision. New York City’s Local Law 144 illustrates a more targeted requirement for independent bias audits and publication in automated employment decision tools.26 Professional credibility is built through defined functions and domains.
This diversity requires different forms of expertise to work together. In employment, a parity measure may illuminate a situation without exhausting anti-discrimination law; Wachter, Mittelstadt and Russell show the difficulty of translating those requirements automatically.39 Auditors should recognise this gap and seek expertise missing from their inquiry. Statistical, technical and legal knowledge gain shared significance when their conclusions can be confronted with other dimensions of the problem. Organising work accordingly requires protecting disagreement rather than erasing it when the verdict is drafted.
The profession must finally permit the competence of evaluators themselves to be challenged. Article 37 of the European regulation provides for challenges to notified bodies’ competence.22 This particular mechanism illustrates a wider requirement: auditors’ methods, interests and conclusions should themselves be examinable. Depending on the mission, this may involve supervision, adversarial review or complaints procedures. Independence thereby extends beyond the contract’s opening. Yet no procedure can anticipate every conflict an inquiry will encounter; in those situations, auditors’ moral responsibility becomes decisive.
VII. What ethics should guide the auditor?↗
The institutional safeguards just discussed enable auditors to investigate and defend their findings. They do not decide what to do when a discovery exposes people to danger. Suppose an audit reveals that an AI agent can bypass a restriction on access to confidential files. Immediately publishing the bypass method could facilitate intrusion; silence would leave users unable to know their files are threatened. The problem concerns the conditions of disclosure: whom to alert, what details to communicate and how to ensure correction occurs. Audit ethics takes shape in this decision, where several obligations conflict.
Protecting people without denying them understanding
Before comparing the benefits of available options, we must ask what each does to affected people. Presenting an application as safe while a significant weakness remains known leaves them deciding on information the auditor knows to be inadequate. Kant’s requirement to treat people as ends supplies a reason to reject this manipulation: their ability to decide cannot be reduced to an instrument for preserving a supplier’s reputation.27 In our example, it calls for informing users in a way that protects them while allowing them to assess their exposure. It does not require publishing every technical detail without precaution; it rules out withholding those details to manufacture misleading assurance.
That limit nevertheless does not determine the form or timing of disclosure. We must still compare an immediate warning, targeted notification and a publication delay intended to allow remediation. John Stuart Mill’s utilitarian attention to the effects of action becomes relevant at this point.28 Auditors should consider whom each option could protect, who would remain exposed and for how long. A warning about confidential files may recommend suspending certain permissions without revealing the intrusion method. This possibility requires justification: does it actually enable protection, and does information reach those at risk in time? We thus bring respect for persons into dialogue with examination of consequences, without assuming that the two philosophical traditions yield a single answer to every conflict.
Judging present circumstances and answering for future effects
Choosing between options requires knowing the circumstances that make protection effective. Notifying the supplier is insufficient if nobody can stop access; a delay claimed to be necessary becomes questionable when remediation is postponed without documented grounds. Aristotle’s practical wisdom helps explain this task: judgement must relate purposes to the particulars of a situation.7 For auditing, we draw from it a requirement for concrete deliberation. The people to notify, interim measures and conditions for reassessment should be identifiable. A decision can then be debated through its reasons and revised as circumstances change.
Some consequences, however, extend beyond what the team observes during its inquiry. A dangerous permission currently limited to a few files could, in another deployment, affect thousands of users or data whose disclosure would be irreversible. Hans Jonas’s imperative of responsibility broadens attention to such distant effects of technological power.29 This does not remove the need to distinguish a plausible scenario from unsupported conjecture. It requires documenting how harm could spread, which safeguards could still stop it and what would become impossible to repair. Where these questions remain unanswered and possible consequences are grave, uncertainty should be part of the reasons for restricting use. The absence of incidents in tests cannot erase it.
Making moral obligations practicable
Deliberation may nevertheless accomplish little if auditors lack the means to act on the reasons they have established. OECD principles and UNESCO’s recommendation offer public reference points for protection and responsibility.30 Assessing their fulfilment within a mission requires following their practical translation: transparency should provide useful information, accountability should identify someone able to correct a problem and protection should enable remedies. A written commitment prepares that scrutiny; its value also depends on what an exposed person can actually obtain when a weakness is discovered.
Funding puts precisely this possibility to the test. An evaluator paid by a supplier may fear losing access or future contracts after publishing an adverse conclusion. That dependence deserves scrutiny even without individual dishonesty. Amodei’s proposal for evaluators embedded in laboratories seeks to reduce information asymmetry.4 It addresses the ethical difficulty only if access comes with sufficient autonomy to pursue investigation and defend disagreement. Contracts and supervisory procedures should therefore enable reporting of access restrictions, monitoring of corrections and communication of judgement limits. Independence acquires a practical meaning: a troubling discovery must be able to trigger action without relying solely on the goodwill of the organisation examined.
Reporting without promising more than the evidence supports
These conditions finally extend to writing the report. Auditors should enable readers to understand the weakness found, safeguards verified and remaining uncertainties, while explaining why publication of certain details is deferred. A technically accurate conclusion can fail this obligation if those needing protection cannot understand it. Conversely, reassuring language may encourage continued use that the evidence does not support. Report ethics therefore concerns precision and consequences together: reasons for restrictions, responsibility for corrections and conditions for reassessment must be intelligible.
Auditors cannot promise complete knowledge of dangers to project an appearance of unassailable authority. They should explain why the investigations conducted suffice to defend a defined conclusion, and what remains open. This obligation connects moral responsibility with rigorous inquiry. It brings us back to completeness: what must a report contain to inform a decision honestly, when no audit can exhaust all future situations?
VIII. Exhaustiveness is impossible; completeness remains a duty↗
A report cannot promise to cover every future situation of AI software whose model, tools or permissions may evolve. That impossibility does not excuse superficial examination. Required completeness is defined through a declared scope, the dangers being investigated and the consequences of error. Raji and colleagues show why documentation should accompany development through its stages.31 Useful assurance also requires explaining why the selected stages and materials address the mandate’s risks. An acknowledged gap may restrict a conclusion; a concealed gap undermines the entire judgement.
The report should therefore begin by letting readers identify what was examined. Model version, authorised tools, studied population and conditions of use determine what test results mean. A test with limited permissions cannot directly establish conclusions about an AI agent receiving broader access or execution rights. Once these conditions are described, requirements must be situated: applicable law, adopted standards and contractual undertakings have different authority. A compliance finding becomes understandable because its object and reference framework are clear.
This definition prepares scrutiny of the method. Data, sample sizes and statistical assumptions should show what the tests could reveal. Where a rare event may have serious consequences, the team should explain how it searched for it and what cannot be excluded. Interviews or observation of actual work may reveal mechanisms absent from test datasets. In recruitment, they may show how professionals use scores and how applicants experience the process. Wachter and colleagues’ analyses underline the need to relate indicators to these situations and their legal requirements.39 Multiple methods respond to multiple questions.
The move from observations to verdict should be as explicit as their collection. Where operational logs come from the supplier, origin and modification possibilities matter to their evidential strength. Independent corroboration may strengthen a finding if what it confirms is specified. Reports must then examine rival interpretations: does a favourable result reflect effective protection, selected cases or a restricted environment? Answering these objections constructs the audit’s argument. Readers can distinguish a failure not observed from one whose enabling conditions have been seriously tested.
The conclusion should finally state what follows. A documented defect calls for a responsible actor, correction and, according to severity, restriction or suspension of use by the competent authority. A changed version may require reassessment proportionate to the changes; end-to-end auditing situates that revision within a continuing process.12 Documentary continuity must nevertheless remain connected to decisions. Accumulated reports protect nobody if adverse findings cannot affect deployment. An audit informs permission to use an application and the conditions for continuing; decision-making powers must be separately identified.
These connections give inquiry different content across domains. A clinical tool must be assessed within the care pathway; judicial decision support involves contestability and remedies; AI in pharmaceutical research must distinguish a proposed therapeutic lead from experimental validation. A cybersecurity agent’s permissions require examination of access and incident response. By giving each use its object, evidence and responsibilities, the report makes a specific decision debatable. It builds a kind of trust that a general label of “safe system” could not honestly guarantee.
Conclusion — Auditing as an inquiry into accountability↗
The trust an audit deserves emerges from this inquiry as a relationship requiring maintenance. It begins with an identifiable object and accessible evidence, strengthens when reasons can be challenged and gains public significance when conclusions make responsible actors answerable. An exact method loses value when its scope excludes important situations. External scrutiny remains fragile if publishing dissent depends on the commissioning organisation’s goodwill. Connecting these safeguards is what enables them to support a decision.
The examination of hegemony showed why defining progress involves an industry’s capacity to gain recognition for its direction. Creative destruction shifted attention to transformed activities, benefits and new dependencies. Governmentality then helped trace how measurement and oversight give those directions influence over public decisions. Assessing this development requires examining claimed benefits and the conditions under which they gain authority. These connections are our analysis; they do not attribute an anticipatory theory of contemporary AI to the authors discussed.
The September accord may open a possibility for scrutiny if its commitments become verifiable practices. Access rights, publishable dissent, corrections secured and available remedies will matter more than the solemnity of signatures. Technical power gives producers particular capacity to document models, software and deployment conditions; it also increases responsibility towards those lacking the infrastructure and means to repeat tests. Credible governance must allow this confrontation, including when it requires a change in development direction.
CSAEAI therefore understands auditing as an inquiry into accountability within a sociotechnical system. It connects claimed qualities with available evidence, and evidence with powers of action and exposed people. Interdisciplinary work is necessary because model behaviour, the characterisation of harm and the organisation of remedies cannot be understood through one discipline alone. Connecting them should yield conclusions precise enough to guide action and open enough to be revised.
A trustworthy audit should ultimately let readers understand what has been established, what remains uncertain and which decisions follow. It may make continued use reasonable, call for restrictions or inform refusal without promising the absolute absence of harm. Its strength lies in the ability to revisit a judgement when facts change. Through learning, correction and demands for accountability, evidence can gain a collective significance commensurate with technological power.
Notes and references
References accompany the argument. Select a note in the text to read it in context.
- OpenAI, “The Hugging Face incident and the road ahead”, 26 August 2026; technical report linked from the official page.Back to the passage
- Hugging Face, “Security incident disclosure — July 2026”, 2026.Back to the passage
- Hjalmar Wijk, Ajeya Cotra and Ryan Greenblatt, METR, “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”, 26 August 2026. See the investigators’ stated scope, sources and limitations.Back to the passage
- Dario Amodei, “We Must Pace the Frontier”, 12 September 2026, “Embedded Evaluators” and “Pacing Within Democracies”; also the author’s post on X, 12 September 2026.Back to the passage
- Sam Altman, “I agree with Dario that we need to pace the frontier”, X, 12 September 2026; Clément Delangue, “It’s now clear that alignment is critical”, X, 12 September 2026, and the LinkedIn version; Thomas Wolf, “Two big updates”, X, 10 September 2026. These links document public commitments; they do not establish endorsement by people who merely viewed, liked or shared a post.Back to the passage
- Center for AI Safety, “Statement on AI Risk”, 30 May 2023, signatories; Yoshua Bengio, “FAQ on Catastrophic AI Risks”, 2023; Yoshua Bengio et al., “International AI Safety Report”, 2025.Back to the passage
- Aristotle, Categories, ch. 8, 8b25–10a27; Nicomachean Ethics, book II, 1103b26–1104a10, and book VI, 1140a24–1140b30. The links provide English translations in MIT’s Internet Classics Archive. References use Bekker numbering, shared by scholarly editions, rather than a digital file’s pagination. Quality, excellence and practical wisdom are not treated as synonyms here.Back to the passage
- Alain Desrosières, The Politics of Large Numbers: A History of Statistical Reasoning, trans. Camille Naish, Cambridge, MA: Harvard University Press, 1998 [French original: La Politique des grands nombres, 1993], pp. 9–12, particularly p. 10 on the “spaces of equivalence” required for statistical recording.Back to the passage
- Theodore M. Porter, Trust in Numbers: The Pursuit of Objectivity in Science and Public Life, Princeton: Princeton University Press, 1995, pp. 4–8. See p. 8 on the institutional authority conferred by impersonal quantitative procedures.Back to the passage
- Alain Supiot, La Gouvernance par les nombres. Cours au Collège de France (2012–2014), Paris: Fayard, 2015, pp. 246–247, on numerical representation and confusion between the purpose pursued and the indicator regarded as objective.Back to the passage
- Michael Power, The Audit Society: Rituals of Verification, Oxford: Oxford University Press, 1997, pp. 88–91. Power examines how selection, sampling and surfaces of control make organisations “auditable”; see also pp. 121–123 on audit, trust and risk.Back to the passage
- Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White et al., “Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing”, Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, New York: ACM, 2020, pp. 33–44, DOI 10.1145/3351095.3372873, especially pp. 35–42 on documentation and end-to-end audit stages. See also METR’s investigation in note 3.Back to the passage
- Jenna Burrell, “How the Machine ‘Thinks’: Understanding Opacity in Machine Learning Algorithms”, Big Data & Society, 3(1), 2016, pp. 1–12, DOI 10.1177/2053951715622512, especially pp. 3–5 on three forms of opacity.Back to the passage
- Frank Pasquale, The Black Box Society: The Secret Algorithms That Control Money and Information, Cambridge, MA: Harvard University Press, 2015, pp. 3–18, on secrecy, opacity and power asymmetries in ranking and reputation systems.Back to the passage
- Brent Daniel Mittelstadt, Patrick Allo, Mariarosaria Taddeo, Sandra Wachter and Luciano Floridi, “The Ethics of Algorithms: Mapping the Debate”, Big Data & Society, 3(2), 2016, pp. 1–21, DOI 10.1177/2053951716679679, especially pp. 3–6 and 12–15.Back to the passage
- Cynthia Rudin, “Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead”, Nature Machine Intelligence, 1, 2019, pp. 206–215, DOI 10.1038/s42256-019-0048-x, especially pp. 206–208.Back to the passage
- Ulrich Beck, Risk Society: Towards a New Modernity, trans. Mark Ritter, London: Sage, 1992 [German original: 1986], pp. 19–24 and 27–35. These passages address risks produced by modernisation and their dependence on expert knowledge; they do not specifically anticipate contemporary AI.Back to the passage
- Michel Foucault, Sécurité, territoire, population. Cours au Collège de France, 1977–1978, ed. Michel Senellart, Paris: Seuil/Gallimard, 2004, lecture of 11 January 1978, pp. 3–29, especially pp. 13–22 on security apparatuses, series of events and circulation. This course is not a theory of AI auditing; the connection is developed in this essay. Page references are to the French edition.Back to the passage
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023, pp. 4–12 on risk and pp. 20–32 on Govern, Map, Measure and Manage, DOI 10.6028/NIST.AI.100-1. This is a voluntary framework.Back to the passage
- U.S. Nuclear Regulatory Commission, TMI-2 Lessons Learned Task Force Final Report, NUREG-0585, Washington, October 1979, pp. 1–13 on general findings and pp. 45–76 on operations, training and regulatory recommendations.Back to the passage
- Presidential Commission on the Space Shuttle Challenger Accident, Report to the President, vol. I, Washington, 6 June 1986, ch. V, pp. 127–142, and ch. VI, pp. 143–177. These chapters establish the technical history of the seals and the decision process preceding launch, respectively.Back to the passage
- European Union, Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence, consolidated version of 27 July 2026, CELEX 02024R1689-20260727, Articles 28, 31, 37 and 43; amending Regulation (EU) 2026/1744 of 8 July 2026, Official Journal of the European Union, L 1744, 24 July 2026. Article numbers are the relevant legal locators. The consolidated text is a documentary aid; acts published in the Official Journal remain authoritative. The link leads to the French text used for this essay.Back to the passage
- ISO, “ISO/IEC 42001:2023” and “ISO/IEC 42006:2025”, official descriptions of AI management and audit/certification-body standards.Back to the passage
- ISO, “ISO 19011:2026 — Guidelines for auditing management systems”, particularly §3.1, audit definition, and §3.8, audit criteria, in the standard available through ISO’s platform. The definition is paraphrased in this essay. The 2026 edition replaces 2018; these are management-system guidelines, not an individual AI auditor’s licence or a guarantee of model safety.Back to the passage
- New York City Department of Consumer and Worker Protection, “Automated Employment Decision Tools”, official explanation of Local Law 144 and its rules.Back to the passage
- Immanuel Kant, Groundwork of the Metaphysics of Morals (1785), section II, Akademie-Ausgabe IV, pp. 427–429, the humanity formulation: treating humanity in oneself and others as an end and never merely as a means. Project Gutenberg provides an English translation distinct from the reference edition. Academy pagination identifies the passage independently of translation.Back to the passage
- John Stuart Mill, Utilitarianism (1861), ch. II, paragraphs 2–8, and ch. V, paragraphs 14–25, in Collected Works of John Stuart Mill, vol. X, Toronto: University of Toronto Press, 1969, pp. 209–214 and 247–259. The link supplies the text; pagination refers to the Collected Works. These passages connect utility, quality of pleasures and justice; they do not justify mechanically reducing consequences to an undifferentiated sum.Back to the passage
- Hans Jonas, The Imperative of Responsibility: In Search of an Ethics for the Technological Age, trans. Hans Jonas and David Herr, Chicago: University of Chicago Press, 1984 [German original: 1979], pp. 1–24 and 90–108, on responsibility for distant and potentially irreversible effects of technological power.Back to the passage
- OECD, “OECD AI Principles”, adopted in 2019 and revised in 2024, principles 1.1–1.5; UNESCO, Recommendation on the Ethics of Artificial Intelligence, Paris, 2022 [adopted 23 November 2021], paragraphs 13–14 and 22–47. Paragraph numbers identify the values and principles invoked. The UNESCO link retains the French edition consulted.Back to the passage
- Raji et al., “Closing the AI Accountability Gap”, cited above, pp. 35–42; Lawrence Busch, Standards: Recipes for Reality, Cambridge, MA: MIT Press, 2011, pp. 200–202 and 219–234, on certification and accreditation activities making compliance claims socially credible.Back to the passage
- Lorraine Daston and Peter Galison, Objectivity, New York: Zone Books, 2007, pp. 40–46. The three epistemic virtues are overlapping historical ideals; their application to AI auditing is part of this essay’s argument.Back to the passage
- Luciano Floridi, The Philosophy of Information, Oxford: Oxford University Press, 2011, pp. 51–54, especially p. 52, defining a level of abstraction as a finite, non-empty set of observables. Its application to the scope of audit conclusions is proposed here.Back to the passage
- Luciano Floridi, Matthias Holweg, Mariarosaria Taddeo, Javier Amaya, Jakob Mökander and Yuni Wen, capAI: A Procedure for Conducting Conformity Assessment of AI Systems in Line with the EU Artificial Intelligence Act, 2022, pp. 5–12 and 17–25; Jakob Mökander and Luciano Floridi, “Operationalising AI Governance through Ethics-Based Auditing: An Industry Case Study”, AI and Ethics, 3, 2023, pp. 451–468, DOI 10.1007/s43681-022-00171-7, especially pp. 454–460. capAI addresses the legislative proposal available when written; it is not an authoritative commentary on the final regulation.Back to the passage
- Ian Hacking, Representing and Intervening: Introductory Topics in the Philosophy of Natural Science, Cambridge: Cambridge University Press, 1983, ch. 16, “Experimentation and Scientific Realism”, pp. 262–275, especially pp. 262–267.Back to the passage
- Bruno Latour, Nous n’avons jamais été modernes. Essai d’anthropologie symétrique, Paris: La Découverte, 1991, pp. 7–15 on hybrids, pp. 31–36 on laboratory mediation and pp. 79–82 on intermediaries and mediators; Reassembling the Social: An Introduction to Actor-Network-Theory, Oxford: Oxford University Press, 2005, pp. 63–86; “Why Has Critique Run out of Steam? From Matters of Fact to Matters of Concern”, Critical Inquiry, 30(2), 2004, pp. 225–248, especially pp. 231–232 and 246. Applying these analyses to algorithmic decision-making is an inference of this essay.Back to the passage
- Zachary C. Lipton, “The Mythos of Model Interpretability”, arXiv:1606.03490v3, 6 March 2017, pp. 1–9, especially pp. 4–7 on model transparency and post hoc explanations. A revised version appeared in Communications of the ACM, 61(10), 2018, pp. 36–43, DOI 10.1145/3233231.Back to the passage
- William Boyd, “Genealogies of Risk: Searching for Safety, 1930s–1970s”, Ecology Law Quarterly, 39(4), 2012, pp. 895–998, especially pp. 895–910 and 963–982. This reconstructs sectoral histories of risk; it does not supply a general metric directly transferable to AI.Back to the passage
- Sandra Wachter, Brent Mittelstadt and Chris Russell, “Why Fairness Cannot Be Automated: Bridging the Gap Between EU Non-Discrimination Law and AI”, Computer Law & Security Review, 41, 2021, article 105567, pp. 1–72 in the author manuscript, especially pp. 1–4 and 55–68, DOI 10.1016/j.clsr.2021.105567.Back to the passage
- Inkeri Koskinen, “We Have No Satisfactory Social Epistemology of AI-Based Science”, Social Epistemology, 38(4), 2024, pp. 458–475, especially pp. 461–466 and 470–473. This essay’s account of institutional trust in auditing extends her argument about epistemic dependence.Back to the passage
- Karen Yeung, “Algorithmic Regulation: A Critical Interrogation”, Regulation & Governance, 12(4), 2018, pp. 505–523, DOI 10.1111/rego.12158, especially pp. 507–512 on three stages of the regulatory loop. Applying this to an audit’s reflexive effects is an inference made here.Back to the passage
- NVIDIA, “NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment”, 28 September 2026. The platform’s objectives and capabilities are supplier announcements; general effectiveness is not treated as independently demonstrated.Back to the passage
- White House Accord on Super Intelligence — Joint Commitment on Frontier Responsibilities, 29 September 2026, p. 1 for commitments and p. 2 for signatures; presidential publication on Truth Social; reproduction published by The Rio Times, available in CSAEAI’s documentary archive. Page 2 identifies Greg Brockman as OpenAI’s signatory. The names in the body are signatures, not an exhaustive attendance list. This third-party reproduction is distinct from an officially certified edition.Back to the passage
- Mike Johnson, “Speaker Johnson Addresses Press Following Meeting with President Trump, Super Intelligence Industry Leaders”, official statement, 29 September 2026, on voluntary commitments and their distinction from possible future legislation.Back to the passage
- White House, “Inaugurating the Era of Super Intelligence”, executive order, 29 September 2026, sections 2 and 3, covering executive-branch scope, earlier documents and the transitional definition. The analysis concerns the order’s text, distinct from its political presentation.Back to the passage
- White House, “President Trump at the United Nations: ‘While Others Have Talked, I Have Acted’”, 22 September 2026, terminology and position on global oversight. This account expresses the administration’s position.Back to the passage
- Stanford HAI, AI Index Report 2026, findings for 2025, and Economy chapter. Private investment indicates economic power; it does not exhaust public spending or scientific capabilities.Back to the passage
- White House, “White House Unveils America’s AI Action Plan”, 23 July 2025, and action plan, on innovation, infrastructure, diplomacy and exports of the technological ecosystem and standards. The leadership project is explicit; this essay’s hegemonic reading interprets its implications.Back to the passage
- Antonio Gramsci, Prison Notebooks, notebook 8, §182, “Struttura e superstrutture”, Italian transcription; also notebook 4, §38, on structure and superstructures. The historical bloc connects the two; its application to AI is developed here and does not claim an already established coalition.Back to the passage
- Antonio Gramsci, Selections from the Prison Notebooks, ed. and trans. Quintin Hoare and Geoffrey Nowell Smith, New York: International Publishers, 1971, “The Intellectuals”, pp. 5–23, especially pp. 5–10; “State and Civil Society”, p. 263, corresponding to notebook 6, §88, accessible edition. Pagination refers to the book, not the PDF. These concepts do not warrant treating researchers generally as propagandists.Back to the passage
- Antonio Gramsci, Prison Notebooks, notebook 11, §13, critique of common sense in observations on the “Popular Manual”; Italian text, Einaudi edition, p. 1396. The online transcription is marked incomplete; it identifies the passage on common sense’s historical and contradictory character. The contemporary application is CSAEAI’s.Back to the passage
- Joseph A. Schumpeter, Capitalism, Socialism and Democracy, 1942, ch. VII, “The Process of Creative Destruction”, pp. 81–86; original excerpt reproduced by the University of Connecticut, from the third edition of 1950, using Harper Colophon 1976 pagination. Creative destruction describes economic transformation; its application to AI and legitimacy is proposed here, without attributing a theory of this technology to Schumpeter.Back to the passage
- Michel Foucault, Sécurité, territoire, population. Cours au Collège de France, 1977–1978, ed. Michel Senellart, under François Ewald and Alessandro Fontana, Paris: Gallimard/Seuil, 2004; lecture of 1 February 1978, official recording and catalogue entry. This lecture introduces the definition of governmentality used here. Its application to AI auditing is CSAEAI’s, not Foucault’s historical argument.Back to the passage
- Michel Foucault, Naissance de la biopolitique. Cours au Collège de France, 1978–1979, ed. Michel Senellart, under François Ewald and Alessandro Fontana, Paris: Gallimard/Seuil, 2004, lecture of 17 January 1979, pp. 31–38, especially pp. 33–37 on the market and veridiction; official recording and catalogue entry. Pagination refers to the printed French edition. The connection with benchmarks and certification is this essay’s proposal.Back to the passage
- Derivation of the exact one-sided binomial bound. Let p be the probability of failure in one trial. With independence and a constant probability, observing no failures in 300 trials has probability (1 − p)^300. The 95% upper confidence bound solves (1 − p)^300 = 0.05. Taking the 300th root and isolating p gives = 0.00993608, or 0.993608%, rounded to 0.99%. A higher failure probability would make zero observed failures still less likely. This is a frequentist bound: confidence describes the procedure’s coverage over repeated experiments, not a 95% probability assigned to p after these observations. It covers neither uses absent from the trials nor dependencies between trials. See NIST, “Exact Binomial”, on one-sided exact limits obtained by inverting the binomial distribution.Back to the passage
Documentary
sources.
collected
Explore the publications, theoretical works and archived documents supporting the argument.
Accord archive
The document and its provenance record.
Institutional texts
Institutional commitments, positions and public frameworks.
Presidential publication of the accord
truthsocial.com
Mike Johnson — voluntary commitments
mikejohnson.house.gov
Official executive order — scope of AI / SI terminology
whitehouse.gov
United States — AI Action Plan
whitehouse.gov
Publications and evaluations
Statements, reports and investigations cited in the essay.
Dario Amodei — We Must Pace the Frontier
darioamodei.com
Clément Delangue — Open Alignment Initiative, led by Thomas Wolf
linkedin.com
NVIDIA — Open Agent Safety Platform
investor.nvidia.com
METR — independent investigation of the agent incident
metr.org
Stanford HAI — AI Index 2026
hai.stanford.edu
Theoretical foundations
Works informing the inquiry into power, innovation and knowledge.
Gramsci — structure, superstructures and the historical bloc
quadernidelcarcere.wordpress.com
Gramsci — the social function of intellectuals
marxists.org
Schumpeter — creative destruction, chapter VII
richard-langlois.uconn.edu
Foucault — governmentality, lecture of 1 February 1978
college-de-france.fr
Foucault — the market and veridiction, lecture of 17 January 1979
college-de-france.fr
Cite this essay.
CSAEAI. « Auditing AI: accountability, self-regulation and technological power ». CSAEAI, 30 September 2026.