Preprint Agents and audit
From Connectors to Harnesses
Runtime Action Spaces, Forensic Readiness, and the Auditability of Agentic AI
A study of how AI agents acquire real-world capability through connectors, skills, permissions, tools, and execution environments — and why auditing them requires reconstructing the runtime that made an action possible.
- Capability, execution, and evidence can change independently.
- A current system snapshot is not enough to explain a historical action.
- The target is minimum sufficient evidence: enough retained material to reconstruct and test a consequential action, with retention and access proportionate to the risk.
- Agentic AI creates an auditability problem as much as a model-safety problem.
- Length
- 17 pages
- Version
- 1.0
- Venue
- Zenodo preprint
- Published
- 19 August 2026
On this page
public preprint agentic AI AI audit
The model is only one part of the system that acts. To understand what an AI agent could do — and what it actually did — we also need the runtime around it.
00 The problem begins after the model
AI governance still often starts with the model.
Which model was used?
How was it evaluated?
What risks did its provider disclose?
What did the model output?
Those questions become incomplete once an AI system can act.
A model connected to a repository may be able to read code. Give the same system write permission and it can modify files. Add a skill and it can follow a reusable workflow. Place it inside a harness with credentials, memory, shell access, network access, approval rules, and external tools, and it can carry a task across multiple systems.
The underlying model may remain unchanged while the operational system around it becomes capable of something materially different.
That creates a different audit question:
What has to be reconstructed when an AI system performs a consequential action?
The answer cannot be recovered from the model name alone.
01 From connectors to harnesses
The paper follows three nested layers through which operational capability becomes increasingly integrated.
External operations become available through tools, APIs, repositories, databases, files, and other connected systems.
Instructions, scripts, references, and resources package reusable ways of carrying out a task.
Tools and procedures are coordinated with permissions, credentials, state, orchestration, execution environments, and approval logic.
These are analytical layers, not a universal history of how every agent is built. The progression matters because each layer adds something that may change what the operational system can actually do. A connector can expose a new external operation. A skill can redirect how an existing operation is used. A harness can coordinate those operations across time, state, authority, and multiple systems.
02 Three different questions hide inside one “agent”
Descriptions of an AI system often flatten several different states into a single account of what the agent can do. For audit, they have to be separated.
What operations were technically available under the active tools, skills, permissions, credentials, approval rules, and environment?
Which trajectory did the system actually take under the context, intermediate results, retries, routing, and state it encountered?
Which parts of that trajectory survived logging, event transformation, retention, export, integrity, and access controls?
The distinction changes how runtime changes should be interpreted. A broader credential can make a previously unavailable operation possible. A modified tool description or skill can redirect behaviour while permissions remain unchanged.
A summary-only logging policy can leave the action itself unchanged while destroying the evidence needed to reconstruct it later.
Capability, execution, and evidence can change independently.
That is why a current system snapshot is not enough to explain a historical action.
03 Why reconstruction matters
Consider a coding agent that damages a production database. Knowing which model generated the commands tells us very little by itself.
An investigation may need to recover:
- which environment the agent was operating in;
- which connectors, tools, and skills were active;
- what permissions and credential scopes were available;
- what instructions and retrieved information entered the task;
- which commands or API calls were executed;
- what retries, routing, or intermediate state changed the path;
- what a human reviewer actually saw and approved;
- what low-level events were retained;
- and which external records can independently corroborate the platform’s account.
The important evidence therefore begins before the damaging command. Historical runtime configuration establishes why an operation was available in the first place. Execution provenance establishes how the system moved from that available operation to the action that actually occurred. This shifts the role of forensic readiness.
For an agentic system, runtime history is itself evidentiary material.
The audit object is no longer only the output or final action. It is the historical configuration, the executed trajectory, and the surviving record connecting the two.
04 The evidence follows a supply chain
Agentic execution rarely belongs to one technical component or one organisation.
A single task can cross:
model provider → agent platform → skill source → connector server → repository → CI runner → cloud account → external API → deploying organisation
The records needed to explain that task may be divided across the same chain. One party may retain model events. Another may hold connector logs. A cloud provider may hold network or resource records. The deploying organisation may hold approval context.
A repository or external API may provide the only independent evidence that a disputed operation actually occurred. The governance problem therefore goes beyond whether logs exist.
It includes:
who retains them → who can export them → who can demand access → which records can be reconciled → which claims can be independently tested
When the acting stack is also the principal source of the evidence used to assess its own conduct, audit becomes dependent on that stack’s recordkeeping choices. This is where forensic readiness meets the existing idea of the algorithmic supply chain.
05 Human oversight also depends on the record
Human oversight is often represented by the presence of an approval step. But an approval button alone says little about whether oversight was meaningful. A reviewer needs enough information to understand the action being authorised and enough authority to stop, redirect, or escalate it.
That means oversight depends on the same runtime and evidentiary architecture as later audit.
If the interface hides credential scope, compresses an action sequence into a summary, omits earlier retries, or presents an automatically generated interpretation of what is about to happen, the human may technically remain “in the loop” while possessing very limited epistemic access to the actual operation.
Human presence acquires governance value through information and intervention power.
For long-running agents, this also makes organisational conditions relevant: expertise, workload, escalation routes, authority, and incentives affect whether a nominal reviewer can intervene in practice.
06 The governance direction
The paper develops a reconstruction-oriented approach to consequential agentic systems. The answer is not to preserve everything. Keeping every prompt, memory, command, credential, network event, and user interaction indefinitely would create a new surveillance and security problem.
The target is minimum sufficient evidence: enough retained material to reconstruct and test a consequential action, with retention and access proportionate to the risk.
In practice, this points toward:
- historical identities for materially different runtime configurations;
- execution provenance linked to the configuration that enabled the action;
- preserved context for consequential human approvals;
- links from dashboards and summaries back to recoverable lower-level events;
- independent corroboration outside the immediately acting harness;
- separation and protection of sensitive material;
- and clear retention, export, interoperability, and access responsibilities across the supply chain.
The broader problem is one of timing. Agent capability can be changed quickly through connectors, permissions, skills, credentials, and orchestration rules. The infrastructure needed to preserve, reconcile, and govern evidence usually changes more slowly.
Agentic AI therefore creates an auditability problem as much as a model-safety problem.
The full preprint develops this argument through contemporary agent architectures, security studies, digital-forensics research, execution-provenance work, accountability scholarship, and documented operational cases. Published as a public preprint on Zenodo, August 2026.