Governance

How to build an audit trail for AI agents.

What to record for every AI agent action so outcomes can be audited: inputs, steps, model versions, human approvals and tamper-evident storage.

An AI agent audit trail records, for every outcome, what came in, what the agent did and why, which model and version it used, what people did, and what came out, in storage that shows if anything was changed later. Design it around the outcome, not the system event, so an auditor can follow one piece of work from start to finish.

Governance is the main thing slowing agentic AI down. In Omdia's polling of MSPs and IT decision-makers, 47% named governance and compliance as the top barrier, far ahead of technical skills at 16%. A good audit trail is the most practical answer to that concern.

Step 1: Decide the unit of record

Make the unit the business outcome: one certificate, one invoice, one policy check. Everything the agent and people did for that outcome hangs off one record ID. System logs organized by server or API call are useful for debugging but hard to audit.

Step 2: Capture inputs as received

Store the documents, emails and system records the agent started from, exactly as they arrived, with a timestamp and where they came from. If an input is later corrected, keep both versions.

Step 3: Record every step and decision

For each step, record what the agent did, the rules or checks it applied and the result: "compared 42 coverage lines, 4 discrepancies". Record confidence checks and where low-confidence items were routed.

Step 4: Record the model and version

Record the model, version, prompt version and tool versions used at each step. When a model is updated, you need to know which outcomes it produced.

Step 5: Record human actions the same way

Exception desk actions and client approvals belong in the same record, with a name, time and what was decided. Regulators and auditors want to see where human oversight was applied, not just that it exists in a policy.

Step 6: Make the log tamper-evident

Use append-only storage where each record carries a hash of the one before it. Corrections are added as new records that point to the original. Any change to past records then becomes detectable.

Step 7: Make it usable by the people who check it

An audit trail no one can read does not help. Link each invoice line to its record. Let clients export records for a period, a process or a single item. Keep the language plain.

Frameworks to align with

The NIST AI Risk Management Framework is a voluntary US framework widely used to structure AI governance. In insurance, about half of US states have adopted the NAIC model bulletin on insurers' use of AI, which expects written governance, risk management and internal controls.

This is how the Evidence File works at Agentic MSP: one record per outcome, hash-chained, with every agent and human action included, and linked to the invoice.

Common questions.

What should an AI agent audit trail contain?

For each outcome: the inputs as received, each step the agent took and the rules it applied, the model and version used, confidence checks and routing, any human actions with names and times, and the final result.

Is logging enough for AI governance?

Logs are necessary but not sufficient. Governance also needs clear ownership, testing before deployment, monitoring after, and a way to act on what monitoring finds. The audit trail is the evidence that those things happened.

Do regulators expect AI audit trails?

In insurance, about half of US states have adopted the NAIC model bulletin on insurers' use of AI, which expects documented governance, risk management and internal controls. The bulletin applies to insurers rather than agencies, but it sets the direction.

Sources

Want this run for you.

Agentic MSP runs back-office work with governed AI agents and bills only for verified outcomes.