An AI agent audit trail records, for every outcome, what came in, what the agent did and why, which model and version it used, what people did, and what came out, in storage that shows if anything was changed later. Design it around the outcome, not the system event, so an auditor can follow one piece of work from start to finish.
Governance is the main thing slowing agentic AI down. In Omdia's polling of MSPs and IT decision-makers, 47% named governance and compliance as the top barrier, far ahead of technical skills at 16%. A good audit trail is the most practical answer to that concern.
Step 1: Decide the unit of record
Make the unit the business outcome: one certificate, one invoice, one policy check. Everything the agent and people did for that outcome hangs off one record ID. System logs organized by server or API call are useful for debugging but hard to audit.
Step 2: Capture inputs as received
Store the documents, emails and system records the agent started from, exactly as they arrived, with a timestamp and where they came from. If an input is later corrected, keep both versions.
Step 3: Record every step and decision
For each step, record what the agent did, the rules or checks it applied and the result: "compared 42 coverage lines, 4 discrepancies". Record confidence checks and where low-confidence items were routed.
Step 4: Record the model and version
Record the model, version, prompt version and tool versions used at each step. When a model is updated, you need to know which outcomes it produced.
Step 5: Record human actions the same way
Exception desk actions and client approvals belong in the same record, with a name, time and what was decided. Regulators and auditors want to see where human oversight was applied, not just that it exists in a policy.
Step 6: Make the log tamper-evident
Use append-only storage where each record carries a hash of the one before it. Corrections are added as new records that point to the original. Any change to past records then becomes detectable.
Step 7: Make it usable by the people who check it
An audit trail no one can read does not help. Link each invoice line to its record. Let clients export records for a period, a process or a single item. Keep the language plain.
Frameworks to align with
The NIST AI Risk Management Framework is a voluntary US framework widely used to structure AI governance. In insurance, about half of US states have adopted the NAIC model bulletin on insurers' use of AI, which expects written governance, risk management and internal controls.
This is how the Evidence File works at Agentic MSP: one record per outcome, hash-chained, with every agent and human action included, and linked to the invoice.