Auditing Agent Actions: Who Did What, On Whose Behalf?

The incident review started well. Someone had pulled 4,200 customer records out of the analytics API in eleven minutes on a Tuesday afternoon, and we had beautiful logs — every read, timestamped to the millisecond, with the record ID, the fields returned, the source IP, and the actor.

The incident review started well. Someone had pulled 4,200 customer records out of the analytics API in eleven minutes on a Tuesday afternoon, and we had beautiful logs — every read, timestamped to the millisecond, with the record ID, the fields returned, the source IP, and the actor.

The actor, on all 4,200 rows, was svc-warehouse-query.

That is a service. It runs in our account, it has read access to the warehouse because that is its entire job, and it made every one of those calls correctly, with a valid token, well within its permissions. The log was not wrong. It was complete, tamper-evident, and retained for seven years. It answered what happened with total precision.

Then counsel asked the only question that mattered — "whose data was this, and who asked for it?" — and we spent nine hours reconstructing the answer from application logs, a Kafka topic with four days of retention, and a Slack thread. The audit system contributed nothing. It had one field for "who," and the thing it named was the last hop.

The request had originated with a support engineer asking an assistant to "check whether the churn report looks right." The assistant called an MCP server. The MCP server called the warehouse service. Three hops, three identities, and our schema had room for one.

"Who" now has three answers, and they are not interchangeable

It is worth being pedantic about the vocabulary here, because the vocabulary is what keeps the design honest.

Every audit schema I have seen has a single actor column, because it was designed when there was exactly one candidate. A human clicked a button; the human is the actor. Service accounts arrived later and got poured into the same column — already a small lie, but a survivable one, since they mostly did scheduled work with no human behind them.

An agent chain breaks the column properly, because there are at least three distinct answers and an investigation needs all of them:

The subject. The human on whose behalf the action ultimately happened. The support engineer. This is the party whose permissions should have bounded the action and whose name appears in the breach notification if this goes badly. In OAuth terms it is sub, and it must be the original subject, not whoever the last hop happened to be authenticating as.

The actor. The component that decided to take this specific action. The assistant. Nobody told it to read 4,200 records; it read a report, formed a plan, and issued calls. This is what you are investigating when you ask "why did this happen at all," and it is completely absent from most logs.

The executor. The component that performed the call and, usually, wrote the audit record. svc-warehouse-query. This is the only one most schemas capture, and the least interesting of the three, because it did what it was told and its identity tells you nothing about intent or authority.

Collapse these into one field and you do not get a lossy record — you get a misleading one. "Actor: svc-warehouse-query" is an affirmative claim that a service did this on its own behalf, which is false. In a regulatory setting a confidently wrong attribution is worse than a gap, because a gap invites investigation and a wrong answer closes it.

There is a fourth thing an investigator wants — the instruction that produced the decision — and it needs separate treatment, because it is the one you should be careful about storing. I will come back to it.

The act chain, read as evidence rather than as protocol

RFC 8693 already gives you the structure. I am not going to re-teach the grant here; what matters for forensics is what the resulting token asserts, and how to read it correctly under pressure.

{
  "iss": "https://auth.example.com/tenant-42",
  "sub": "[email protected]",
  "aud": "https://warehouse.internal",
  "scope": "warehouse.read",
  "jti": "8b1c2f...",
  "txn": "req_01J9Z8Q3",
  "exp": 1754400300,
  "act": {
    "sub": "mcp-analytics",
    "act": {
      "sub": "agent-support-assistant"
    }
  }
}

Three properties carry the forensic weight.

sub is the support engineer and stays the support engineer at every hop, forever. It never becomes the agent, and it never becomes the service. If you find a token in your estate where sub changed identity mid-chain, you have impersonation, not delegation, and attribution is already gone regardless of how good your logging is.

The act chain is nested, with the most recent actor outermost. Read the payload above as: this is the engineer's request, currently carried by mcp-analytics, which received it from agent-support-assistant. Each exchange wraps the previous chain rather than appending to a list. Everyone gets this backwards at least once, and getting it backwards in an audit pipeline is uniquely nasty — you end up with a log that is confidently wrong about the order of causation, which is exactly what an investigator will build a timeline on.

may_act is the other half of the picture, and it is about permission rather than history: a claim stating who is allowed to act for that subject. It is rarely populated, because it requires knowing the delegation graph at issuance time, but when present it is forensically valuable in a way people miss — it records what was permitted, which is a different question from what occurred, and reconciling the two is most of an investigation.

The critical operational point: a token is a transient artifact and your audit record is not. The chain exists for a few minutes. If your resource server validates the token, authorizes on it, then writes actor = <whatever the auth middleware put in the principal object>, you have thrown away the evidence at the exact moment you held it. That middleware almost certainly surfaces sub, and sub alone is the most misleading single field available — it looks like a correct human attribution while omitting that a machine made the decision.

What a usable audit record actually carries

flowchart TB
    U["Engineer<br/>eng-4471"] --> A["Agent<br/>support-assistant"]
    A --> M["MCP server<br/>mcp-analytics"]
    M --> W["Warehouse service<br/>svc-warehouse-query"]

    A -.-> RA["records: sub=eng-4471<br/>act=[assistant]<br/>txn, task, tool chosen<br/>token jti issued"]
    M -.-> RM["records: sub=eng-4471<br/>act=[mcp, assistant]<br/>txn, tool + arguments<br/>authz outcome, scopes"]
    W -.-> RW["records: sub=eng-4471<br/>act=[warehouse, mcp, assistant]<br/>txn, rows returned<br/>scope exercised, jti presented"]

Note what is constant down the right-hand column and what changes. sub and txn never move; the act chain grows; the authorization basis is specific to each hop. That invariance is what makes the records joinable at 2am.

Five things belong in every record:

The full chain, not the last hop. Store the ordered actor list as a first-class field — subject plus every actor, in causal order. Not a JSON blob dumped for completeness; a field you can query. If you store only one derived value, store the first actor, the originating agent, because that is the one nobody has and everybody eventually needs.

A correlation identifier that survives hops and organizational boundaries. Teams reflexively reach for the distributed-tracing ID. It is the wrong instrument here: trace headers get dropped at trust boundaries, and a third-party MCP server has no obligation to propagate your header conventions. Put the correlation ID in the token instead. txn is a registered JWT claim — RFC 8417 defines it as a transaction identifier — and if your authorization server mints it at the first exchange and copies it through every subsequent one, it arrives at hop four whether or not anyone in between cooperated. Very few deployments do this and it is close to free.

The authorization basis, not just the outcome. "Access granted" is not evidence. Granted on the strength of what? Record the jti of the token presented, the scopes actually exercised (the ones this call consumed, not every scope in the token), and which delegation permitted the exchange. That is the difference between "the system allowed it" and "the system allowed it because policy X says agents of class Y may act for users in tenant Z with scope warehouse.read." Only the second is auditable.

The task. An agent action is unintelligible without the unit of work it belonged to — not the prompt, but an identifier for the job the agent was executing, with the tool name and arguments at each call. Without it an investigator sees 4,200 independent reads. With it they see one task that fanned out, which is a different finding and takes ten minutes rather than nine hours.

The decision and the enforcement, separately. The model decided to call read_customer. The authorization layer decided whether it could. Two facts, they frequently disagree, and the disagreement is the highest-value signal in the system. Most pipelines log successes at audit grade and denials at debug grade, which is precisely inverted for agents: one agent generating authorization denials across many different subjects in a short window is the clearest available signature of prompt injection, and it is invisible if denials are sampled into a 14-day log store.

The prompt provenance question, and where I would draw the line

Sooner or later the investigation reaches "what instruction caused the agent to do this," and someone proposes logging prompts into the audit trail.

Don't. An audit store is compliance-scoped, append-only, broadly readable by security and legal, and retained for years. Prompt content is customer data — frequently the most sensitive in the system — and in an injection scenario it is attacker-controlled content you are now permanently storing in your most privileged log. You have created a retention obligation you did not analyze and a stored-XSS surface aimed at your own SIEM console. The equally common failure is logging nothing, which leaves you with a chain of actors and no account of why any of it happened.

The position I would defend: the audit record carries a reference, and the content lives in a separately-governed store with its own retention, access control, and deletion path. Store a content hash and an interaction ID in the audit event. The hash gives you the property that actually matters — you can prove later that a specific stored interaction is the one that produced this action, and that it has not been altered — without the content crossing into the compliance boundary. The interaction store is then governed like the customer data it is: shorter retention, tighter access, subject to deletion requests, encrypted with keys the audit pipeline does not hold.

The trade-off is real and worth stating plainly: if the interaction store's retention is 90 days and your investigation starts on day 100, you have a hash that proves nothing. That is a policy decision someone should make deliberately, with the privacy and forensic costs both on the table, rather than discovering it during an incident.

The query nobody's schema supports

Here is the part that determines whether any of this pays off.

Audit systems are indexed for the queries that funded them, which are compliance queries: everything that happened to this resource, or everything this user did. Both assume the interesting axis is the resource or the human.

Incident investigation with agents needs two queries, and they run in opposite directions:

"Everything done on behalf of this subject, by any actor, in this window." The breach-notification query — the one that determines whose data was touched. It needs sub indexed, which you probably have, and it needs records written by hop three to still carry the original sub, which is the whole point of the previous sections.

"Everything this agent did, across all subjects." The containment query. An agent is suspected of being injected, compromised, or simply buggy, and you need its blast radius now: which users, which resources, which tenants. Nobody builds this index, because until agents arrived there was no reason to — a service account's activity was uninteresting by construction, and a human's activity was already the primary key.

The second is where the schema cost lands. Indexing the actor chain means indexing an array of high-cardinality values, and the cheap version — indexing only the outermost actor — answers the wrong question, because the outermost actor is the executor service. You want the innermost actor, and probably every intermediate one too. That is a multi-valued high-cardinality index on your highest-volume table.

The Hidden Cost of Audit Logs argues that every searchable high-cardinality field is a permanent multiplier on your bill, and that stands. The mitigation specific to this case is that agent-initiated actions are a classifiable subset: records carrying a non-empty act chain are a small fraction of volume today and deserve a different indexing tier. Split the index rather than widening it — the containment query needs to be fast, the token-refresh records do not.

Four failure modes worth hunting for

Chain truncation by re-authentication. An intermediary that does not exchange the token it received, and instead calls downstream with its own client credentials, silently severs the chain. Everything past that hop is attributed to a service, permanently. This is not hypothetical — it is the default behaviour of every service written before anyone thought about delegation, and "stop forwarding my own credentials" is on nobody's roadmap.

Detection is easier than the fix and worth wiring up first. At the authorization server, look for clients that obtain tokens by both token_exchange and client_credentials: a mixed grant profile means some code path is bypassing delegation. At the resource server, alert when a request arrives with a sub resolving to a machine principal rather than a human, from a client that participates in delegation chains elsewhere. This is invisible in most schemas precisely because a service sub and a user sub land in the same column and look equally valid.

Clock skew making ordering unreliable. Timelines get built by sorting on timestamps, and hops sit on different hosts with different drift. Forty seconds of skew is enough to make the effect appear before the cause, and an investigator who does not know that will build a theory on it. Never order a chain by wall clock when causal structure is available — the act nesting depth gives ordering no clock can corrupt — and record a per-hop sequence number within the txn. Exchanges are serialized through the authorization server, which makes its issuance sequence the only reliable ordering authority in the architecture.

Retries producing duplicates with no idempotency key. Agents retry, and loops retry harder. An investigator counting three send_email records cannot tell whether three emails went out or one went out and two attempts failed after the side effect — and that difference is a customer-notification decision. Idempotency keys have to be minted at the tool-invocation layer and carried into the audit record; there is no retrofit, because the information exists nowhere else. Most likely to bite you, least likely to be on anyone's plan.

Recording the token's subject instead of the delegation. The quiet one. Middleware hands the application a principal object with a user ID on it; the application logs that user ID; the record now says a human did something a machine decided to do. It passes review, it looks correct, and it is the failure behind the incident at the top of this article.

What audit can and cannot prove

Two limits worth stating before anyone builds a compliance narrative on top of this.

An audit log proves what your system was told and what it enforced. It does not prove what a model intended. You can demonstrate that a token carrying this chain was presented, that policy evaluated to permit, and that the action executed. You cannot demonstrate why the model chose it, and no amount of logging changes that, because the reasoning is not a fact your system holds — it is at best a post-hoc explanation generated by the component under investigation. Treat model-produced rationales as interesting, never as evidence, and never store them somewhere that implies otherwise.

If the same service both acts and writes the log, you have attribution but not independent evidence. A compromised or buggy executor writes whatever it writes. This is where the authorization server earns its keep as a forensic component: it is a different system, it issued every token in the chain, and its exchange records — which token, to which client, on which subject token, with which scopes — reconstruct the chain independently of anyone downstream. Almost nobody retains those at audit grade. They are low volume, they are the spine of every agent investigation, and they are written by a party that did not perform the action. If you do one thing from this article, retain them.

A note on where the model helps

Disclosure: I work on ClavionX, and its design here is a convenient illustration rather than a recommendation — it is a design-phase model, not a shipped and battle-tested one.

The choice with the most direct audit consequence is a small one. An MCP server is registered as a resource server with kind=MCP, governed by an MCP policy, and that policy carries requiresDelegatedUser, defaulting to true — it refuses tokens representing nothing but a machine acting for itself. As an authorization control that is unremarkable. As an audit control it is more interesting: it makes the un-investigable case structurally impossible at that boundary, because a call arriving with no human at the root of the chain is never served, so no record can exist naming only a service. The whole thing composes because there is no agent special-casing in the runtime — delegation is plain RFC 8693, sub stays the original user and actors accumulate in act, and an Agent is a first-class registered object realized as an ordinary OAuth2 client. The audit chain is not a feature; it is a side effect of not inventing a parallel mechanism.

The uncomfortable summary

Most audit systems answer the question "what happened to this resource" and were never asked to answer "who caused this, through what, on whose authority." Those were the same question for thirty years, and they stopped being the same question roughly the moment something in your architecture started deciding what to call next.

The fix is not more logging. It is a schema change — the actor field becomes a chain, the correlation identifier moves into the token, the authorization basis gets recorded alongside the outcome — and an index nobody has budgeted for.

The test is cheap enough to run this week. Take one action an agent performed yesterday. Ask your audit system who the human behind it was, and what task it was part of.

If the answer is a service account, you do not have an audit trail for agents. You have a very expensive record of your own infrastructure talking to itself.