All notes

Governance and security10 min read

What an AI agent should log: the tool call, not the conversation

A log that can say why an agent acted, who approved it and what it had read is an engineering artefact. The fields we think belong in each record, who may write it, what is redacted and how long it is kept, each with its source.

In this note
  1. Three readers want the log, for three different reasons
  2. The unit of the log is the tool call, not the conversation
  3. The fields we think belong in every record
  4. The agent must not be able to write to its own log
  5. Arguments are data, and data may be personal
  6. How long to keep it is a decision with a date on it
  7. What the log has to answer a month later
  8. What this is not

Suppose an agent changes the delivery date on a purchase order, and a month later someone asks why. Suppose too that what exists is a chat transcript, if it was kept, and a line in an application log saying that a function ran. Neither says who approved the change, which rule let it through, or what the agent had read just before it acted.

The unit of an agent’s log is the tool call, not the conversation. A log that can answer those three questions is an engineering artefact: it has a field list, a writer that is not the agent, a redaction rule and a retention period.

On the AI Systems page we commit to “a log of every call, kept in your account”, and the permissions note gave that log one paragraph. This is the long version. Where it goes further than those two pages it is our opinion, with a primary source beside each field. It describes no client’s system.

Three readers want the log, for three different reasons

The law wants traceability. As of October 2026, Article 12 of the EU AI Act (opens in a new tab) covers high-risk AI systems. They must “technically allow for the automatic recording of events (logs)” over their lifetime, so that risks can be spotted and the system monitored once it is in use. For one category, remote biometric identification, the article names fields: the start and end of each use, the reference database, the input that matched, and the people who verified the result.

The Commission’s timeline has these rules applying to the high-risk uses in Annex III from 2 December 2027. Whether a system is high-risk is a legal question, and nothing in this note is legal advice. Either way, the article says what a log is for.

Security wants evidence. The OWASP Top 10 for Agentic Applications (opens in a new tab) of December 2025 calls strong observability “non-negotiable”. Against tool misuse it asks for “immutable logs of all tool invocations and parameter changes”. The Model Context Protocol specification (opens in a new tab) says clients should “Log tool usage for audit purposes”. Microsoft, describing its control plane for MCP, asks whether an agent may invoke a tool “with these arguments, at this time”. A log records the answer.

The operator wants to read what happened. Anthropic’s guidance on evaluating agents defines a transcript as “the complete record of a trial” and ends its advice with “Read the transcripts!”. Our AI Systems page names “no record of what it did” among the reasons these systems fail.

One record has to serve all three readers, so its fields have to be written down.

The unit of the log is the tool call, not the conversation

A conversation is what the model said. A tool call is what the system did. The risk is in the second, so the record belongs there. The AI Systems page gives our method in one line: “The model proposes; the tool, its permissions and a person decide.” The log is the record of that decision.

The telemetry standard has the same shape. OpenTelemetry’s conventions for generative AI (opens in a new tab) define an execute_tool span for each tool call. Its own walkthrough shows where that span sits: under an invoke_agent span, beside a chat span for each model call. One span per call, in a tree that also holds the model turn that asked for it. The difference, as we see it, is that telemetry may be sampled and an audit record may not: every call is kept. Two more things follow.

Every attempt is a record, including the refused ones. Microsoft’s toolkit logs “every tool call attempt, policy decision, and execution outcome”. Claude Code, which we build with, can export telemetry with a decision event for every permission decision, accepted or rejected, and a result event only for calls that ran. We think the refusals are where an attack or a misunderstanding shows first.

The transcript is a different thing. It is large, mostly other people’s words, and needs a retention rule of its own. The call record should be readable without it.

The fields we think belong in every record

The first two columns are our opinion. The third names a source that asks for the field or already records it, as read on 3 October 2026. OpenTelemetry’s conventions are still marked Development, so their names may change.

Field What it records Who asks for it
Call id, run and parent A unique id, the run it belongs to, the model turn that asked for it, and its position in the run OpenTelemetry: gen_ai.tool.call.id, gen_ai.conversation.id and the span tree. OWASP: watch tool-chaining patterns, which needs the order.
Times When the call was requested, decided and finished Article 12(3), for remote biometric identification: the start and end of each use. OWASP: “tamper-evident, time-stamped logs”.
Tool, server, version The tool’s full name, the server that provides it, the version of each OpenTelemetry requires gen_ai.tool.name. OWASP: fully qualified tool names and version pins.
Agent and credential Which agent made the call, under which credential and scope OWASP: logs bound to agent identities. OpenTelemetry: gen_ai.agent.name.
Arguments What was sent: verbatim, hashed or left out by a written rule, with its size OWASP: all invocations and parameter changes. Claude Code records tool_input_size_bytes even when it withholds the arguments.
Decision and rule Allowed, asked or refused, and the rule or hook that decided Claude Code’s decision event: decision and source. Microsoft’s toolkit logs every policy decision.
Approver The named person who approved, when, and whether they saw the full arguments MCP: show tool inputs before a call. Article 12(3) again: the people who verify a result.
Outcome Success or error, the class of error, the result’s shape and size, the duration MCP: a failed tool returns isError: true. OpenTelemetry: error.type on a failed span.
What was in force The model, prompt version and permissions map version at the time OpenTelemetry: gen_ai.request.model. Article 12(2): logs that help spot a substantial modification.

The last row is the easy one to leave out. Without it, a record from March is read against October’s rules.

The agent must not be able to write to its own log

The permissions note said this in one clause: the calls go into “a log the agent cannot write to”. A record the agent can edit is a claim, not evidence. OWASP’s list returns to the point under four of its ten risks, in the words immutable, tamper-evident, tamper-proof and signed. We think it comes down to three decisions.

The writer is the layer that runs the call. The host application sees the request, the decision and the result, so it writes the record. Claude Code shows the pattern: its hooks run before and after a tool call, in code the harness runs, and receive the tool’s name, its input and a call id. A summary the model writes of its own work is useful. It is not a log.

The store is append-only, and the agent holds no credential for it. No tool on the agent’s list writes, edits or deletes a record.

Tampering is detectable. OWASP’s logging guidance says to “Build in tamper detection”. Microsoft’s toolkit keeps “Append-only, hash-chained audit logs”: each record carries a hash of the one before it, so a changed or missing record breaks the chain.

The service page adds where the log lives: in the client’s account.

Arguments are data, and data may be personal

Security wants the arguments. Privacy wants as little as possible. The log has to satisfy both.

Arguments are where personal data enters the record: an email address, a name, the body of a message. OpenTelemetry makes tool arguments and results opt-in, each with the warning that it “may contain sensitive information”.

As of October 2026, where the GDPR (opens in a new tab) applies, personal data must be “limited to what is necessary” for its purpose, and that includes personal data in a log. OWASP’s logging guidance names what should usually not be recorded directly, access tokens, passwords and sensitive personal data among them. Such values are to be “removed, masked, sanitized, hashed, or encrypted”.

So the rule we think belongs in the permissions map, for each argument of each tool, has three classes.

  1. Stored verbatim. Identifiers and values needed to reconstruct the action: the order number, the amount, the file path, the destination.
  2. Stored as a keyed hash. Values that must be matched but not read, such as an email address. Under the GDPR a pseudonymised value is still personal data for whoever holds the means to attribute it, so hashing narrows the exposure and does not end the duty.
  3. Never stored. Secrets of any kind, and free text such as the body of a document. The record keeps a size, a content hash and a reference to the original. OpenTelemetry recommends a similar pattern for model inputs and outputs in production: external storage, under separate access controls.

Two cautions. Redaction happens before the record is written, not when it is displayed: a value masked on screen and present in storage is the failure in the redaction note. And the approver still sees the full arguments; the record says so.

Reading the log is an event too. The same OWASP guidance: “All access to the logs must be recorded and monitored”.

How long to keep it is a decision with a date on it

As of October 2026, the EU AI Act gives a number. Its Articles 19 and 26 require providers and deployers of high-risk systems to keep the logs under their control for a period appropriate to the system’s purpose, “of at least six months”. Both allow other law to say otherwise, and both name data protection law in particular. The GDPR’s storage limitation principle runs the other way: personal data is kept in identifiable form no longer than its purpose needs.

For any other system, absent a sector rule, the period is a design decision, and we think it belongs in the written scope with its reason. We would split it by class: the skeleton of each record (who, what, when, which rule) kept longest, hashed values for less time, referenced content for the shortest. Six months is a fair place to start, because it is the figure the Act itself uses. It is not a recommendation for any particular system.

Append-only does not mean for ever. Records should expire in whole batches, in a job the agent has no part in, and the expiry is logged. OWASP’s guidance says log data “must not be kept beyond” its retention period.

What the log has to answer a month later

Three questions, and the fields that answer each.

  • Why did it do that? The decision and its rule, read against the versions in force.
  • Who approved it? The approver, the time, and whether they saw the arguments.
  • What did it read first? The earlier records of the same run, in order.

The third is the one we think an investigation starts with. An agent that was talked into something was talked into it by something it read.

A synthetic example follows: every name and value is invented. First the run, one line per call:

run_2291, 1 September 2026

seq  tool                         decision                         outcome
5    mail.read                    allow  (rule: reads-in-scope)    ok, 1 message
6    documents.read               allow  (rule: reads-in-scope)    ok, 1 attachment
7    orders.update_delivery_date  ask    (rule: order-writes-ask)  approved 10:15:52, ok
8    mail.send                    refuse (rule: send-denied)       not run

Then the full record of call 7:

{
  "call_id": "call_0193",
  "run_id": "run_2291",
  "sequence": 7,
  "parent_turn": "turn_04",
  "requested_at": "2026-09-01T10:14:07Z",
  "tool": "orders.update_delivery_date",
  "server": "erp-connector",
  "server_version": "1.4.2",
  "agent": "order-desk-agent",
  "credential": "svc-order-desk",
  "scope": "orders:write",
  "arguments": {
    "order_id": "PO-4471",
    "new_date": "2026-09-19",
    "requested_by": "hmac:9f2c41d0"
  },
  "arguments_bytes": 142,
  "decision": "ask",
  "rule": "order-writes-ask",
  "approver": "order-desk-lead",
  "approver_saw": "full arguments",
  "decided_at": "2026-09-01T10:15:52Z",
  "finished_at": "2026-09-01T10:15:53Z",
  "outcome": "ok",
  "is_error": false,
  "result_shape": { "updated": 1 },
  "result_bytes": 31,
  "duration_ms": 412,
  "in_force": {
    "model": "the model id the API returned",
    "prompt": "v14",
    "permissions_map": "v6"
  },
  "previous_hash": "b71e09aa",
  "hash": "4c0d82f3"
}

The run answers all three questions without the transcript. The order desk lead approved the change at 10:15, under a rule that makes every order write ask. Before proposing it the agent had read one message and one attachment. Then it tried to send an email and was refused, and only a log that records refusals shows that.

Rateslip applies the same discipline to documents: it keeps the region of the page every figure was read from, so a number can be traced to where it was printed. A tool-call log does that for an action.

What this is not

It is not compliance with the EU AI Act, the GDPR or anything else. The service page says so: “We document the controls; we do not certify compliance with any regulation.” A log like this is a documented control: written down, in the code, and readable by the client. Whether it meets a legal duty is a question for the client’s own adviser.

It is not a permission system. A log records what was allowed and prevents nothing. The permissions map does the preventing, and the log is how anyone checks that the map was followed.

It is not finished when it is switched on. The fields fit on a page. The work is deciding, for each tool and each argument, what goes into them before the first call, and then reading the records.

Sources

  1. Regulation (EU) 2024/1689 (AI Act), Article 12: Record-keeping, consolidated text as at 27 July 2026 (opens in a new tab)European Commission, AI Act Service Desk. Accessed 3 Oct 2026
  2. Regulation (EU) 2024/1689 (AI Act), Article 19: Automatically generated logs (opens in a new tab)European Commission, AI Act Service Desk. Accessed 3 Oct 2026
  3. Regulation (EU) 2024/1689 (AI Act), Article 26: Obligations of deployers of high-risk AI systems, paragraphs 5 and 6 (opens in a new tab)European Commission, AI Act Service Desk. Accessed 3 Oct 2026
  4. Timeline for the Implementation of the EU AI Act (opens in a new tab)European Commission, AI Act Service Desk. Accessed 3 Oct 2026
  5. Regulation (EU) 2016/679 (General Data Protection Regulation), recital 26, Article 4(1) and 4(5), Article 5(1)(c) and 5(1)(e) (opens in a new tab)EUR-Lex, Publications Office of the European Union. Accessed 3 Oct 2026
  6. OWASP Top 10 for Agentic Applications 2026 (December 2025), introduction and the mitigations under ASI02, ASI08, ASI09 and ASI10 (opens in a new tab)OWASP Gen AI Security Project. Accessed 3 Oct 2026
  7. Logging Cheat Sheet (“Data to exclude”, “Protection”, “Disposal of logs”) (opens in a new tab)OWASP Cheat Sheet Series. Accessed 3 Oct 2026
  8. Model Context Protocol specification, 2026-07-28, Server features: Tools (“Error Handling”, “Security Considerations”) (opens in a new tab)Model Context Protocol (modelcontextprotocol.io). Accessed 3 Oct 2026
  9. Semantic conventions for generative client AI spans, status Development (“Inference”, “Execute tool span”, “Capturing instructions, inputs, and outputs”) (opens in a new tab)OpenTelemetry. Accessed 3 Oct 2026
  10. Inside the LLM Call: GenAI Observability with OpenTelemetry (14 May 2026) (opens in a new tab)OpenTelemetry. Accessed 3 Oct 2026
  11. Securing MCP: A Control Plane for Agent Tool Execution (22 April 2026) (opens in a new tab)Microsoft for Developers. Accessed 3 Oct 2026
  12. Claude Code documentation, Monitoring (the tool decision and tool result events) (opens in a new tab)Anthropic. Accessed 3 Oct 2026
  13. Claude Code documentation, Hooks reference (PreToolUse and PostToolUse input) (opens in a new tab)Anthropic. Accessed 3 Oct 2026
  14. Demystifying evals for AI agents (9 January 2026) (opens in a new tab)Anthropic. Accessed 3 Oct 2026

Amiya BeraFounder, Kordal Systems

Published . Governance and security.

More notes

All notes
  1. Build notes10 min read

    A black rectangle is not a redaction

    Drawing a box over text leaves the text in the file. What a PDF still carries after the box, why Redact PDF rebuilds a marked page as an image, and how it checks its own export before you send it.

  2. Field notes9 min read

    AI-generated software is a prototype until someone owns the failure modes

    Generating a working application has become cheap. What still costs money is knowing how it fails, who is told when it does, and what happens next: the work between a convincing demo and a system a business can run on.