18 December 2025 / Agent operations

The agent’s activity log is part of the deliverable

When an agent changes code, data or configuration, I need the work record alongside its final answer. The next person should be able to see what was attempted, what passed and where the system was left.

A hand drawing in a spiral notebook at a marble-topped table.
Photo: Kaizen Nguyễn (opens in a new tab)

An agent can finish with a short message saying the change is complete. That message does not show which files it read, which command failed on the first attempt, whether it changed course or whether the final verification inspected the actual target. Those details matter when somebody reviews the work or returns to it later. A full chat transcript is not a good substitute. It mixes reasoning, source content and tool output into one long record. It can expose private data while still failing to identify the exact artefact that was changed. Store the important events as structured records. I want the activity log delivered with the result, much like tests and a diff accompany a code change. The log should be concise enough to inspect and detailed enough to reconstruct the operational path without replaying the whole conversation. The basic unit is a run: one attempt to complete one named job. A retry creates a new run linked to the same job. This avoids merging separate attempts into a timeline that implies uninterrupted progress. Each event should include a timestamp, run identifier, event type, actor or component, target reference and outcome. Depending on the job, useful event types include input resolved, file read, tool called, change proposed, approval received, change applied, check run and target verified. Sequence numbers help when events arrive from several services or share a timestamp. I also want the instruction version, tool configuration version and code revision attached to the run. They can be references rather than copies, as long as the referenced material is retained.

The event format should be append-only during normal operation. Corrections can add a new event that points to the mistaken record. Editing old entries makes it harder to explain what a worker believed at the time. A command exit code or successful API response does not prove the intended state now exists. The log should separate the request, provider response and independent verification. For a file change, that may look like an edit event followed by syntax checks and a diff inspection. For an external record, the worker can read the exact target back and compare the fields that were meant to change. A database update may need affected-row counts plus a query that checks the destination conditions. This separation is useful when a tool reports success but acts on the wrong environment, when a write is accepted asynchronously or when another process changes the target immediately afterwards. The final run status should depend on the verification event, not merely on the attempted action. When verification fails, the agent should stop, preserve evidence and avoid further compensating changes unless the job allows them. "Tests passed" is too compressed for a handover. I want each check identified by its command or procedure, result and relevant output reference. If the run skipped an expected check, the log should state why.

For code changes, checks may include targeted tests, type checks, formatting, a build or a browser acceptance step. Running every possible check is not always practical, so the activity record should distinguish what passed from what was not run. That lets the next person judge the remaining risk. Data and configuration changes need their own acceptance checks. A dry run, query plan, schema validation, service reload and health check each answer different questions. Their evidence should be attached near the change they validate rather than collected in a vague closing paragraph. Large outputs can live as separate artefacts with checksums or stable references. The event itself only needs the command, exit status, duration and a useful summary. Secrets and private payloads should be redacted before storage, not after they spread through logs. The next person needs to know where the system was left. For a repository, that includes the branch, commit or uncommitted files, generated artefacts and any process still running. For a data job, it includes the last accepted checkpoint, rejected records and whether retries remain queued. For infrastructure, it includes service state and any temporary rollback option. I want a final state snapshot generated from the environment rather than written from memory. Commands such as repository status, service status or a targeted database query can provide that snapshot. The log should include the exact target so a check against a staging service is not mistaken for production evidence.

Temporary work deserves attention. Agents often create scratch files, branches, containers or test records. The run should either remove them or list them with a reason and owner. Silent leftovers become confusing inputs to later work. If the job stops early, the same snapshot is still useful. An incomplete run can be handed over cleanly when it states what changed before failure and what has not yet been attempted. Prompts, retrieved documents and raw tool output can contain personal data, client material, tokens or internal paths. Copying all of it into the activity log increases the exposure. The durable activity record should keep operational facts and controlled references while leaving sensitive content in the system that already governs it. Access should match the underlying job. A person allowed to review run timing may not be allowed to open every source document. References can preserve that boundary by resolving through the source system's existing permissions. Retention also needs a rule. Security investigation, operational debugging and project handover may require different periods. Keeping every detailed payload indefinitely increases risk and makes the log harder to search. A compact event record can last longer than verbose diagnostic output.

Redaction itself should be tested. Common secret formats, authentication headers and sensitive fields can be filtered at the tool boundary. If raw output must be retained for a short investigation, it should use a restricted store with an expiry rather than the ordinary activity stream. The human-facing report can be generated from the structured events. It should name the job and run, changed artefacts, checks and outcomes, verification result, unresolved items and final state. Links can open the diff, check output or source record when the reviewer needs detail. The report should not imply certainty the events do not support. If a build passed but no browser check ran, say exactly that. If a write was attempted and the read-back failed, the run remains unverified even if the proposed payload looked correct. My immediate test for the activity log is a small controlled change. I want to compare the structured report with the repository and tool state after the run, then hand it to someone who did not watch the work. Any question they cannot answer without opening the transcript points to an event or final-state field that is still missing.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs