Blog
What I want from an always-on AI agent
I am treating an always-on agent as a bounded worker. It should accept a named job, leave an activity record and put any proposed change somewhere a person can review it.

An always-on agent should not wake up with a general instruction to be useful. I want it to receive a job with an identifier, objective, input references, permitted tools, time limit and completion condition. Those fields make the run inspectable and tell the worker when to stop. The job can be created by a person, a schedule or another system, but the contract should look similar. A request such as "review these open issues and draft a prioritised note" identifies a source set and an output. "Keep the project moving" does not. It invites the agent to decide which work matters and how much authority it has. Inputs should be references where possible, not large copied payloads. A file path, issue identifier or database key lets the run record preserve where information came from. The worker can then record exactly which version it read. I also want an explicit expiry time. Work can become irrelevant while it waits in a queue. A daily summary that starts two days late should fail as stale rather than arrive as if it were current. A background worker needs more than a list of prompts. Each job should move through named states such as queued, claimed, running, waiting for approval, completed, failed or cancelled. The transition history is useful when the final output never appears.
Claiming has to be atomic so two workers do not perform the same job. A lease can allow recovery if a worker disappears, but the lease duration must account for long tool calls. The worker should renew it while healthy. If another worker takes over, it should create a new run rather than pretending to continue the lost process. Retries need a limit and a reason. A network timeout may justify another attempt. A rejected permission request probably does not. I want retry policy attached to the error class, with increasing delay for transient failures and a dead-letter state for jobs that require inspection. The queue view should show what is waiting, what is currently claimed, which jobs are late and which error is repeating. These fields allow operators to inspect the worker without reconstructing its state from logs. The agent may know how to use many tools, but a particular job should expose only the ones it needs. A research job can have read access to named sources and a place to write a draft. It does not need deployment credentials. A maintenance job may run diagnostics while keeping restart or configuration changes behind approval. I prefer capabilities that are narrow in both action and target. "Read repository A" is safer and easier to review than a token that can read every repository the account can access. Where an integration cannot issue narrow credentials, the agent wrapper should enforce an allowlist and record the resolved target before executing.
Write tools should support dry runs, draft states or sandboxes where the external system allows them. The job contract can require the worker to stop after preparing a proposed change. A person then reviews the diff, query plan or message before a separate approved action applies it. Tool output also needs limits. A command that returns a huge log can overwhelm the current task. The wrapper should cap output, save the full artefact separately and give the agent a reference for targeted follow-up. Put a proposed code change or data update where its evidence can be inspected, rather than leaving it in a final chat response. For code, that may be a branch and diff. For configuration, it may be a patch. For database work, it may be a dry-run report with affected identifiers and counts. The review surface should show the original job, inputs used, proposed change, checks performed and unresolved concerns. It should not rely on the reviewer reconstructing the run from a transcript. Approval must bind to a specific artefact. If the branch changes after review, the approval should no longer apply. A commit hash, patch checksum or versioned change set can provide that binding. The execution step then verifies it is applying the reviewed version.
Some jobs can finish without approval because their outputs are read-only reports. That does not make their evidence unimportant. A report should still link to the source versions it used so a later reader can tell whether the conclusion is stale. For every run, I want timestamps, state changes, tool calls, approvals, output references and a final status. Keep timestamps, state changes and references as operational data. Sensitive prompts and tool responses can remain in their source systems behind controlled references. The record should distinguish an attempted action from a completed one. A tool returning success is only one event. If the job changes an external record, the worker should read the target back and attach that verification before marking the action complete. Classify failures more precisely than "Agent error". The record should say whether the failure came from invalid input, unavailable dependency, denied permission, expired job, failed verification or an internal exception. That classification determines whether a retry is sensible. I also want the exact instruction and tool configuration version used for the run. An agent's behaviour can change even when the named job stays the same. Version links make comparisons possible without stuffing configuration into every log entry.
An always-on process needs resource and authority limits that end work cleanly. I am setting maximum runtime, tool-call count, cost or token budget where available, and a ceiling on retries. Reaching a limit should produce a stopped state with partial artefacts, not a vague failure. The worker should pause when an input is ambiguous, a required source is missing, a proposed action exceeds scope or verification disagrees with the expected result. In those cases the useful output is a compact request for a decision, with the current evidence attached. There should also be a global way to stop new claims while allowing active runs to finish or be cancelled. That is needed for maintenance, credential incidents and broken tool behaviour. I do not want to log into a process and kill it blindly while it may be halfway through a write. Start with a read-only acceptance test: submit one job, watch it move through every expected state, confirm its source and output references, then restart the worker and verify the completed job is not claimed again.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.