AI Agents blogs
Featured blogs
Generated image: Villar David editorial
Blogs
AI Agents
Posts filed under AI Agents. View every category.
2026

What stays yours when the agent harness is managed
OpenAI's Agents API now provides hosted sessions, environments, context management, tool loading and subagents. Product teams still have to define the work, authority, evidence and recovery path around that managed harness.

Where agent supervision has to happen
Anthropic's recent analysis suggests experienced users approve more automation while also interrupting agents more often. That points to supervision built around visible plans and consequential boundaries instead of a confirmation attached to every command.

The coding-agent interface is becoming a control plane
Parallel agents, isolated worktrees and queued review change the developer’s role from typing every change to supervising work in motion. The difficult part is setting boundaries, spotting drift early and keeping enough evidence to understand each result before integration.

Production agents need institutional context
OpenAI's internal data-agent write-up shows how much context sits around the model: schemas, code, permissions and previous corrections. The useful design question is who maintains that context and how a user can correct it.
2025

The agent’s activity log is part of the deliverable
When an agent changes code, data or configuration, I need the work record alongside its final answer. The next person should be able to see what was attempted, what passed and where the system was left.

What I want from an always-on AI agent
I am treating an always-on agent as a bounded worker. It should accept a named job, leave an activity record and put any proposed change somewhere a person can review it.

Operating local AI infrastructure after the migration
A successful migration is followed by less glamorous work: services starting in the right order, data paths staying stable, backups being restorable and failures being visible. I am treating those operating details as part of the AI product because continuity disappears when the infrastructure is opaque.

What has to survive when an AI assistant moves machines
Before moving an AI collaborator to a new host, I am listing the parts that have to survive and the tests that will prove they did. The prompt is easy to copy; the harder state sits in records, tools and the links back to source material.

Trying to remember everything made the agent worse
A total-recall experiment has made the problem clear: more retrieved history can crowd out the task in front of the agent. I am separating identity, durable decisions, project state and searchable archives so memory is selected by purpose instead of poured into every session.

What I am testing in long-term AI memory
I am testing whether a graph can recover a decision through its relationships to a person, project and source. The immediate question is which records improve the next session and which simply make retrieval noisier.

Writing work orders for coding agents
Coding agents can now accept an issue and return a pull request, which makes the issue itself part of the engineering system. A useful work order needs a bounded objective, relevant context, acceptance criteria, permitted scope and checks that show whether the change actually works.

Scoping tool access for an AI agent
OpenAI's Responses API combines models with search, files and computer use. Each tool gives the agent a different kind of authority, so I define its identity and approval rules before connecting it.
2024

What computer use exposes in business software
Anthropic's computer-use beta can click, type and navigate software without a purpose-built API. That makes every ambiguous label, hidden state and irreversible button part of the agent's operating environment.

Why I’m testing a graph for long-term AI memory
A long-running AI collaborator needs to recover relationships between people, projects, decisions and evidence. I am testing a graph-backed memory layer because those connections are difficult to preserve in a pile of transcripts, and I want the stored reasoning to remain inspectable.