Blogs
Featured blogs
Generated image: Villar David editorial
Blogs
Practical notes from the work.
Software, applied AI, product decisions and the operating details that decide whether a system holds up.
2026

Updating a procedure after a workaround
Coding agents are running for longer and reading more context before they return. That makes the current procedure more important because an agent can follow an old workaround much further before anyone notices.

Checking an integration after the job says it finished
A long agent run can complete most of its work and still miss an important requirement. I want the same distinction in production integrations: a completed run and a verified result are separate facts.

What stays yours when the agent harness is managed
OpenAI's Agents API now provides hosted sessions, environments, context management, tool loading and subagents. Product teams still have to define the work, authority, evidence and recovery path around that managed harness.

Recheck the permission boundary when the model changes
OpenAI has classified GPT-6 Astra at its Critical cybersecurity capability threshold. A team moving to a materially more capable model should review the identity, tools and approvals around it before treating the upgrade as a drop-in replacement.

Coding agents are changing the economics of technical debt
Some migrations remain postponed because their manual cost is larger than their visible business value. Parallel agents and fast verification can change that calculation, although only where the repository has clear conventions, reliable tests and a definition of done that a reviewer can check.

What I changed in agent memory
I now keep working context separate from durable records and searchable history. The change came from watching old material enter sessions where it was technically related but useless to the task in front of the agent.

A task-board export needs an experiment index
A task board can show that work moved while losing the experiment behind it. Add an index that links each test to its question, evidence and result so later work can tell what is safe to reuse.

Write the handoff before clearing the agent context
Before ending an agent session, I write down the current state and the command that should run next. Without that, the next session has to reconstruct the work and can repeat changes that are already complete.

A score needs a methodology readers can inspect
If software gives an organisation a capability score, the reader needs to see what was measured and which rule version produced it. Build that explanation beside the scoring logic so it changes with the method.

Write the evaluation set while the AI feature is still moving
Teams often postpone evaluation until the prompt and interface feel finished. I prefer collecting difficult real cases during development because they reveal whether a change improves the underlying task or simply makes the latest demonstration look cleaner.

Govern the whole agent system
I review an agent by tracing what it can read, what it can change and how a person can stop or reverse it. That includes the harness and credentials as well as the model.

A conversational assessment still needs deterministic scoring
Keep the conversation and the scoring engine separate. The conversation captures evidence; a versioned rule set calculates the result and leaves a reviewer able to reproduce it.

What embedded AI delivery uncovers in the workflow
Working beside the team using an AI system exposes details the brief rarely contains: distrusted source data, exceptions handled outside the software and approvals nobody wrote down.

Build the searchable specification before the application gets expensive
When requirements are scattered across PDFs, email and meeting notes, implementation starts by repeatedly rediscovering the product. Consolidate the source material into a searchable specification and keep each decision linked to the evidence behind it.

Where agent supervision has to happen
Anthropic's recent analysis suggests experienced users approve more automation while also interrupting agents more often. That points to supervision built around visible plans and consequential boundaries instead of a confirmation attached to every command.

The coding-agent interface is becoming a control plane
Parallel agents, isolated worktrees and queued review change the developer’s role from typing every change to supervising work in motion. The difficult part is setting boundaries, spotting drift early and keeping enough evidence to understand each result before integration.

Production agents need institutional context
OpenAI's internal data-agent write-up shows how much context sits around the model: schemas, code, permissions and previous corrections. The useful design question is who maintains that context and how a user can correct it.

Reconciling portals, permissions and integrations before release
In a multi-portal product, the final release checks often fail between components: one role sees the wrong state, an integration updates late or two screens apply different rules. The examples in this post are a composite of recurring delivery problems.
2025

The agent’s activity log is part of the deliverable
When an agent changes code, data or configuration, I need the work record alongside its final answer. The next person should be able to see what was attempted, what passed and where the system was left.

What I want from an always-on AI agent
I am treating an always-on agent as a bounded worker. It should accept a named job, leave an activity record and put any proposed change somewhere a person can review it.

Operating local AI infrastructure after the migration
A successful migration is followed by less glamorous work: services starting in the right order, data paths staying stable, backups being restorable and failures being visible. I am treating those operating details as part of the AI product because continuity disappears when the infrastructure is opaque.

What has to survive when an AI assistant moves machines
Before moving an AI collaborator to a new host, I am listing the parts that have to survive and the tests that will prove they did. The prompt is easy to copy; the harder state sits in records, tools and the links back to source material.

Trying to remember everything made the agent worse
A total-recall experiment has made the problem clear: more retrieved history can crowd out the task in front of the agent. I am separating identity, durable decisions, project state and searchable archives so memory is selected by purpose instead of poured into every session.

One bulk query can beat thousands of careful loops
I have been working through data pipelines where individually reasonable queries create unreasonable runtimes at scale. Moving the work into bulk operations changes throughput dramatically. A bad bulk update can damage far more records before anyone sees it, so the job needs a dry run, checkpoints and a way to reconcile the result.

What I am testing in long-term AI memory
I am testing whether a graph can recover a decision through its relationships to a person, project and source. The immediate question is which records improve the next session and which simply make retrieval noisier.

Writing work orders for coding agents
Coding agents can now accept an issue and return a pull request, which makes the issue itself part of the engineering system. A useful work order needs a bounded objective, relevant context, acceptance criteria, permitted scope and checks that show whether the change actually works.

If software produces a score, the rules need a version
A score can look objective while hiding a changing set of weights, thresholds and assumptions. I want the rule set versioned alongside the result so a reviewer can reproduce what happened, explain it to the person affected and distinguish a changed answer from changed evidence.

Scoping tool access for an AI agent
OpenAI's Responses API combines models with search, files and computer use. Each tool gives the agent a different kind of authority, so I define its identity and approval rules before connecting it.

What DeepSeek-R1 changes in a model evaluation
DeepSeek-R1 gives teams another open reasoning model to test against their own work. I would compare task success, latency and operating requirements on the same evaluation set before treating its published benchmarks as a product decision.

Long-running data jobs should be designed to stop
Large imports rarely fail at a convenient boundary. They need deliberate stopping points, a record of what completed and a restart path that does not duplicate accepted work.
2024

Dates, dashboards and the work that has not processed
A dashboard can look current while timezone assumptions or hidden processing limits tell a different story. I am putting unprocessed records, date rules and result limits into the interface because operational users need to see incomplete work rather than discover it through a customer complaint.

Designing AI products around actions and state
Gemini 2.0 is being presented around native tool use and agent-style tasks. A product built on that direction has to show what the model is doing, which action is waiting and what the user can still stop.

Accounting integrations need explicit sync state
An integration becomes difficult to support when the application cannot say whether a record is waiting, synced, rejected or changed afterwards. I am making sync state visible so retries and reconciliation have evidence and staff do not have to compare two systems by hand.

What computer use exposes in business software
Anthropic's computer-use beta can click, type and navigate software without a purpose-built API. That makes every ambiguous label, hidden state and irreversible button part of the agent's operating environment.

What I am testing with o1-preview
OpenAI's o1-preview spends more time reasoning before it answers. I want to test it on tasks where the reasoning can be checked independently, because a longer process is only useful when the result survives verification.

AI governance has entered the product backlog
The EU AI Act is now in force, with obligations arriving in stages. Product teams should start by recording where AI is used and who is responsible for each use.

Paid modules change more than the checkout screen
Adding paid access to a learning platform touches enrolment, permissions, account state, completion records, certificates and support. I am breaking the work into those consequences because a clean payment screen does not help if the learner receives the wrong access or the records disagree afterwards.

Why I’m testing a graph for long-term AI memory
A long-running AI collaborator needs to recover relationships between people, projects, decisions and evidence. I am testing a graph-backed memory layer because those connections are difficult to preserve in a pile of transcripts, and I want the stored reasoning to remain inspectable.

Using voice and images in an existing workflow
GPT-4o and other recent demonstrations show text, images, audio and live visual input moving into the same interaction. I am interested in the ordinary workflow consequence: what becomes quicker to show, what still needs confirmation and which input has to be kept as evidence.

When I would test an open-weight model
Llama 3 makes open-weight models worth testing against a defined product task. I would compare it with a hosted model using the same cases, then account for the hardware and operational work needed to run it.

The workflow that starts when a payment clears
Recording a successful payment is only one part of the job. The surrounding workflow has to update the right account, send the right message, avoid duplicate handling and leave enough state for support staff to work out what happened when one step fails.

How I decide whether an AI idea deserves a build
A convincing demo is easy to mistake for a product opportunity. I first ask who would notice a wrong result, what they could check and whether the time saved survives the review work the feature creates.

Start with the work people are already doing
AI workflows built from the documented procedure usually miss how the job is done. I watch the handoffs and workarounds first, especially where people leave the official system to finish the task.

What applied AI means in my work
Applied AI starts with a piece of work that already exists. I want to understand the decision, the information available, the cost of delay and what happens when the answer is wrong before choosing a model or building an interface.

Why I’m writing this down
I keep notes while I work, usually because a decision will need explaining again months later. This journal is where I will keep the implementation details, failures and changes of mind that are worth returning to.