Articles

Long-form documents, written in chapters.

Detailed guides on applied AI, software architecture, technical leadership and the operating work around them.

2026

22 September 2026 / Agent operations / 8 chapters

Recovering an automated workflow after a partial failure

When an automated workflow stops halfway through, some of its changes may already exist in other systems. Before restarting it, establish what completed and which actions can be repeated safely.

27 August 2026 / Applied AI / 8 chapters

Budgeting and governing an AI workflow after the prototype

I start with the full operating cost: review, support, failed runs and the work needed to keep source data usable. A cheap model call can still sit inside an expensive process, and that process needs an owner after the prototype is approved.

30 June 2026 / Software delivery / 8 chapters

Managing review when coding agents outpace it

The pressure point appears when agents can produce changes faster than people can inspect them. This article starts there and works back through task definition, isolated execution and the merge controls needed to keep review meaningful.

30 April 2026 / Applied AI / 8 chapters

Designing an auditable AI capability assessment

An AI-assisted assessment can collect conversational evidence while a deterministic ruleset produces the score. This guide develops that architecture through versioned rules, fixtures and an audit record a reviewer can reproduce.

19 February 2026 / Software delivery / 8 chapters

Releasing a multi-portal platform without losing state between components

The examples in this article are a composite of late-stage failures between roles, portals and integrations. The guide shows how to isolate each failure for agent-assisted implementation and human review.

2025

27 November 2025 / Applied AI / 8 chapters

Self-hosting an AI agent: architecture, observability and recovery

A self-hosted agent can appear healthy while its memory store, tool credentials or job runner is unavailable. This guide starts with those partial failures and builds the health checks and recovery tests around them.

25 September 2025 / Applied AI / 9 chapters

Designing an audit trail for tool-using agents

When an agent reads private material or changes an external system, a chat transcript is an unreliable audit trail. This article defines the events worth recording, how to connect them to a task and how to keep the log useful during review and recovery.

30 June 2025 / Applied AI / 8 chapters

Long-term AI memory: records, relationships and retrieval

A transcript records what was said. Durable memory also needs current facts, links to their sources and a correction path; this article develops the records and relationships that support that work.

30 April 2025 / Applied AI / 8 chapters

Deciding what a tool-using agent can read and change

Start with one task and give the agent only the access needed to complete it. This article follows the permission decisions from read-only source material through to an action that requires a person to approve it.

27 February 2025 / Data systems / 8 chapters

Designing resumable large-data pipelines

This guide begins with a job stopped halfway through an import. It develops the state and restart rules needed to identify accepted work, isolate failed records and continue without rebuilding the entire run.

2024

19 December 2024 / Applied AI / 8 chapters

Production readiness for applied AI

The first production question I ask is what happens after a wrong answer reaches the workflow. This guide works backwards from detection and recovery to the permissions, source checks and release evidence needed before launch.

17 October 2024 / Applied AI / 8 chapters

Evaluating an AI workflow before customers depend on it

An AI feature needs an evaluation set that reflects the work it will actually perform. This article shows how to collect representative tasks, define acceptable outputs, test failure modes, measure human review effort and decide whether the workflow is ready for a controlled release.

31 July 2024 / Applied AI / 8 chapters

Building retrieval-augmented generation for real use

A RAG demo can answer prepared questions from a clean document set. Production material brings mixed formats, stale documents, conflicting permissions and questions the source cannot answer; the guide follows those problems from ingestion through evaluation.

30 May 2024 / Technical leadership / 8 chapters

How technical decisions break down as a team grows

As more people join a technical team, the same decision starts being made in several places. This article examines what has to be written down, which decisions can stay local and where review prevents incompatible versions of the product from taking hold.

29 February 2024 / Product delivery / 8 chapters

Map the work before replacing its software

When a team asks for replacement software, I begin with one real piece of work and follow it from request to completion. The missed handoffs and private workarounds usually explain more than the current feature list.