Applied AI blogs

Applied AI

Posts filed under Applied AI. View every category.

2026

3 September 2026 / AI governance

Recheck the permission boundary when the model changes

OpenAI has classified GPT-6 Astra at its Critical cybersecurity capability threshold. A team moving to a materially more capable model should review the identity, tools and approvals around it before treating the upgrade as a drop-in replacement.

20 August 2026 / Applied AI

Coding agents are changing the economics of technical debt

Some migrations remain postponed because their manual cost is larger than their visible business value. Parallel agents and fast verification can change that calculation, although only where the repository has clear conventions, reliable tests and a definition of done that a reviewer can check.

23 July 2026 / Applied AI

What I changed in agent memory

I now keep working context separate from durable records and searchable history. The change came from watching old material enter sessions where it was technically related but useless to the task in front of the agent.

25 June 2026 / Applied AI

A task-board export needs an experiment index

A task board can show that work moved while losing the experiment behind it. Add an index that links each test to its question, evidence and result so later work can tell what is safe to reuse.

11 June 2026 / Applied AI

Write the handoff before clearing the agent context

Before ending an agent session, I write down the current state and the command that should run next. Without that, the next session has to reconstruct the work and can repeat changes that are already complete.

28 May 2026 / Applied AI

A score needs a methodology readers can inspect

If software gives an organisation a capability score, the reader needs to see what was measured and which rule version produced it. Build that explanation beside the scoring logic so it changes with the method.

14 May 2026 / Applied AI

Write the evaluation set while the AI feature is still moving

Teams often postpone evaluation until the prompt and interface feel finished. I prefer collecting difficult real cases during development because they reveal whether a change improves the underlying task or simply makes the latest demonstration look cleaner.

23 April 2026 / Applied AI

Govern the whole agent system

I review an agent by tracing what it can read, what it can change and how a person can stop or reverse it. That includes the harness and credentials as well as the model.

26 March 2026 / Applied AI

What embedded AI delivery uncovers in the workflow

Working beside the team using an AI system exposes details the brief rarely contains: distrusted source data, exceptions handled outside the software and approvals nobody wrote down.

2025

2024

12 December 2024 / AI products

Designing AI products around actions and state

Gemini 2.0 is being presented around native tool use and agent-style tasks. A product built on that direction has to show what the model is doing, which action is waiting and what the user can still stop.

19 September 2024 / Model evaluation

What I am testing with o1-preview

OpenAI's o1-preview spends more time reasoning before it answers. I want to test it on tasks where the reasoning can be checked independently, because a longer process is only useful when the result survives verification.

20 August 2024 / AI governance

AI governance has entered the product backlog

The EU AI Act is now in force, with obligations arriving in stages. Product teams should start by recording where AI is used and who is responsible for each use.

16 May 2024 / Multimodal AI

Using voice and images in an existing workflow

GPT-4o and other recent demonstrations show text, images, audio and live visual input moving into the same interaction. I am interested in the ordinary workflow consequence: what becomes quicker to show, what still needs confirmation and which input has to be kept as evidence.

23 April 2024 / Model evaluation

When I would test an open-weight model

Llama 3 makes open-weight models worth testing against a defined product task. I would compare it with a hosted model using the same cases, then account for the hardware and operational work needed to run it.

20 February 2024 / Applied AI

How I decide whether an AI idea deserves a build

A convincing demo is easy to mistake for a product opportunity. I first ask who would notice a wrong result, what they could check and whether the time saved survives the review work the feature creates.

18 January 2024 / Applied AI

What applied AI means in my work

Applied AI starts with a piece of work that already exists. I want to understand the decision, the information available, the cost of delay and what happens when the answer is wrong before choosing a model or building an interface.