Applied AI blogs
Featured blogs
Generated image: Villar David editorial
Blogs
Applied AI
Posts filed under Applied AI. View every category.
2026

Recheck the permission boundary when the model changes
OpenAI has classified GPT-6 Astra at its Critical cybersecurity capability threshold. A team moving to a materially more capable model should review the identity, tools and approvals around it before treating the upgrade as a drop-in replacement.

Coding agents are changing the economics of technical debt
Some migrations remain postponed because their manual cost is larger than their visible business value. Parallel agents and fast verification can change that calculation, although only where the repository has clear conventions, reliable tests and a definition of done that a reviewer can check.

What I changed in agent memory
I now keep working context separate from durable records and searchable history. The change came from watching old material enter sessions where it was technically related but useless to the task in front of the agent.

A task-board export needs an experiment index
A task board can show that work moved while losing the experiment behind it. Add an index that links each test to its question, evidence and result so later work can tell what is safe to reuse.

Write the handoff before clearing the agent context
Before ending an agent session, I write down the current state and the command that should run next. Without that, the next session has to reconstruct the work and can repeat changes that are already complete.

A score needs a methodology readers can inspect
If software gives an organisation a capability score, the reader needs to see what was measured and which rule version produced it. Build that explanation beside the scoring logic so it changes with the method.

Write the evaluation set while the AI feature is still moving
Teams often postpone evaluation until the prompt and interface feel finished. I prefer collecting difficult real cases during development because they reveal whether a change improves the underlying task or simply makes the latest demonstration look cleaner.

Govern the whole agent system
I review an agent by tracing what it can read, what it can change and how a person can stop or reverse it. That includes the harness and credentials as well as the model.

What embedded AI delivery uncovers in the workflow
Working beside the team using an AI system exposes details the brief rarely contains: distrusted source data, exceptions handled outside the software and approvals nobody wrote down.
2025
2024

Designing AI products around actions and state
Gemini 2.0 is being presented around native tool use and agent-style tasks. A product built on that direction has to show what the model is doing, which action is waiting and what the user can still stop.

What I am testing with o1-preview
OpenAI's o1-preview spends more time reasoning before it answers. I want to test it on tasks where the reasoning can be checked independently, because a longer process is only useful when the result survives verification.

AI governance has entered the product backlog
The EU AI Act is now in force, with obligations arriving in stages. Product teams should start by recording where AI is used and who is responsible for each use.

Using voice and images in an existing workflow
GPT-4o and other recent demonstrations show text, images, audio and live visual input moving into the same interaction. I am interested in the ordinary workflow consequence: what becomes quicker to show, what still needs confirmation and which input has to be kept as evidence.

When I would test an open-weight model
Llama 3 makes open-weight models worth testing against a defined product task. I would compare it with a hosted model using the same cases, then account for the hardware and operational work needed to run it.

How I decide whether an AI idea deserves a build
A convincing demo is easy to mistake for a product opportunity. I first ask who would notice a wrong result, what they could check and whether the time saved survives the review work the feature creates.

What applied AI means in my work
Applied AI starts with a piece of work that already exists. I want to understand the decision, the information available, the cost of delay and what happens when the answer is wrong before choosing a model or building an interface.