Blog
Designing AI products around actions and state
Gemini 2.0 is being presented around native tool use and agent-style tasks. A product built on that direction has to show what the model is doing, which action is waiting and what the user can still stop.

Google announced an experimental Gemini 2.0 Flash model yesterday with native tool use, including Google Search, code execution and user-defined functions. Its launch material also shows research prototypes that can act through a browser and work on coding tasks. Tasks that call tools need more interface state than a chat box provides. The product has to display when a task is waiting on a tool, asking for approval, retrying a failed action or leaving a partial result behind. Every agent-style run should have a task record outside the chat transcript. Give it an identifier, owner, stated objective, creation time, current status and links to the inputs it is allowed to use. The status needs to describe the work honestly. Queued, running, waiting for input, waiting for approval, blocked, completed, cancelled and failed are useful starting points. Do not display "working" while the system is waiting indefinitely for a tool or has lost its worker. Keep the user's original request, but add a structured task definition where the product can. A request to "update the records" still needs a target set, permitted fields and completion condition. If the system cannot resolve those safely, it should ask before acting.
Durable task state also makes refresh and reconnection ordinary. The user can leave the page and return without relying on one browser session to remember what happened. A model request to use a tool is a proposed action. The application should turn that proposal into a typed record with a tool name, arguments, target, risk level and current state. Validate arguments against a schema before execution. Resolve human-friendly references to stable identifiers, then show those identifiers in a form the user can understand. Check permissions with the task's execution identity, not the model's claim that an action is allowed. Action states may include proposed, approved, executing, succeeded, failed, rejected and cancelled. Store the result and error against the action. This gives the interface something concrete to display and prevents the transcript from becoming the only source of truth. Use an idempotency key for actions that may be retried. A network timeout after a successful external call should not cause a second payment, message or record creation. Approval should appear immediately before an action that crosses a meaningful boundary. Asking for confirmation on every read trains users to accept prompts without inspection. Asking once at the start of a broad task may authorise changes the user could not yet see.
Show the proposed effect. For a message, display recipient and final content. For an update, show the record and changed fields. For a bulk action, show the selection rule, count and a sample or downloadable review set where appropriate. Approval records should include who approved, what exact action they saw and when. If the proposal changes afterwards, require approval again. A permission to draft does not imply permission to send. Some tools can remain automatic within bounded conditions. Reading an approved knowledge source or running isolated calculation code may not need a person each time. The product still needs to log the action and stop if the tool goes outside its allowed scope. Long tasks spend a surprising amount of time waiting. They may need user input, external capacity, a rate-limit window, a supplier response or approval from somebody else. That is state the product should show. A waiting task needs a reason, the event that can resume it and an expiry or escalation rule. "Waiting for approval from the account owner" is useful. "Pending" is not. Notify the person who can unblock it and avoid repeatedly telling everybody else that the agent is thinking.
Failures should retain the last successful boundary. If four of five actions completed, the interface must say which four. A single red "failed" status can lead a user to rerun the whole task and duplicate completed work. Classify whether a failure is retryable, requires changed input or needs human repair. Preserve the original error for operators, while giving the user a concise explanation and safe choices. A stop button cannot reverse an action already accepted by an external system. The product should distinguish stopping future work from compensating for completed work. When a user cancels, mark the task as cancellation requested and prevent new actions from starting. Let an in-flight operation reach a known boundary or time out. Then list completed, interrupted and unstarted actions. If compensation is available, propose it as a separate action with its own approval. Design this before connecting tools. Ask each tool what cancellation means, whether a request can be interrupted, how to detect final state and whether the effect is reversible. Sending an email and generating a local draft have very different answers. The stop control must remain available outside the model's own interface and permissions. An agent should not be able to hide or disable the user's route to stop it.
Show the plan at a useful level, sources consulted, tools called, actions proposed, outputs received and changes made. Do not expose a theatrical stream of internal reasoning as though it were an audit record. Attach citations to retrieved material where the task depends on it. For code execution, show the code or a bounded summary plus the result and environment. For record changes, retain before and after values. Mark model-generated interpretation separately from tool-returned facts. The activity view should be readable in sequence and filterable by action status. Operators may need more detail than end users, including request identifiers and raw errors, but sensitive arguments and credentials should be redacted. Take one existing tool-enabled prototype and draw its state diagram before adding another tool. Include task states, action states, approval, timeout, retry, cancellation and partial completion. For every transition, write the screen state and operator record that will show what happened. Add those missing views before expanding the tool set.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.