Blog
Using voice and images in an existing workflow
GPT-4o and other recent demonstrations show text, images, audio and live visual input moving into the same interaction. I am interested in the ordinary workflow consequence: what becomes quicker to show, what still needs confirmation and which input has to be kept as evidence.

OpenAI announced GPT-4o on 13 May as a model working across text, audio, images and video. The demonstrations are striking, but the useful product question begins with an existing point of friction. Some information is awkward to turn into a text prompt. A person may need to describe a damaged component, read a label while their hands are occupied, identify a field on a printed form, or explain a visual problem in software. Letting them show the object or speak naturally can shorten that input step. A camera or microphone request needs a named outcome. “Capture the serial number for confirmation” and “point out the control that needs attention” can be tested. The interface can then ask for the specific image, voice or screen input required without collecting a broad stream by default. Start with one moment where typing causes delay or loses useful detail. Keep the rest of the process unchanged until that input method proves useful. Voice input feels direct, but speech can be misheard and the speaker may change direction mid-sentence. Images can contain several plausible values or omit the context needed to interpret them. A model response should therefore present important captured facts for confirmation before they affect a record or action.
For a spoken address, show the parsed address in fields the person can edit. For a photographed label, display the extracted identifier beside the relevant crop. For a visual support flow, ask the user to confirm which object or region they mean instead of continuing from an uncertain reference. Confirmation should match consequence. A casual search query may need no extra step. An amount, identity detail, appointment time or instruction sent to another person deserves explicit review. The interface should make correction easy rather than asking a generic “Is this right?” beneath a long transcript. Where the system has low confidence or conflicting cues, it should ask a focused question. Repeating the entire interaction wastes time and may produce the same ambiguity. A multimodal interaction can produce original audio, images or video, a transcript, extracted fields, model responses and user corrections. Retention should follow the job and its risk. A consequential workflow may need the original capture and correction history, while a low-risk search field may need only the corrected text. Define the record for the task before collecting input. If a photograph supports an asset inspection, the original image may be required alongside the accepted observation. If voice is only a convenient way to fill a low-risk search field, the corrected text may be enough and the audio may not need retention.
The policy should cover purpose, access and retention. Users need to know when recording starts and what will be kept. Background speech, faces, documents and screens can introduce information about other people who did not intend to join the interaction. Corrections are part of the record where they change an operational fact. Store the value proposed by the system, the value accepted by the user and the resulting action. This is more useful than a raw transcript when someone later asks why a field changed. Camera and microphone permission at the device level does not answer who may view or process the captured material inside the product. The application needs its own access rules tied to the task. A user who can submit a photograph may not be allowed to browse every image attached to an account. A support operator may need the extracted value but not the full recording. A model service should receive only the content needed for the current step, with credentials and logs configured for the relevant data boundary. Live visual input adds another issue: the frame can change while the model is responding. The product should show which image or moment the answer refers to. If an action depends on it, capture a deliberate frame and ask for confirmation rather than relying on an unmarked live stream.
Test permission withdrawal and interrupted capture. If the user denies microphone access, the same task should have a text path. If an upload stops halfway through, the interface should not present a partial analysis as complete. The GPT-4o announcement describes broad multimodal capability, with access rolling out in stages. A product evaluation should use only the inputs and outputs available in the intended environment, not assume every demonstration capability is ready in the same API or interface. Create task cases that include poor lighting, background noise, accents, overlapping speech, rotated documents, small text and irrelevant objects in frame. Add privacy cases, such as another person's details appearing beside the target information. Define when the system should continue, request another capture or stop. Measure more than model accuracy. Record time to complete the task, number of recaptures, user corrections, abandoned attempts and reviewer effort. Compare this with the current typed or manual route. A voice interaction that transcribes well but is difficult to correct may be slower overall. Run tests on the devices and network conditions people will actually use. Upload speed, microphone quality and mobile permission behaviour can dominate the experience before model quality becomes relevant.
A useful first release does not need a continuous conversation using every modality. Add the input mode that addresses the identified friction and keep a visible alternative. For image input, define framing guidance, file limits, preview, retake and confirmation. For voice, define start and stop controls, live indication, transcript review and a way to correct specific fields. For either, log processing failures without storing sensitive content unnecessarily. Keep model output separate from confirmed workflow state. An extracted value can remain proposed until a person accepts it. A suggested action can remain a draft until the authorised user approves it. Downstream services should consume the confirmed record, not an unstable conversational message. Take one existing form or support step and mark the field people find hardest to express in text. Test voice or image input beside the current method. Record completion time, corrections and recaptures, then decide whether that input should remain available.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.