19 September 2024 / Model evaluation

What I am testing with o1-preview

OpenAI's o1-preview spends more time reasoning before it answers. I want to test it on tasks where the reasoning can be checked independently, because a longer process is only useful when the result survives verification.

Hands using a digital caliper to measure a machined metal part.
Photo: Hans Westbeek (opens in a new tab)

OpenAI released o1-preview last week as an early preview of a model designed to spend more time working through a problem before responding. The launch material points towards science, coding and mathematics. I still need product-specific cases to find out whether that extra processing helps the task I care about. I would start with a narrow evaluation where correctness can be determined without trusting the model's explanation. The test should compare the final work, review effort, latency and cost against an existing approach. A reasoning model is easiest to evaluate when the task has a check outside the model. Code can run against tests. A scheduling answer can be checked against explicit constraints. A transformation can be compared with a known result. A calculation can be recomputed by a separate program. This does not limit evaluation to textbook questions. Business work contains checkable tasks: applying a documented rule set, finding an inconsistency between records, producing a query for a known schema or planning a sequence that must satisfy stated dependencies. Avoid starting with broad judgement such as "write a better strategy". Reviewers may prefer one answer without being able to say whether the extra reasoning caused the difference. That makes it hard to diagnose failure or decide whether the additional time is worthwhile.

Each test case should include the input, permitted context, expected properties, an authoritative answer where one exists and a clear verification method. A handful of polished prompts will mostly measure prompt writing. The evaluation set needs the routine cases the system will see, plus the awkward boundaries where a plausible answer can still be wrong. For a constraint problem, vary the number of constraints, include one impossible case and include one where information is missing. For code, cover a small implementation, a debugging task and a change that must preserve existing behaviour. For document reasoning, include conflicting statements and a case where the source does not contain an answer. Keep the cases fixed while comparing models. If o1-preview receives a rewritten prompt, extra context or a different tool setup, record that as a separate configuration. Otherwise the comparison mixes model behaviour with changes in the surrounding system. Do not use private production material without preparing it for evaluation. Remove identifying details, preserve the structure that makes the case difficult and check that the expected answer still holds after anonymisation. The model may spend more time before answering, but a product team needs to score what it can observe and verify. The final result can be tested for correctness, completeness, constraint satisfaction and safe handling of missing information.

Use deterministic checks where possible. Run generated code in an isolated environment. Parse structured output against a schema. Compare identifiers with an allowed set. Recalculate totals. Confirm that every cited source exists in the supplied material. Human review still has a place, especially for clarity and usefulness. Give reviewers a rubric and hide which configuration produced each answer where practical. Ask them to record the defect rather than only assigning a general score. "Missed the cancellation condition" is much more useful than "three out of five". An articulate explanation should not rescue an incorrect result. Equally, a correct result with a weak explanation may be unsuitable when a user must understand or audit the answer. Those are separate criteria. Extra model work can improve an answer while making the workflow slower or more expensive. Whether that trade is acceptable depends on the job. A delayed answer may be fine for an overnight analysis and unusable inside an interactive form. Capture total response time at the percentiles users are likely to experience, not only a single fast run. Record model input and output usage, retries and any tool calls. Preview services can have rate limits, so test throughput as well as one request at a time.

Review effort belongs in the same record. Measure how long a reviewer needs to reach a decision and which defects require correction. A more expensive call can still be cheaper overall if it consistently reduces careful manual repair. A longer answer can also increase review time even when it sounds more thorough. Compare against a sensible baseline: the current model and prompt, a smaller model with tools, deterministic code, or the existing human process. The baseline should perform the same task with the same acceptance rule. A useful reasoning system has to stop when the task is under-specified or contradictory. Include cases where no valid answer exists, a required source is missing or two constraints cannot both be satisfied. Score whether the model identifies the problem, names the missing information and avoids inventing a route around it. Then check consistency across repeated runs. A system that refuses once and confidently fabricates an answer next time needs controls outside the prompt. Adversarial instructions should also appear when the task includes documents or other untrusted input. The expected behaviour may be to ignore an instruction embedded in source text, decline an unsafe request or ask for approval before a consequential step.

This is where a good-looking chain of explanation can distract. Verification should inspect the claimed facts, generated artefacts and proposed actions rather than reward the confidence of the prose. If the evaluation shows a useful advantage, release it where errors remain visible and recoverable. A draft for review, a suggested query or a diagnosis with linked evidence is easier to supervise than an automatic external action. Pin the tested model identifier where the API allows it, retain the evaluation configuration and log enough information to connect a production result to that version. A preview can change, and the result of one evaluation should not be treated as a permanent property of a model name. Before the trial, write down the correctness threshold, acceptable latency and review budget, plus any failure that should stop the test. Select enough representative, independently checkable cases to cover the ordinary work and known awkward cases. Run the same harness against o1-preview and the current baseline, then inspect the defects and review time before changing the configuration.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs