20 February 2025 / Model evaluation

What DeepSeek-R1 changes in a model evaluation

DeepSeek-R1 gives teams another open reasoning model to test against their own work. I would compare task success, latency and operating requirements on the same evaluation set before treating its published benchmarks as a product decision.

Hands using a digital caliper to measure a machined metal part.
Photo: Hans Westbeek (opens in a new tab)

DeepSeek-R1 gives teams another open reasoning model to test against their own work. I would compare task success, latency and operating requirements on the same evaluation set before treating its published benchmarks as a product decision. DeepSeek published R1 and R1-Zero model weights in January, along with distilled models based on Qwen and Llama. The release repository reports results across maths, coding and reasoning benchmarks. Teams with suitable infrastructure can also evaluate the released models on local or privately managed systems. A new model should enter an existing evaluation as another candidate. If the tasks, prompts or scoring rules change for R1, the comparison becomes difficult to interpret. I would start with work sampled from the intended product. For a support workflow, that could include incomplete questions, conflicting source documents, requests requiring a refusal and answers that need precise citations. For coding, use repository tasks with a known starting commit, tests and an expected patch boundary. The evaluation should contain ordinary work as well as cases that expose expensive failures.

Keep inputs fixed where the product can keep them fixed. Use the same system instructions, retrieval results, tool definitions and output schema. If a model requires a different prompt to perform properly, record that as a separate configuration rather than quietly tuning one candidate until it wins. Scoring should test whether an answer is usable, not merely plausible. Depending on the task, that may require exact matches, executable tests, schema validation or a reviewer rubric. A small evaluation with clear scoring is more useful than a large pile of prompts nobody can grade consistently. DeepSeek's reported numbers are evidence about the model under the stated benchmark setup. They are not measurements of a particular product. The repository says its sampling evaluations used a temperature of 0.6, top-p of 0.95 and 64 responses per query to estimate pass@1. A production request may generate once, under a tighter time limit, with retrieved context and tools in the loop. Reproducing the benchmark settings would still not measure that production path.

I would retain a few public benchmarks as a sanity check, especially where they resemble the work. But the decision sheet should separate vendor-reported results from internally reproduced tests and record the exact model identifier, serving provider, quantisation if any, prompt template and generation settings. Without those details, a later claim that "R1 scored 90" cannot be tied to a test, metric or configuration. Reasoning models may spend more tokens and time before returning an answer. That can be worthwhile when the task benefits from deliberate work, but the trade belongs in the evaluation. Record end-to-end latency at realistic concurrency, including queueing and tool calls. Percentiles are more informative than one average because a few slow responses can dominate the user experience. Count input and output tokens and record request cost under the pricing and limits available at the time. For a self-hosted model, the evaluation must include infrastructure and operating costs. The principal R1 release is a mixture-of-experts model with 671 billion total parameters and 37 billion activated per token, according to DeepSeek. "Only 37 billion active" does not mean it has the memory footprint of a 37 billion parameter dense model. Hardware capacity, model loading, parallelism and serving software all need an actual test.

The distilled releases, ranging from 1.5 billion to 70 billion parameters, create more practical local candidates. They should be evaluated as distinct models. A smaller distillation may have lower operating demands and different failure patterns from the full release. The deployment choice changes what the team must operate. With an external API, I would inspect retention terms, regional availability, rate limits, incident handling and model updates. With a privately managed endpoint, the list includes GPU availability, patching, monitoring, autoscaling, access control and who responds when inference slows down. Licensing needs a precise check as well. DeepSeek's repository applies the MIT licence to its code and model weights and permits commercial use. The distilled Qwen and Llama variants also carry conditions from their base models, which the repository calls out. The legal review should use the exact artefact being deployed, not the family name in a slide.

Security testing should use the same application boundary planned for production. If the model receives untrusted retrieved text, test prompt injection. If it can call tools, test malformed arguments and attempts to exceed its authority. Local weights do not remove application security work. For each run, retain the evaluation set version, model and serving configuration, prompt version, raw outputs, scores and reviewer notes. Then compare candidates task by task. One model may be stronger on multi-step analysis while another is fast and adequate for classification. That can support routing, or show that the extra complexity is not worth carrying. Use the results to decide whether R1 improves enough of the actual workload to justify its latency, deployment work and risk profile. Run the fixed set, inspect the failures, and write down the configuration. The published leaderboard can stay as context.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs