23 April 2024 / Model evaluation

When I would test an open-weight model

Llama 3 makes open-weight models worth testing against a defined product task. I would compare it with a hosted model using the same cases, then account for the hardware and operational work needed to run it.

Hands using a digital caliper to measure a machined metal part.
Photo: Hans Westbeek (opens in a new tab)

Meta released the first Llama 3 models on 18 April in 8 billion and 70 billion parameter versions. They can now be included in a product evaluation. Self-hosting still needs a specific operating or product reason. I would start with the constraint that makes an open-weight model relevant. It might be a requirement to keep inference inside a controlled environment, a need to tune behaviour for a narrow task, predictable high-volume use, or a product that must keep operating without dependence on one hosted endpoint. Each reason leads to a different test. If the reason is “more control”, I break that down into data location, access, uptime, output behaviour and release control. Possession of the weights does not settle any of those. The surrounding service still needs authentication, request logging, deployment rules and an owner. A locally running model can leak sensitive input through careless logs just as easily as any other application. If there is no specific operating or product constraint, a hosted service may remain the simpler baseline. The evaluation should be allowed to reach that answer. Benchmark results are useful for understanding broad capability. Product selection needs cases taken from the intended job. A support classification task, document extraction task and code assistant place different demands on the model.

Build a small evaluation set before changing infrastructure. Include normal inputs, difficult inputs, missing evidence, long material, unusual formatting and cases where the correct response is to ask for clarification. Remove private data or run the evaluation in an environment approved for it. Define acceptable output for each case. For extraction, compare fields against verified source values. For question answering, require source support and mark unsupported claims. For generation, identify facts and instructions that must be preserved, then have reviewers record corrections rather than giving a single impression score. Run the open model and hosted baseline with equivalent task instructions and retrieval material. Exact prompts may need adjustment because model formats differ, but do not quietly give one option more context or more attempts. Keep the configuration with each result so a later rerun is meaningful. The Llama 3 release includes pretrained and instruction-tuned variants. Model size and format change hardware needs, latency and behaviour. An evaluation against a remote demonstration says little about the version the product can operate. Choose a realistic serving configuration. Record the weights, numerical precision, quantisation if used, inference engine, hardware, context length and generation settings. Quantisation can reduce memory requirements, but its effect on the target task should be measured rather than assumed. The smaller model may be sufficient for a bounded classification job and inadequate for another task with longer instructions.

Test concurrency, not only one prompt at a time. Measure time to first token where interaction matters, total completion time, tokens processed, queueing under expected load and memory use. Run long and short inputs because context size affects throughput. Also test failure behaviour. What happens when the model process restarts, the request exceeds the supported context, or the queue reaches its limit? A product needs a timeout, a retry rule and a clear response when inference is unavailable. Running weights is one component of an inference service. The deployment also needs an API boundary, identity checks, resource limits, monitoring and a release process. Requests should be authorised for the underlying task, not merely for access to a generic model endpoint. If the application retrieves private documents, enforce permissions before those passages enter the prompt. Decide which request and response data can be logged, how long logs are retained, and how an operator investigates a failure without exposing content unnecessarily. Monitoring should cover service health and task health. GPU memory, request latency and error rates show whether the service runs. They do not show whether a prompt change increased unsupported answers. Keep the task evaluation available for every model, prompt, retrieval or serving change.

Model artefacts need provenance. Store where the weights came from, their licence, checksum, configuration and any adapters applied. A deployment should be reproducible without relying on a manually edited server directory. A hosted model usually exposes a usage price. A self-hosted model spreads cost across hardware, idle capacity, engineering time, power, monitoring and incident response. Comparing only the price of tokens with the hourly price of a GPU leaves out the work needed to keep each option usable. Estimate expected volume and peak concurrency. A dedicated machine may sit mostly idle at low volume, while an on-demand host may have startup delay and availability limits. Larger models may require multiple accelerators or lower precision. Storage and network transfer matter during deployment even when inference is local. Include maintenance. Someone needs to patch the host, update the serving stack, test new model versions, manage capacity and respond when the endpoint slows down. If that work belongs to a team with no available operating time, it is a product constraint. Hosted options have operating risks too, including rate limits, provider changes and external data handling. Put both options in the same cost and risk record rather than treating one as infrastructure and the other as a simple API call.

The result should be a decision table grounded in observed runs. Record task quality, reviewer effort, latency under load, expected operating cost, data boundary, implementation effort and recovery path. Keep unresolved items visible. An open-weight model may win for one component and lose for another. A local model could classify routine internal material while a hosted model handles occasional complex drafting, provided the data rules allow it. There is no need to force the whole product onto one model family. Before committing, run the candidate as a limited internal service and repeat the evaluation after packaging it exactly as production would use it. Confirm that monitoring catches a stopped worker, a saturated queue and an invalid request. Confirm that the model version can be rolled back. I would choose one task, gather ordinary and failure cases, and price one serving configuration that could actually be deployed. If those numbers justify a local trial, I would then package and run the selected weights.

Continue the thinking.

Comments are public and hosted in an open-source GitHub Discussions repository.

Loading comments connects your browser to GitHub. A GitHub account is required to post.

All blogs