An agent’s result depends on the software coordinating its work, the model generating its responses, and the infrastructure running inference. Separating those layers makes it easier to investigate slow requests, failed edits, or inconsistent results.
The harness coordinates the work
The harness calls the model and manages its tools, context, and task state. Coding tools such as Claude Code, Codex, and OMP each make choices about what the model sees and how its proposed actions are executed.
Repeated context, failed tool calls, and oversized outputs can add tokens and delay. When an agent loops, inspect its prompts, tool results, and stopping conditions as well as the model’s responses. The transcript is more useful than assuming one layer caused the failure.
The model generates responses
The model uses trained weights to generate responses from the context it receives. Evaluation results also depend on the agent setup and tests around it. SWE-bench distinguishes model and agent configurations, so compare results under matching conditions.
Serving runs inference
Serving is the system that runs inference and returns tokens. Hardware and scheduling affect latency; numerical precision and runtime configuration can also affect outputs. Investigate the settings and measurements available for the service you are using:
- Latency: measure time to the first token and the rate of generation separately from tool execution.
- Precision: use documented configuration or controlled measurements to assess quantization effects; do not infer a change from one poor response.
- Reliability: record failed requests, rate limits, and retries when comparing endpoints.

Compare runs under recorded conditions
If a run is slower than expected, record request latency, token counts, tool time, and errors. A slow response alone does not establish that a provider changed precision or used a weaker model.
When comparing agents, hold the task and evaluation criteria steady. Record the harness version and configuration, model identifier, and serving endpoint alongside the result. The formula below is a conceptual shorthand, not a quantitative model.
Performance = Harness (Prompts & Tools) + Model (Brain Quality) + Serving (Latency & Precision)Change one layer at a time and repeat the same task. That gives you a better basis for deciding whether a tool change, model change, or hosting change helped.
This discussion was inspired by Roy (@usr_bin_roygbiv on X). The source link appears below.
Source attribution: Roy’s analysis of the harness, model, and serving layers.
