The Harness Is Part of the Product — Not a Substitute for the Model
A useful claim about coding-agent benchmarks hardened into a much larger one this week. The useful claim is that a score belongs to a model-and-harness system: prompts, tools, retrieval, retry logic, context management and loop control all affect the result. The larger claim — that harness choice now eclipses model choice and turns the model into a commodity input — does not follow. Even the paper advancing the "Binding Constraint Thesis" restricts it to long-horizon tasks and models of comparable frontier…