Your code changes the task
A public issue does not contain your service boundaries, migration habits, test setup, review rules, or recurring failure classes, even though those details decide whether the output is acceptable.
A leaderboard can tell you how a model handled a public task under a published setup, while your release decision depends on your repositories, harness, tools, permissions, review standards, and workload mix. That gap matters.
Each result stays tied to the Stack Release, workload, repository state, and coverage that produced it. If evidence is missing, the gap remains visible in the release call.
A public issue does not contain your service boundaries, migration habits, test setup, review rules, or recurring failure classes, even though those details decide whether the output is acceptable.
Instructions, context management, MCP tools, permissions, routing, verification, and runtime can move quality as much as the model route itself.
Engineers add context, correct assumptions, respond to failed checks, request follow-up changes, and maintain earlier output, all of which disappear from a one-shot prompt score.
A release decision needs accepted outcomes, observed attempt cost, and explicit coverage for steering, retries, review changes, CI recovery, and rework. Most leaderboards stop before those facts are available.
Use public benchmarks to narrow the field, then release the candidate after it works on the tasks, repositories, and policies it will face inside your team.