Best Coding LLM: Choose by Task, Then Govern the Workflow
The best coding LLM depends on the task, constraints, and workflow around it. Compare coding strengths, then decide what should run as deterministic software instead of being re-reasoned.

The short answer
There is no best coding LLM in general, only a best one for a given job under given constraints. Match the model to the task: a frontier reasoning model for long-horizon work across a repository, a mid-tier model for routine generation and review, a fast small model for high-volume triage. Then test candidates against your own code, and wrap the winner in tests, permissions, and logs.
That last part is where most teams lose. Model leadership rotates every few weeks. The workload split underneath it barely moves, and neither does the review, routing, and approval work that surrounds every accepted change.
What a coding LLM is actually good at
Four capabilities get bundled under "coding" and they fail in different ways.
Generation is writing new code against a spec. Models are strong here when the target is a well-documented library and weak when the target is your internal framework with no public examples. Explanation is summarizing unfamiliar code, which is the most reliable of the four because the source of truth sits in the context window. Debugging is the hardest, because it needs the model to reason about runtime state it cannot observe. Give it a stack trace and a failing test and it does well. Give it "the job hangs sometimes" and it guesses.
Refactoring sits in between. A model can apply a mechanical rewrite across many files accurately, and it will quietly drop an edge case if the rewrite is not actually mechanical. That is the failure mode to design around.
How to compare coding models
Public benchmarks tell you how a model performed on a fixed, mostly open-source task set. Your repository is not that task set. Build your own.
Pull 30 to 50 tasks from your merged pull request history, weighted toward the work your team actually does. Include the boring ones. For each task, record the starting commit, the prompt, and the diff a human shipped. Then run every candidate model against the same set and score on five things:
- Correctness. Does the generated change pass the existing test suite without modifying tests?
- Regression safety. Run the full suite, not the tests touching the diff. Silent breakage elsewhere is the expensive failure.
- Scope discipline. Compare the size of the model diff to the human diff. Large overshoot is a signal the model is rewriting rather than fixing.
- Latency at your context size. Measure with your real repository payload, not a short prompt.
- Cost per accepted change, not cost per call. A cheaper model that needs three attempts is not cheaper.
Add a sixth criterion if you are in a regulated environment: whether the vendor's data handling and residency terms clear your policy. That one is binary and it eliminates candidates before any quality score matters.
Re-run the set when you change models. Scores from a previous generation do not carry forward, and a model swap is a production change even when nothing in your code moved.
Best coding LLM by job
The table below pairs common engineering jobs with candidates whose vendors position them for that work. The evidence column points at official documentation, because that is the only claim about a model that is current by construction. Everything in the limitation and governance columns is editorial judgment from working with these workflows.
| Job | Candidate positioned for it | Evidence source | Limitation | Governance note | | --- | --- | --- | --- | --- | | Long-horizon agentic work across a repo | Claude Fable 5.1, positioned for demanding reasoning and long-horizon agentic work | Anthropic models overview | Slowest tier by the vendor's own latency comparison; wrong choice for anything interactive | Long runs touch many files. Scope the repository credential to the branch, not the org | | Complex agentic coding on enterprise codebases | Claude Opus 5, positioned for complex agentic coding and enterprise work | Anthropic models overview | Strong general choice, which makes it easy to over-apply to jobs a cheaper tier handles | Log the model ID with every accepted diff so you can attribute a bad change later | | Difficult end-to-end implementation | GPT-6 Astra, recommended for complex reasoning and coding | OpenAI models docs | End-to-end output is harder to review than incremental output. Reviewer fatigue is real | Require a human approval event before merge, recorded with the reviewer identity | | Software engineering and multi-step agent workflows | Gemini 3.8 Flash and Gemini 3.7 Flash, positioned for software engineering, agentic workflows, and multi-step execution | Gemini API models | Vendor positioning is not evidence of fit on your stack. Run the test set | Multi-step runs need per-step logging or you cannot reconstruct what happened | | High-volume, latency-sensitive triage | Claude Haiku 4.5 or GPT-5.6 Luna, positioned for speed and cost-sensitive high-volume work | Anthropic, OpenAI | Near-frontier is not frontier. Use for classification and routing, not for authoring changes | Cheap models get called constantly. Rate-limit at the app layer, not in the prompt | | Code that cannot leave your environment | An open-weight model you host yourself | Your security and data-residency policy | Self-hosting moves the cost from tokens to infrastructure and evaluation effort | The strongest control available, and the one that most often decides the shortlist | | Reviewing a diff for a specific class of defect | Any mid-tier model with a narrowed prompt | Your own test set results | Broad "review this PR" prompts produce broad, low-signal output | Store findings as rows, not chat messages, or nothing accumulates |
Two things this table deliberately does not do. It does not rank the models against each other, because the ordering would be wrong within a quarter. And it does not print benchmark scores, because a score measured on someone else's task set is not evidence about yours. For a closer read on model-level coding strengths, see our comparison of coding models.
Put the model in a governed coding workflow
Take pull request triage. The job is to look at an incoming PR, decide whether it needs a senior reviewer, flag anything that touches auth or billing, and route it. Most teams build this as one long prompt and re-run the whole thing on every PR.
Watch what that costs. The prompt refetches the diff, re-reads the routing rules, re-derives which paths are sensitive, re-decides the reviewer, and produces a fresh judgment every time. Only one of those steps needs a model.
Split it. Fetching the diff is an API call. Matching changed paths against a sensitive-path list is a lookup. Picking the reviewer from a rotation is a function. Recording the decision, the approval, and the timestamp is a database write. What actually needs judgment is reading the diff and deciding whether the change is riskier than its size suggests. That is one bounded model call against a narrow prompt, and it returns a structured verdict rather than prose.
The app holds the state. A triage record looks like this:
{ "pr": 4821, "repo": "billing-service", "sensitive_paths": ["src/billing/refund.ts"], "model_verdict": "elevated_risk", "model_id": "claude-opus-5", "routed_to": "senior-review", "approved_by": "j.okafor", "approved_at": "2026-09-08T14:22:07Z"}
That row is the audit event. It survives the conversation, it names the model that produced the verdict, and it makes the approval attributable to a person. When someone asks in six months why a refund path shipped without senior review, the answer is a query rather than an archaeology project. This is the difference between a coding agent that acts and one you can account for.
The economics follow the same split. The first run pays for the model to work out the triage logic. After the deterministic parts move into the app, every later run pays only for the bounded judgment call, so token cost is front-loaded and then close to flat as PR volume grows. Re-reasoning the whole workflow on every PR gives you the opposite curve.
What this article doesn't cover
Model-specific prompt engineering, fine-tuning on a private codebase, and inline completion latency tuning in the editor are separate problems with separate literature. This article is about selection and the workflow around the selection. It also assumes you have a test suite worth running. If you do not, no model choice fixes that.
The Major take
Even the best model output goes through the same repeatable sequence before it reaches production: fetch context, apply policy, run tests, route for review, capture approval, record what happened. That sequence does not benefit from judgment. It benefits from running identically every time. When it lives inside a prompt, it is re-derived on every execution, it costs tokens on every execution, and it leaves no record you can audit.
Major resolves that by letting the agent build the app instead of repeating the reasoning. The agent works out the triage logic once, then builds a PR review app for it: a managed database holding review state and verdicts, scoped repository credentials applied where the agent acts, role-based access on who can approve, and logs written as the work happens. From then on the agent runs the app and spends model calls only on the risk judgment. Reason once, run forever.
The honest scope. If you are one developer using a model in your editor, you do not need any of this, and you should not build it. The case turns when multiple people depend on the output, when a change touches something regulated, or when the same workflow runs often enough that re-reasoning it becomes the dominant cost. At that point the model choice matters less than whether the surrounding work is code you can inspect, which is also the point where agent governance stops being a policy document and becomes an architecture question.
Pick your model on the test set. Build the workflow around it as software. If you want to see the second half working, describe your PR triage rules to Major and let it build the review app your team routes through. Build your PR triage and review app on Major.
Related articles
Frequently asked questions
- What is the best coding LLM?
- No single model wins every coding job. Frontier reasoning models suit long-horizon work across a repository, mid-tier models handle routine generation and review, and fast small models fit high-volume triage. Pick by matching the task to vendor-documented strengths, then confirm the choice against 30 to 50 tasks from your own merged pull requests. Model leadership rotates; your selection method should not.
- Which LLM is best for writing code?
- For writing code specifically, choose a model whose documentation positions it for complex or agentic coding, then check three things on your repository: whether it handles your internal frameworks without public examples, whether its diffs stay close in size to human diffs, and whether generated changes pass your full suite untouched. Vendor positioning narrows the shortlist. Your test set decides it.
- How do you evaluate a coding LLM?
- Build a task set from merged pull requests, recording the starting commit, prompt, and shipped diff for each. Score candidates on correctness against untouched tests, regression safety across the full suite, scope discipline versus the human diff, latency at your real context size, and cost per accepted change rather than per call. Data residency terms are a separate pass or fail gate.
- Can a coding LLM build production software?
- It can produce production code, but reaching production safely takes more than generation. The change needs tests, policy checks on sensitive paths, a routed review, a recorded human approval, and a durable log of what shipped and which model produced it. Run those steps as deterministic application code with scoped credentials and stored state, and reserve the model for judgment calls.