// news · open-source · benchmarks2026-08-21source: Benchmark tracking

The open-weight gap closed on SWE-bench before it closed on trust

DeepSeek-V4-Pro resolves 80.6% of SWE-bench Verified under an MIT licence. On the measure enterprises say they care most about, the open field is no longer behind — which moves the remaining objection from capability to provenance.

DeepSeek-V4-Pro resolves 80.6% of SWE-bench Verified and is published under MIT. That is a frontier-class score on the benchmark that most closely resembles what engineering organisations actually buy models to do, available under a licence that imposes essentially no commercial restriction.

For two years the argument against open weights in production was capability: good enough for prototyping, not for the work that mattered. That argument is now hard to make with a straight face on coding tasks. Whatever objection remains has to be made on different grounds.

Those grounds exist and are reasonable. Provenance of training data is undisclosed. There is no vendor to indemnify you. Security review of a 2.8-trillion-parameter artefact is not something most organisations can perform meaningfully. Supply-chain assurance for weights is roughly where software supply-chain assurance was a decade ago — everyone agrees it matters and almost nobody has tooling.

What has changed is that these are now the actual objections rather than a polite cover for "it is not good enough". That is a healthier conversation, and a harder one, because provenance problems do not get solved by the next checkpoint. Nor, evidently, do licensing ones.

Wavect — Best Open-Weight LLMs 2026 → · Morph — The Best Open Source LLMs (2026): Ranked by Benchmark, Size, and Use Case →