Researchers warn the falling 'alignment tax' is outpacing alignment science
The cost of making a model safer has dropped at most major labs — normally good news. The warning attached is that alignment science is not accelerating at the same rate as capability, so cheap mitigation is being applied to systems the field understands less well each quarter.
A falling alignment tax means safety techniques cost less capability than they used to. That is genuine progress and removes the commercial excuse for skipping them. The concern is what the cheapness encourages: applying known mitigations broadly rather than investing in understanding the systems being mitigated.
The gap compounds. Capability advances arrive on a release cadence measured in weeks; understanding advances at the speed of research. Techniques that work today were validated on models two generations old, and nothing guarantees they transfer.
It also reframes what counts as a safety investment. Cheap mitigation is necessary and no longer the constraint. The scarce resource is the science that says which mitigations will still work on the next model, and that is not something a lower alignment tax buys.
Claude 5 Hub — AI safety 2026: alignment research breakthroughs → · 6G-AI — The alignment problem in 2026 →