A technical agenda for AGI safety and security, written to be argued with
A structured account of technical approaches to AGI safety and security — the kind of document whose value is that it commits to specific positions other researchers can disagree with precisely, rather than surveying everything and concluding more work is needed.
The safety literature has no shortage of surveys. What it has been short of is documents that take positions specific enough to be wrong, which is the precondition for a field making progress rather than accumulating citations.
An agenda's usefulness is measured by what it excludes. Naming which threat models are being addressed, which mitigations are considered promising, and which problems are treated as out of scope gives other groups something to contest — and contested agendas are how research communities allocate effort without a central planner.
The security framing alongside safety is doing real work. Alignment failures and adversarial attacks have been studied by largely separate communities with different assumptions, and deployed systems experience both, often together. A document treating them as one problem surface is describing the world more accurately than either literature alone.
The recurring difficulty is that technical agendas are written against capabilities that do not exist yet, so their threat models are forecasts. That is unavoidable and worth stating: the alternative is waiting, and the empirical results now arriving suggest the forecasts have not been wild.
arXiv — An Approach to Technical AGI Safety and Security → · arXiv — Frontier AI Risk Management Framework in Practice v1.5 →