| 1 | Capability benchmarks designed to last years are now lasting months. On SWE-bench Verified — the canonical coding benchmark — frontier models rose from sixty percent to roughly one hundred percent of the human baseline in a single year. Humanity's Last Exam, built explicitly to resist saturation, gained thirty percentage points in twelve months. PhD-level science question-answering, multimodal reasoning, and competition mathematics all crossed human-baseline territory in the same window. The bench-design community is now permanently behind the model-release cycle, and the conversation is shifting from static benchmarks to continuous evaluation harnesses that regenerate problems on every run. |