A new model lands, the changelog says it is better at everything, and the migration is one string. Then something answers as the old model because a fallback you forgot about is still wired, or a refusal arrives as a successful 200 and your error handling never sees it. The interesting failures are not in the benchmark table. This forces the boring receipts: prove which model answered, re-run the prompts that already hurt, and only then talk about quality.
Fill in: old_modelnew_modelworst_promptsverify_cmd
Shared as is, for reference. Read it and decide what it will do before you run it; using it is your responsibility, under the terms.