Language models answer a little differently every time. That makes it hard to tell a real change from a bad roll. These steps help you collect evidence that holds up.
On launch day
Pick five to ten tasks from your real work that the new model handles well. Save each prompt exactly, including any files, settings and the model name you selected. Save the answers too, with the date.
Prefer tasks with an answer you can check: code that has to pass tests, a calculation, an extraction from a document. “It feels worse” is hard to compare. “It passed 9 of 10 tests in March and 5 of 10 in May” is not.
Every week or two
Run the same prompts again, in a fresh session, with the same settings. Run each prompt at least three times, because a single answer can be an outlier. Note the date, the time of day and how long each answer took.
When something looks different
Compare the results side by side. Write down what changed: correctness, length, instruction following, speed. If you can, repeat the test through the API with a pinned, dated model version. If the pinned version still behaves like launch day and the app does not, that difference is worth reporting.
Share it carefully
Describe what you measured, not what you suspect. Share prompts and numbers rather than single screenshots. Remove anything confidential first.
Community trackers
A few people measure models continuously. Each one watches one particular setup, not everything a provider serves, so read them as monitors, not as proof. We link only trackers whose method can be checked (reviewed 1 October 2026):
- livenerf measures Claude Opus 5.5 through Claude Code on a subscription, with a pinned tool version, a fixed question set and a statistical plan published before the first data. It is a one-person project, still collecting its baseline; a first verdict is expected around late October 2026, and the daily raw data is not public yet.
- Margin Lab trackers run a daily set of 50 SWE-Bench-Pro coding tasks through Claude Code and Codex and show confidence intervals. Because they always use the latest tool release, a drop can come from the tool as well as from the model. After a model switch the tracker collects a new baseline, and numbers from before and after are not comparable.
Three things to keep in mind with any tracker: a change in the measuring setup can look like a nerf, a switch to a new model resets the comparison, and a single bad day on a small sample proves nothing.
This is exactly why we ask for pinned versions and public changelogs: so that nobody has to do this detective work.