The GPT-6 Astra benchmarks figure that circulates most widely was read in September, and it reads 52.7 — with a second September reading at 54.7. The harness’s live page for GPT-6 Astra reads 51.83 on 2026-10-07. The model did not change between those dates. The number did.
That two-point drift is smaller than it sounds and larger than it should be. It is smaller than the gap between the top and bottom of this comparison set, which is why nobody noticed it. It is larger than the margin any serious decision should be made on, which is why it matters.
One metric, three readings, two dates
Here is the same measurement written out three times, with nothing rounded off:
| reading | Intelligence Index | read |
| first September reading | 52.7 | September 2026 |
| second September reading | 54.7 | September 2026 |
| live page | 51.83 | 2026-10-07 |
All three are readings of the same metric on the same model, and the spread between the highest and the lowest is 2.87 points — wider than the distance between two of the four models this comparison covers on that index. That single fact is the argument of this article: on this metric, the change in the reading over four weeks is larger than the difference between models that people are choosing between.
Note also the direction. Both September readings sit *above* the live one, so a page written in September is not merely old — it is flattering. Anyone comparing that page against a competitor’s fresh figure is running a race with one runner’s clock set a month fast.
What moves a score without moving the model
A benchmark result is not a property of a model. It is the output of a harness, and every component of a harness can be revised between two runs.
The task set. Evaluations add tasks, retire tasks and rebalance categories. A score computed over a different mix of items is a different measurement even where the metric name is untouched, and the name is almost always untouched.
The configuration. Every number in this cluster was taken at a stated reasoning effort — Max for GPT-6 Astra and GPT-6.1 Sol, Xhigh for Grok 4.7, read 2026-10-07. A reading taken at a different rung is a reading of a different configuration, and it will move without the model moving at all.
The scoring path. Open-ended tasks are graded, whether by a rubric, a judge model or a human panel. Revise the grader and every historical score computed with it becomes a score under a different rubric.
Published limits. A serving stack’s advertised context or output ceiling can change after a launch. Where an evaluation occupies a long window, a change to that window changes the result.
None of these is exotic, and none is announced with the number. The reading simply appears on the page, next to a fresh date, and a reader comparing it against a quotation from four weeks earlier has no way to see that the measurement rather than the model is what changed.

The digits that collide
There is a trap in the numbers above that deserves its own section, because it is the reason a figure without its metric label is close to useless.
54.7 is not one number in this cluster. It is a September Intelligence Index reading for GPT-6 Astra, and it is also that model’s pass rate on a separate long-form reasoning exam on the same board — 0.5468, rounded, read 2026-10-07. Same model, same digits, two unrelated measurements with two different scales.
Read the two side by side and the problem is obvious: one is an aggregate index in the low fifties, the other is a percentage of exam questions answered. Compare them as though they were the same quantity and the arithmetic looks meaningful while being nonsense. That is not a hypothetical failure — it is the ordinary way a benchmark table degrades as it travels, one hop at a time, from a labelled column to an unlabelled bullet in a slide deck.
The rule that follows is cheap and it prevents the whole class of error: carry the metric name, the reading date and the configuration rung in the same sentence as the number, or do not use the number. “54.7” is not a fact. “Astra’s pass rate on the long-form reasoning exam, at Max effort, read 2026-10-07” is.
What a moved reading does and does not mean
Three things it does not mean, all of them commonly asserted:
• It does not mean the model regressed. Nothing was shipped. The measurement was redone.
• It does not mean September’s pages were wrong. 52.7 was the right reading for September. It is a dated fact, not an error, and treating it as an error is as bad as treating it as current.
• It does not mean the newer reading is the final one. A number that moved down can move back up on the next revision. The live figure carries 2026-10-07 for exactly that reason.
What it does mean is narrower and more useful: a benchmark figure has a shelf life, and the shelf life is shorter than the interval on which most content is refreshed. A page written on launch day and left alone is quoting a measurement that has probably already been revised.
Three checks before you quote a number off a page
Find the read date first. If the page does not carry one, the number is undated and therefore unusable for comparison. Every figure in this cluster carries the day it was read, and this section of the series is the reason.
Find the metric name and the effort rung. “Astra scores 54.7” could be two different things at two different scales, and the score itself is a configuration as much as a model. Without both labels the figure cannot be compared to anything.
Check whether the page predates the last revision. A dated page from September is a September measurement. Where the live reading has moved, the honest thing to do is quote both and give both dates, not pick the flattering one.

The takeaway
GPT-6 Astra’s Intelligence Index has been read at 52.7, then 54.7, and now reads 51.83 on the harness’s live page, read 2026-10-07 — a 2.87-point spread on one model, wider than the gap between two models in the same comparison. The model did not move; the measurement did.
Two habits defuse the whole problem. Quote a number only with its metric name, its reading date and its effort rung attached, because 54.7 on this model is both an index reading and a separate exam pass rate and the digits alone cannot tell you which. And where a page is older than the last revision, quote both readings with both dates rather than the one that reads better.
A benchmark figure is a reading, not a property. Treating it as a property is how a four-week-old page ends up quietly overstating a model to everyone who copies it.
OrcaRouter carries GPT-6 Astra at list rate on the same key as the rest of the OpenAI line, so re-running your own comparison after a revision is a model-string change rather than a second integration.
Sourcing note: the Intelligence Index readings for GPT-6 Astra — 52.7, 54.7 and the live 51.83 — are Artificial Analysis’ own published values, read 2026-10-07; the earlier readings circulated in September 2026 and are quoted here as dated readings rather than as errors. The long-form reasoning pass rate of 0.5468, and the Max reasoning setting recorded for GPT-6 Astra and GPT-6.1 Sol against Xhigh for Grok 4.7, are from the same publisher’s live pages, read the same day. The observation that the spread across the model’s own readings exceeds the gap between two models in this comparison, the list of harness components that move a score, and the two-reading rule are this article’s own analysis of those published figures, not published claims. No vendor claim is cited in this article. All checked 2026-10-07; re-check before 2026-11-01.









































































