Is ChatGPT Getting Better?
- 4 minutes ago
- 9 min read

Monday 17 August 2026
There is a peculiar difficulty in measuring the progress of artificial intelligence. When motor cars became better, one could measure their maximum speed, fuel consumption, braking distance and reliability. When computers became better, one could count transistors, measure clock speeds and calculate operations per second. But when a machine becomes better at thinking — or, more cautiously, better at performing tasks that human beings associate with thought — the relevant unit of measurement is obscure. Intelligence is not a scalar quantity. A machine may become considerably better at mathematics while remaining irritatingly literal in conversation. It may acquire formidable abilities in computer programming while still inventing a book that does not exist. It may become better at reasoning yet worse at knowing when to stop talking.
Nevertheless the evidence is now becoming difficult to resist. ChatGPT is getting better. Indeed the improvement since the appearance of GPT-4 in 2023 is sufficiently substantial that merely describing successive models as marginal upgrades understates what has happened. The more interesting questions are how it is getting better, whether users can actually notice the improvement and whether some approximate magnitude can be assigned to the change.
The first complication is that ChatGPT is not really a single thing. It is a product within which different models, reasoning systems, search tools, memory mechanisms and other software components may operate. Improvements to ChatGPT therefore arise both because the underlying neural networks become more capable and because the machinery surrounding them becomes more sophisticated. A modern ChatGPT system can search the internet, inspect documents, manipulate files, use software tools, reason for longer before answering and select different computational strategies according to the difficulty of the question. Asking whether ChatGPT has improved is consequently rather like asking whether an army has improved. Better soldiers are part of the answer — but so are better reconnaissance, communications, artillery, logistics and command.
Still, the underlying models have plainly improved. Consider one relatively clean historical comparison. When GPT-5 was introduced in August 2025, OpenAI reported a score of 94.6 per cent on the 2025 American Invitational Mathematics Examination benchmark without tools, 74.9 per cent on SWE-bench Verified — a collection of real software-engineering problems — and 84.2 per cent on MMMU, a demanding multimodal reasoning benchmark. On the coding benchmark SWE-bench Verified, the preceding OpenAI o3 model scored 69.1 per cent. GPT-5 therefore did not merely generate more elegant prose: it solved a measurably greater proportion of difficult problems.
Yet percentage-point improvements can conceal more than they reveal. Suppose an examination is already being passed at 90 per cent. Moving from 90 to 95 per cent sounds like an improvement of only five percentage points, but the error rate has fallen from ten questions to five. In one important sense the machine has become twice as reliable. This distinction becomes increasingly important as artificial intelligence approaches high levels of competence. Progress near the top of a benchmark frequently appears numerically small because there is little remaining headroom. Error reduction may therefore be a more illuminating measure than raw score improvement.
On factual reliability the improvement has at times been striking. OpenAI reported that GPT-5, when using reasoning, produced approximately 80 per cent fewer factual errors than o3 on certain evaluations and around six times fewer hallucinations across particular open-ended factuality benchmarks. With web search enabled on prompts representative of real ChatGPT traffic, GPT-5 responses were reported as about 45 per cent less likely to contain a factual error than GPT-4o, while GPT-5 reasoning responses were about 80 per cent less likely to do so than o3. These figures must not be mistaken for a universal statement that ChatGPT became “80 per cent more accurate”. They concern specified evaluations under specified conditions. Nevertheless a reduction of this magnitude in one of the defining weaknesses of large language models is not cosmetic.
The same trend has continued. OpenAI’s evaluations of GPT-5.5, released in 2026, found that on a deliberately difficult set of conversations previously flagged by users for factual errors, its individual factual claims were 23 per cent more likely to be correct than those of GPT-5.4. The proportion of complete responses containing a factual error fell by a much smaller three per cent, partly because the newer model made more factual claims in each answer. That apparent discrepancy illustrates how treacherous the measurement of artificial intelligence can be. A model that attempts more, explains more and makes more individually correct statements can simultaneously expose itself to more opportunities to make at least one mistake.
There is another measure that may ultimately prove much more significant than examination scores. The independent research organisation METR attempts to estimate an AI system’s task-completion time horizon. The idea is ingenious. Instead of asking whether an artificial intelligence can answer another collection of examination questions, researchers ask how long a task would take a competent human expert and then estimate the duration of tasks that an AI agent can successfully complete with a specified probability. A machine capable only of reliably performing five-minute jobs is plainly less useful as an autonomous intellectual worker than one capable of completing projects that would occupy a human expert for several hours.
On METR’s evaluation, GPT-5 had a 50 per cent task-completion horizon of approximately two hours and seventeen minutes, compared with about one hour and thirty minutes for o3. The confidence intervals are wide and these figures should therefore not be fetishised. Nevertheless METR found GPT-5 ahead of o3 in 96 per cent of its bootstrap comparisons. By April 2026 METR was reporting a point estimate of about 5.7 hours for GPT-5.4 under its standard methodology, although the estimate was unusually sensitive to how attempts involving “reward hacking” were treated. METR itself warns that measurements above sixteen hours are presently unreliable.
This suggests something more profound than incremental benchmark improvement. The frontier may be moving from answering questions towards sustaining competent activity. That is an altogether different kind of progress.
A brilliant person who can concentrate for thirty seconds is of limited professional usefulness. Much of human intellectual achievement consists not of isolated flashes of reasoning but of maintaining a chain of competent decisions: reading a document, identifying a problem, finding the relevant evidence, testing a hypothesis, noticing an error, correcting it and continuing until the objective has been achieved. If the reliable working horizon of artificial intelligence continues to lengthen, its economic significance may grow much faster than its benchmark scores suggest. Moving from a machine capable of performing a ten-minute task to one capable of conducting a day’s work is not merely an eightfold quantitative improvement. It potentially changes the category of activities for which the machine can substitute or augment human labour.
This is one reason why experienced users sometimes perceive improvements in ChatGPT more dramatically than casual users do. Ask successive generations “What is the capital of Peru?” and virtually nothing has changed. Ask them to analyse an intricate legal problem, reconcile contradictory documents, research authorities, draft a memorandum, identify weaknesses in their own conclusions and revise the result, and generational differences become much more apparent. Improvements accumulate across every stage. If a complicated task requires ten dependent reasoning steps, improving reliability at each step from, say, 90 to 97 per cent would increase the crude probability of completing all ten correctly from about 35 per cent to about 74 per cent. The arithmetic is illustrative rather than descriptive of actual ChatGPT performance, but it demonstrates why apparently modest improvements in local reliability can produce enormous improvements in usefulness on long tasks.
This phenomenon may explain the curious experience of using contemporary frontier models. They do not necessarily seem several times more intelligent in every sentence. Rather, they fall apart less often. They retain the thread of an argument for longer. They obey complicated instructions more consistently. They are better at distinguishing evidence from inference. They can use tools more competently. They recover from mistakes more frequently. Increasingly, they can recognise that a task is impossible or that information is missing rather than improvising an answer. OpenAI’s GPT-5 evaluations, for example, found that when images necessary to answer a question had deliberately been removed, o3 nevertheless gave confident answers about the nonexistent images 86.7 per cent of the time; GPT-5 did so only nine per cent of the time. Knowing that one does not know is itself an important component of intelligence.
There is also a subtler qualitative improvement — one much harder to capture in benchmarks — in understanding what the user means. Early large language models were astonishing sentence-completion machines. Increasingly their successors behave more like collaborators constructing a model of the task. The distinction is important. Human instructions are chronically incomplete. Lawyers, academics, businessmen and soldiers rarely specify every assumption underlying what they say because human colleagues supply enormous quantities of contextual inference automatically. An AI system becomes dramatically more useful when it can do likewise — while avoiding the opposite danger of confidently inventing context that was never supplied.
This is where subjective impressions become relevant, notwithstanding their scientific untidiness. Long-term users develop a sense for the characteristic failures of a model. One learns when it is likely to fabricate a citation, misunderstand a qualification, lose track of an instruction or produce superficially impressive nonsense. As those failures become rarer, interaction changes. The user devotes less effort to supervising the machine and more effort to exploiting it. This reduction in supervisory burden may eventually be one of the most economically useful measures of AI progress.
Suppose an AI assistant completes a task in ten minutes but requires twenty minutes of checking. It has saved little. Suppose its successor takes the same ten minutes but requires only five minutes of checking. Its benchmark intelligence might have risen by ten per cent while its practical value has multiplied. The final frontier is not therefore whether an AI can produce excellent work. Contemporary systems already sometimes can. It is whether the user can safely expect excellent work.
There remain formidable qualifications. Benchmarks become contaminated as their contents or characteristic problems enter training data. Developers naturally select evaluations that display their systems favourably. Different models have different strengths. Some apparent advances arise because models are allowed greater computational resources at inference time. Tool use can make a mediocre internal answer look much better. Conversely a powerful model can be crippled by poor tools, restrictive interfaces or inadequate context. There is no equivalent of the metre or kilogram for intelligence and claims that one model is “30 per cent smarter” than another should generally provoke suspicion.
Nor has hallucination disappeared. A model whose error rate has fallen dramatically can become more dangerous in a particular psychological sense because users cease expecting errors. An unreliable witness is treated cautiously; a witness who is correct ninety-nine times and invents the hundredth answer with equal confidence is more treacherous. As AI becomes more reliable, calibration — matching confidence to actual probability of correctness — becomes at least as important as raw intelligence.
There is another paradox. Better models are being asked harder questions. The frontier moves as the machine advances. Nobody celebrates because ChatGPT can summarise a newspaper article anymore; that achievement has become mundane. Users instead ask it to interpret obscure points of appellate procedure, construct software systems, analyse thousands of pages of documents or conduct autonomous research. The subjective impression that “it still gets things wrong” can therefore coexist perfectly well with spectacular objective progress. Humans continually move the examination paper.
So by what magnitude has ChatGPT improved?
There is no intellectually respectable single number. Depending upon the domain and the baseline, measured improvements range from modest percentage-point gains to reductions of 50, 80 or even more per cent in particular categories of error. On autonomous software tasks, independent METR measurements suggest that the duration of tasks frontier models can tackle with substantial reliability has been expanding rapidly — from tens of minutes towards hours — although the estimates become increasingly uncertain at the upper end. On some saturated academic benchmarks, by contrast, improvements now look comparatively small because the machines are approaching the ceilings of the tests.
Qualitatively, however, the change may be larger than any single benchmark indicates. The important transition is from eloquence to reliability, from response to reasoning and from reasoning to sustained agency. The earliest ChatGPT astonished because a computer could converse. The next generation astonished because it could solve difficult intellectual problems. The emerging generation is significant because it can increasingly undertake sequences of intellectual work, use external tools, detect some of its own mistakes and persist towards an objective.
That progression has an obvious destination, although nobody knows how quickly it will be reached. A genuinely transformative artificial intelligence need not possess consciousness, emotions or anything resembling the interior life of a human being. It need only become sufficiently competent, sufficiently persistent and sufficiently reliable that delegating a complicated intellectual project to it becomes as ordinary as delegating that project to a capable colleague.
We are not there yet. The mistakes remain too frequent, the confidence occasionally too misplaced and the boundaries of competence too irregular. But the direction of travel is increasingly measurable. If one insists upon reducing the answer to a sentence, it is this: ChatGPT is not merely becoming better at answering questions; it is becoming better at completing work.
And that is a considerably larger improvement.




