top of page

The Contest of the Machines: ChatGPT and the New Generation of Artificial Intelligence

  • 2 minutes ago
  • 9 min read

By Matthew Parish


Monday 24 August 2026


The difficulty with writing comparisons of large language models is that they become obsolete with extraordinary speed. Only a short time ago, serious comparisons revolved around GPT-4, Claude 3 and Gemini 1.5. By August 2026, the competitive landscape has moved several generations beyond them. OpenAI is deploying GPT-5.6 in ChatGPT, Anthropic’s principal Claude family has advanced through Claude 4 and beyond, Google’s Gemini family has reached the Gemini 3 generation and xAI has developed Grok into a serious competitor. Meanwhile Meta’s Llama project remains important for a different reason: it represents one of the principal alternatives to the closed proprietary model.


Comparisons have also become more difficult because the expression “large language model” is increasingly inadequate. The leading systems no longer merely predict and generate text. They reason for variable periods, search external sources, analyse images, manipulate files, write and execute computer code and increasingly undertake sequences of actions using tools. Google is pushing particularly heavily into multimodality, including Gemini Omni, while OpenAI describes GPT-5.6 as directed towards end-to-end knowledge work. Grok 4.6 is explicitly designed for long-running agents and complex tasks extending across multiple stages.


We are therefore no longer comparing electronic conversationalists. We are beginning to compare embryonic general-purpose intellectual working environments.


ChatGPT: the accomplished generalist


⁠ChatGPT and OpenAI remains the standard against which other general-purpose artificial intelligence products are commonly judged, partly because it established the modern consumer market for conversational artificial intelligence and partly because OpenAI has persistently attempted to combine reasoning, writing, multimodality and tool use within a single product.


As of August 2026, GPT-5.6 is at the centre of OpenAI’s frontier offering. GPT-5.6 Sol is available to Plus and Pro users in ChatGPT, while GPT-5.6 Luna has become the lighter default offering for Free and Go users. OpenAI has also introduced greater user control over the amount of computational reasoning devoted to an answer. The underlying philosophy is significant. Instead of forcing the user to choose continually between a quick conversational model and a laborious reasoning model, ChatGPT is increasingly becoming a system capable of deciding, or being instructed, how much intellectual effort a particular problem deserves.


Its greatest strength is consequently not that it is invariably the best model at every individual task. Almost certainly it is not. Its advantage is breadth. ChatGPT is unusually competent at moving between activities that human professionals regard as quite different: legal analysis, computer programming, mathematical reasoning, document preparation, historical research, image interpretation, correspondence and sustained argumentative prose. GPT-5.6 has been presented by OpenAI specifically as an advance in knowledge work, science, cyber capabilities and design.


There is another, less easily measured advantage. ChatGPT is particularly effective as an intellectual collaborator. A good conversational model must do more than answer questions correctly. It must understand what sort of answer the user is trying to obtain. It must distinguish between requests for explanation, criticism, drafting, investigation and speculation. It must retain stylistic instructions and contextual assumptions while remaining capable of challenging them where appropriate. On these softer qualities, which conventional benchmarks measure poorly, ChatGPT remains exceptionally strong.


Its weakness is partly the consequence of that versatility. A general-purpose assistant is necessarily constrained by safety systems, product design and the need to serve an enormous variety of users. It can sometimes be overcautious. Like every contemporary model, it can still hallucinate. Fluency remains dangerous because an elegantly expressed falsehood can appear more authoritative than a hesitant truth. The enormous progress of the technology has diminished this problem without eliminating it.


Claude: the scholarly rival


⁠Anthropic has developed Claude according to a somewhat different intellectual temperament. Claude has long distinguished itself through extended textual analysis, programming and carefully structured reasoning. By August 2026 the older Claude Opus 4.1 has been retired from Anthropic’s API in favour of substantially newer models, with Anthropic identifying Claude Opus 4.8 as its replacement. The rapidity of this succession itself illustrates how little value should be attached to permanent league tables of artificial intelligence models.


Claude’s distinctive virtue remains its seriousness. It is often excellent when confronted with a large body of material that must be read carefully before conclusions are drawn. Lawyers, programmers, academics and analysts can therefore find Claude particularly attractive. It tends towards orderly argument and is often good at identifying qualifications to propositions that another model might state too confidently.


There is nevertheless something recognisably “Claude-like” about Claude’s prose. It can be earnest, cautious and occasionally over-structured. What constitutes a virtue in analysing a contract may become a vice in writing an essay, satire or polemic. This distinction should not be exaggerated, because all the leading models have become substantially more adaptable. Yet model personality remains real. Different systems presented with identical prompts frequently produce recognisably different kinds of prose.


For rigorous textual analysis, Claude consequently remains one of ChatGPT’s most formidable competitors. There will be tasks for which an experienced user reasonably prefers it.


Gemini: Google’s enormous advantage


The strategic strength of ⁠Google Gemini lies not merely in the quality of Google’s models but in the Google network surrounding them. Google possesses Search, Android, YouTube, Gmail, Maps, Workspace and an extraordinary collection of research and computational infrastructure. Artificial intelligence can therefore be integrated into activities that hundreds of millions of people already undertake every day.


Gemini itself has advanced considerably. Google introduced Gemini 3.1 Pro in February 2026 as a natively multimodal reasoning model directed towards difficult problems, while subsequent Gemini development has continued rapidly. Google’s more ambitious Gemini Omni project points towards an artificial intelligence in which text, images, sound and video cease to be separate categories of information. Omni can accept combinations of these media and, initially, produce and edit video conversationally.


This may ultimately prove enormously important. Human cognition is not fundamentally text-only. We see, hear, speak and manipulate physical objects. An artificial intelligence that treats these different forms of information as parts of one representation of the world may acquire capabilities qualitatively different from those of the original language models.


Gemini’s historic weakness relative to ChatGPT and Claude has often been less about raw capability than about personality and consistency. A model may perform spectacularly on a benchmark yet feel less satisfactory during an hour-long intellectual conversation. Google has steadily reduced that distinction. Gemini must now be regarded as a frontier competitor rather than merely Google’s answer to ChatGPT.


Its greatest prospective advantage may be distribution. If artificial intelligence becomes embedded invisibly throughout ordinary computing, the winner may not be the model that users consciously decide is cleverest. It may be the system already present in their telephone, browser, email, documents and search engine.


Grok: the increasingly serious outsider


Any contemporary comparison should now include ⁠Grok. Earlier versions of Grok were sometimes perceived principally through their association with X and Elon Musk. That is increasingly an inadequate description of the project.


Grok 4.6, released in August 2026, is explicitly designed for long-running agents, coding, research and complex interactive and visual work. xAI describes the model as capable of remaining engaged with tasks extending across many steps, including research and work across large codebases. Its 500,000-token context window and adjustable reasoning effort demonstrate how far the architecture of contemporary AI products has travelled from the simple question-and-answer chatbot.


Grok also possesses an interesting cultural characteristic. It has deliberately been positioned as less conventional in tone than some competitors. This can make it refreshing, particularly where users dislike excessively sanitised corporate prose. Yet irreverence and epistemic reliability are different virtues. The long-term test for Grok will be whether it can combine its distinctive personality with the extraordinary reliability demanded of an artificial intelligence entrusted with serious professional work.


Nevertheless, it belongs in the first rank of competitors in a way that was considerably less obvious during the first years of the generative AI revolution.


Llama and the importance of open models


⁠Meta AI and Llama occupy a different strategic position. Llama’s significance has always extended beyond the question of whether Meta’s latest model defeats GPT or Claude on a particular benchmark. Meta’s central contribution has been to make powerful model weights available for developers to adapt, fine-tune and deploy themselves.


Llama 4 introduced natively multimodal open-weight models in Scout and Maverick. This matters because there are many circumstances in which control is more important than possession of the absolute frontier model. Governments may not wish to place confidential material in an external commercial service. Companies may want models operating on private infrastructure. Researchers may wish to modify a model in ways impossible with a closed API.


Open models also provide an important competitive restraint upon the proprietary laboratories. If sufficiently capable artificial intelligence can be operated locally, the leading companies cannot assume that access to machine intelligence will remain permanently concentrated in a handful of subscription services.


The price of this freedom is that Llama is not directly comparable with ChatGPT. ChatGPT is an integrated product; Llama is, to a substantial extent, technological infrastructure from which products can be built. A carefully configured open model may be superb at a specialised task while a casually deployed version may be substantially inferior to a frontier commercial assistant.


What do the comparisons actually tell us?


The temptation is to construct a league table. That exercise is increasingly misleading. Models now vary not merely in intelligence but in context length, reasoning time, search facilities, tool access, multimodality, memory, latency, cost and integration with external software. The “best model” may therefore change halfway through a working day.


A lawyer analysing several thousand pages of disclosure might prefer Claude for one stage of the work and ChatGPT for another. A researcher deeply embedded in Google’s ecosystem might find Gemini overwhelmingly more convenient. A programmer building autonomous agents might discover particular advantages in Grok. A defence contractor or financial institution requiring complete control over its infrastructure might prefer a customised Llama deployment even if a proprietary frontier model performs better on abstract tests.


There is also the inconvenient question of benchmarks. Artificial intelligence companies naturally publish measurements on which their new models perform impressively. Benchmarks are useful but they are not identical to intelligence. They are especially poor at measuring judgement, taste, intellectual curiosity, conversational adaptability and the capacity to recognise what a human interlocutor really wants.


These qualities become more important as the models become more capable. When every leading model can summarise a document competently, the interesting question becomes which one notices the strange paragraph on page 437 that changes the interpretation of everything preceding it.


A tentative ranking


Any ranking in August 2026 should therefore be provisional. Nevertheless, if forced to make qualitative judgements, I would place ChatGPT and Claude in the foremost group for sustained intellectual and professional work, with Gemini extremely close and potentially superior for particular multimodal and ecosystem-dependent applications. Grok has developed into a genuine frontier competitor and deserves considerably more serious consideration than its earlier reputation sometimes attracts. Llama occupies a separate category because its greatest virtue is openness and adaptability rather than the uniform experience of a centrally operated assistant.


Amongst them, ChatGPT probably remains the most accomplished all-rounder. Its particular strength is not dominance of every benchmark but the combination of reasoning, writing, research, multimodal capabilities, tool use and conversational adaptability within one environment. OpenAI’s August 2026 updates to GPT-5.6 have concentrated expressly upon factual reliability, focus and controllable reasoning effort, while its new Ultrafast service demonstrates another emerging frontier: making sophisticated reasoning sufficiently rapid for interactive professional applications.


Claude may still have an advantage for certain kinds of prolonged textual scrutiny and austere analytical work. Gemini has perhaps the greatest structural advantage because Google can integrate artificial intelligence into an unparalleled existing digital ecosystem. Grok may prove particularly formidable if long-running autonomous agents become the next decisive battleground. Llama and its successors remain indispensable because they ensure that the future of machine intelligence is not necessarily synonymous with dependence upon a small number of closed systems.


Yet there is a larger conclusion. The differences between the leading models are becoming less important than the difference between all of them and their predecessors.


The first generation of generative artificial intelligence was impressive because machines could converse. The second was impressive because they could reason. The emerging generation is significant because the systems can increasingly research, create, use tools and pursue objectives through sequences of actions. The transition is from answering questions towards performing intellectual work.


This makes the competition between OpenAI, Anthropic, Google, xAI, Meta and others more than a contest over which company produces the cleverest chatbot. They are competing over the architecture through which human beings may increasingly undertake intellectual labour. The ultimate victor may therefore not be the model with the highest examination score. It may be the system that best learns how to complement human judgement: knowing when to calculate, when to search, when to doubt, when to create and, perhaps most importantly, when the human being using it has asked the wrong question.


For the moment, ChatGPT probably retains the crown as the most convincing general-purpose intellectual collaborator. But the margins are narrow, the competitors are formidable and the crown has never been less secure. In artificial intelligence, six months is becoming an historical epoch.

 
 

Note from Matthew Parish, Editor-in-Chief. The Lviv Herald is a unique and independent source of analytical journalism about the war in Ukraine and its aftermath, and all the geopolitical and diplomatic consequences of the war as well as the tremendous advances in military technology the war has yielded. To achieve this independence, we rely exclusively on donations. Please donate if you can, either with the buttons at the top of this page or become a subscriber via www.patreon.com/lvivherald.

Copyright (c) Lviv Herald 2024-25. All rights reserved.  Accredited by the Armed Forces of Ukraine after approval by the State Security Service of Ukraine. To view our policy on the anonymity of authors, please click the "About" page.

bottom of page