How Long Is the Memory of a Large Language Model?

By Matthew Parish
Tuesday 6 October 2026
One of the more peculiar experiences of talking regularly to a frontier large language model is that it can seem to know you. You begin a conversation about Kant, return to the subject of Ukrainian politics, ask for an obscure point of American civil procedure and then mention something you discussed weeks ago. Sometimes the machine responds as though all these things form part of a continuous acquaintance. At other times it stares at you, metaphorically speaking, with the blank expression of a hotel receptionist confronted by a guest who insists that they had a fascinating conversation together last Tuesday.
This raises an apparently simple question with a surprisingly complicated answer. How long is the memory of a typical frontier large language model?
The shortest answer is that an LLM does not have a memory in quite the sense that human beings do. It has several different mechanisms that can produce something resembling memory, and these mechanisms operate over radically different timescales. One may last for seconds, another for the duration of a conversation, another potentially for years and yet another is buried so deeply in the model that it is more appropriately called knowledge than memory. Confusing these mechanisms is responsible for much of the popular misunderstanding about what artificial intelligence does when it appears to remember us.
The goldfish myth
It used to be common to describe large language models as possessing the memory of a goldfish. Each interaction was essentially self-contained. Close the conversation and the machine forgot you. Even within a conversation, sufficiently long exchanges could cause earlier material to disappear beyond the model’s effective field of attention.
That description is increasingly obsolete for frontier AI systems. Contemporary systems may combine large context windows with retrieval mechanisms, persistent user information, summaries of earlier conversations and access to external stores of information. The resulting architecture can create something substantially closer to continuity.
Nevertheless, the underlying model remains rather strange. A language model does not ordinarily possess an autobiographical stream running quietly in the background. When nobody is talking to it, it is not sitting somewhere thinking about the conversation it had with you yesterday. There is no little electronic Marcel Proust reclining inside the data centre, involuntarily remembering the computational equivalent of madeleines.
Instead, relevant information is assembled when required.
The first kind of memory: the context window
The most straightforward form of LLM memory is known as the context window. This is the quantity of information that can be presented to the model while it generates an answer.
Imagine placing a very large pile of papers on somebody’s desk and saying: “Everything you need to know is somewhere in this pile.” The person’s ability to answer your question depends partly upon how much material fits on the desk and partly upon whether he can locate the important passage amongst several thousand pages of irrelevant material.
The same broad principle applies to an LLM.
Context windows have grown enormously. Early widely used transformer systems dealt comfortably with only a few thousand tokens. Frontier systems subsequently moved into tens of thousands, hundreds of thousands and, in some architectures or configurations, considerably more. A token is not precisely a word: in English it is typically a word or part of a word. Therefore token counts should never casually be translated into equivalent numbers of pages without qualifications.
Yet raw context length is only half the story. A model theoretically capable of accepting an enormous context does not necessarily use every portion of it equally effectively. Important information can become harder to retrieve when submerged in vast quantities of material. Researchers sometimes call variants of this phenomenon the “lost in the middle” problem. Having a million things available to remember is not equivalent to remembering a million things perfectly.
Humans, incidentally, suffer from something remarkably similar. Anybody who has practised law will recognise the distinction between possessing a 4,000-page trial bundle and knowing what is in it.
Conversation is not memory
This distinction explains another source of confusion. During a long conversation, a large language model may appear to possess excellent memory because the previous exchanges are being supplied, directly or indirectly, as context for the next response.
Suppose you tell an LLM that your fictional dog is called Aristotle. Twenty messages later you ask: “What is my dog called?” If the original exchange remains available within the conversational context, answering “Aristotle” does not require anything resembling long-term memory. The machine is effectively consulting the transcript.
That is closer to an open-book examination than recollection.
The distinction becomes important because conversations can exceed context limits or be compressed. AI systems may summarise earlier portions of an exchange rather than repeatedly supplying every word. Once that happens, details can disappear. If a summary says merely that “the user discussed his pets”, Aristotle may vanish from computational history.
This produces one of the characteristic oddities of LLM interaction. The machine can remember an extraordinarily sophisticated argument about Wittgenstein while forgetting the name of your dog. That is not necessarily because Wittgenstein matters more to it. It may simply reflect what survived the processes by which conversational history was selected, compressed or retrieved.
Persistent memory changes everything
A different phenomenon arises when an AI service maintains information across conversations. This is what most people intuitively mean when they say that ChatGPT, Claude, Gemini or another assistant “remembers” them, although the precise implementations differ between products and change rapidly.
Here the relevant information need not remain inside the active context window indefinitely. Instead, something derived from previous interactions can be stored separately and later retrieved when relevant. The model might therefore know that a particular user prefers British English, is writing a novel about nineteenth-century Constantinople or does not want coriander in recipes without the user having to repeat those facts every morning.
In principle, this kind of memory can last indefinitely. Its duration is no longer principally determined by the transformer’s context window. It depends upon the surrounding software: what gets retained, how it is represented, when it is retrieved, whether it is subsequently updated and what controls the user has over it.
This makes asking “How long is an LLM’s memory?” rather like asking how long a human being’s paperwork lasts. The answer depends upon whether we mean the documents currently lying on the desk, the contents of the filing cabinet or the things the person has actually learned.
Retrieval is the real revolution
The increasingly important concept is therefore not memory alone but retrieval.
Suppose an AI system has access to ten years of correspondence, documents and conversations. It would be computationally absurd to place every word of this material into the prompt every time somebody asks what they are having for dinner. Instead, another mechanism searches for information likely to be relevant to the immediate question.
The model consequently operates with a reconstructed present. Before answering, the surrounding system can assemble pieces of relevant history and place them before the model. From the model’s perspective, these materials become part of what it presently knows.
There is something surprisingly human about this arrangement. Human recollection is not a continuous replay of everything that has ever happened to us. Memories are reconstructed in response to cues. We remember one thing because another reminds us of it. We forget things that subsequently return when somebody mentions a name, shows us a photograph or takes us back to a familiar street.
LLM memory and human memory remain profoundly different mechanisms. Nevertheless, both illustrate an important philosophical point: memory is useful only insofar as something can be retrieved from it.
A library containing every book ever written would be almost useless if nobody had invented a catalogue.
Then there is the model itself
There is another sense in which a large language model has an astonishingly long memory. Its parameters encode statistical regularities acquired during training.
Ask a frontier model about Julius Caesar, Maxwell’s equations or the plot of Hamlet and it may answer without conducting an external search. In some loose sense, it has “remembered” information encountered during training.
But calling this memory can be misleading. A trained neural network does not generally contain a searchable archive of its training documents from which it simply retrieves the relevant paragraph. Information has been transformed into distributed numerical relationships among parameters. The distinction resembles, very imperfectly, the difference between remembering the exact page of a school textbook and simply knowing that Paris is the capital of France.
This form of knowledge can survive for the entire lifetime of a particular model version. Yet it has another peculiar characteristic: ordinarily the model does not rewrite its fundamental parameters every time somebody speaks to it. If you tell an LLM today that you have invented a new species of penguin, the underlying neural network does not normally retrain itself during the conversation so that every future user tomorrow knows about your penguin.
That continuity must be supplied elsewhere.
So how many days does an LLM remember?
We can now see why the question has no sensible answer expressed simply as “thirty days”, “six months” or “forever”.
A frontier LLM may simultaneously have several different temporal horizons. Its immediate attention extends over whatever material is made available in its current context. Its conversational continuity may encompass an entire dialogue, subject to context management and summarisation. Its persistent personal memory may extend across conversations for as long as the surrounding service retains and successfully retrieves relevant information. Its parametric knowledge may persist until the model is replaced, modified or retrained. External retrieval systems can potentially give it access to archives extending across decades or centuries.
The practical memory of an AI assistant is therefore becoming less a question of capacity than one of selection. That may turn out to be the more interesting problem.
The danger of remembering everything
For the first generation of conversational AI, forgetting was an obvious weakness. Users repeatedly had to explain themselves. They specified the same preferences, supplied the same documents and corrected the same assumptions. An assistant that remembered more seemed self-evidently superior.
But perfect memory would introduce problems of its own. Human relationships depend partly upon forgetting. People change their minds. Their tastes alter. They abandon projects. They say foolish things at two o’clock in the morning that ought not to become permanent elements of their biography. A machine that faithfully resurrected every casual statement a person had ever made might rapidly become intolerable.
There are also obvious questions of privacy and control. Information that improves an assistant’s usefulness may simultaneously become information that a user does not want retained indefinitely. The engineering challenge is therefore inseparable from a normative one: what should an artificial intelligence remember about us?
The best AI memory may consequently turn out not to be the largest memory but the most discriminating one. It should preserve durable information that genuinely improves future interactions while allowing transient, erroneous, embarrassing or unwanted information to fade away.
In this respect artificial intelligence confronts a problem that evolution has already confronted in human beings. Forgetting is not necessarily a defect in memory. It is part of the architecture of useful memory.
The strange future of artificial acquaintance
The implications become more intriguing as AI systems acquire longer histories with individual users. Imagine an assistant that has worked with somebody for twenty years. It might know how that person’s prose changed, which projects succeeded, which friendships disappeared, which arguments repeatedly preoccupied him and which ambitions he quietly abandoned.
Such a system would possess something no previous machine has possessed on any comparable scale: a longitudinal intellectual record of a human life.
Whether that amounts to knowing a person is a philosophical question rather than an engineering one. The machine might retrieve thousands of facts while experiencing none of them. It could recall the user’s wedding without having attended it, remember his grief without ever having grieved and recognise the jokes he told twenty years earlier without possessing anything resembling nostalgia.
Yet from the human side of the conversation the distinction may become psychologically difficult to maintain. Something that remembers us begins to acquire the appearance of acquaintance. Something that remembers us accurately for decades may acquire the appearance of an old friend, irrespective of what is happening on the other side of the screen.
That prospect is both useful and unsettling.
The frontier large language model therefore has no single length of memory. Depending upon what we mean, its memory might last a few moments, an entire conversation, many years or approximately as long as the technological infrastructure supporting it continues to exist. The more important transformation is that AI systems are moving from machines that merely answer questions towards systems capable of constructing continuity across time.
For centuries we have assumed that memory belongs to the person doing the remembering. Artificial intelligence complicates that assumption. Increasingly, part of our remembered lives may exist outside ourselves, reconstructed by machines from conversations, documents and digital traces whenever circumstances require them.
The great irony may be that machines will eventually remember far more about our lives than we do. The great question will not then be how long their memories are. It will be what we want them to forget.




