Introduction
Latent learning is learning that happens without reinforcement and stays invisible until a reason to use it appears. A rat wanders a maze for ten days, finds nothing, and looks like it has learned nothing. Then food appears at the end. The next day it runs the maze almost perfectly. The knowledge was there the whole time. Only the motive was missing.
You have met the standard version of this story. It is on every psychology revision site and in most introductory textbooks. Edward Tolman put rats in a maze, three groups, one fed, one never fed, one fed only from day eleven. The third group appeared to learn nothing until the food arrived, then their errors collapsed overnight. Conclusion: reinforcement is not necessary for learning, behaviourism was wrong, and the cognitive map was born.
Almost every clause in that paragraph is wrong.
Tolman did not run the first latent learning experiment and did not name the phenomenon. There were four experimental conditions in the 1930 paper, not three, and the fourth one has almost vanished from the internet even though the paper's own title announces it. The rats that supposedly learned nothing were measurably improving from the first week. The group that got food on day eleven did not catch up with the always-fed rats. It passed them. And the result did not defeat behaviourism. The leading behaviourists explained it, argued about it for a generation, failed at least once to replicate it, and walked away from a question nobody had answered.
None of this is hidden. Most of it sits in a paper Tolman published in 1948 that is free to read online, and the rest in an open access review from 2006 that reproduces the original figures. The 1929 and 1930 papers are print only, so what comes from them here comes through those two, which both quote them directly. It has simply never reached the version of the story most people meet.
This article puts the record straight, then follows the idea into the last three years, where latent learning turns up in sleeping mice, in ants, in human brain scans and, in 2025, in a Google DeepMind paper naming it as something machine learning still cannot do.
What Latent Learning Actually Is
Start with the distinction the whole thing rests on, because it is the part worth keeping.
Learning and performance are not the same thing. Learning is a change inside the animal. Performance is what the animal does. You can have the first without the second, and the gap between them is where latent learning lives.
Nicholas Soderstrom and Robert Bjork put the modern version of it in a 2015 review that opens on exactly these rat studies: although reinforcement is necessary to reveal learning, it is not required to induce learning [1].
That sentence is doing a lot of work. It concedes the behaviourist point that reward controls behaviour. It denies the stronger claim that reward is what builds the knowledge in the first place.
There is a second distinction that almost every popular account collapses, and it produces a definition you will see everywhere: that latent learning is learning without reinforcement or motivation. The rats were motivated. They were food deprived for the entire study, and Tolman and Honzik published a companion paper in the same 1930 volume, "Degrees of hunger, reward and non-reward, and maze learning in rats", in which hunger was the variable being manipulated. What was missing was reinforcement at the goal box, not drive. Drive and reinforcement are different things, and if you treat them as one a precise finding turns into a vague one.
The Man Who Ran It First
His name was Blodgett, he worked at Berkeley, and he published in 1929.
H. C. Blodgett ran three groups of rats through a six unit alley maze, one trial per day. One group found food at the end of every run. A second group was not fed in the maze for the first six days and first found food on the seventh. A third group first found food on the third day. When the food appeared, the error curves of both delayed groups dropped sharply, and Blodgett named the effect.
We know all of this because Tolman himself wrote it down, in the single most cited paper on the subject.
In his 1947 presidential address to the American Psychological Association, published the following year as "Cognitive maps in rats and men", he is completely explicit [2].
"The first of the latent learning experiments was performed at Berkeley by Blodgett. It was published in 1929. Blodgett not only performed the experiments, he also originated the concept." A few lines later he describes the delayed group's improvement and adds that this learning, which did not show itself until after the food had been introduced, was what Blodgett called latent learning.
Then, describing his own famous study, Tolman writes one of the more disarming sentences in the psychological literature. "Honzik and myself repeated the experiments (or rather he did and I got some of the credit) with the 14-unit T-mazes shown in Fig.1, and with larger groups of animals, and got similar results."
He gave the credit away twice in one page. To Blodgett for the discovery, and to Charles Honzik for the work. Neither name appears in most of what is written about latent learning today.
Soderstrom and Bjork are among the few to say it in print, describing the 1930 study as essentially a replication of Blodgett's [1].
If you take one correction from this article, take that one. Search for who discovered latent learning and the answers you get are quiz pages and homework sites naming Tolman. The man who ran it, named it and published it first said otherwise, in writing, in 1948.
Inside the Maze
The apparatus matters, because the version you will have read says "a complex maze" and leaves it there.
Tolman and Honzik used a 14 unit T-maze with the blind alleys numbered one through fourteen [3]. Each rat ran it once a day. Food was placed in the goal box for the delayed group on day eleven, and the chart everyone reproduces covers the first 17 days [3].
Two details there look like a contradiction and are not. The paper's own text puts the reward period from the twelfth to the twenty-second day inclusive [3]. So the first rewarded run was day eleven, the first full reward day was day twelve, and the study ran well past the end of the chart you will have seen.
The group labels below are Tolman's own, from his 1948 paper [2].
They are worth using because the numbering conventions on the open web contradict each other. One site calls the never-fed rats Group 2, another calls them Group 3, and if you check two sources against each other you get two different experiments.
One more thing about the figure everybody reproduces. Its vertical axis is not a raw count of wrong turns.
It is labelled the average number, with a constant multiplier, of blind alleys entered [3].
So the shape of the curves is real, and the numbers printed up the side are not something you can quote as errors. This article gives none, for that reason.
Nor will you find a sample size here. The number of rats per group in the 1930 study is not recoverable from any source available now, and the closest anyone gets is Tolman's own phrase about using larger groups of animals than Blodgett had. Inventing a figure would be worse than admitting the gap.
The Rats Were Never Aimless
Here is the first place the story you were told goes wrong in a way that changes the meaning.
Wikipedia says the never rewarded rats wandered the maze without preferentially heading for the end. Read the actual curves and something different appears.
Robert Jensen's 2006 review reproduces all three of them and reports that the average error rate for all three groups of rats declined over the first five days [3]. All three. Including the two that had never seen food.
Think about what that rules out. If food were the only driver, the unfed rats should have shown no decline, or the fed rats a far bigger one. Neither happened. Something in the maze itself was shaping behaviour before any reward appeared.
That something is not mysterious. A blind alley, once entered, stops you. Rats enter it less often afterwards. Getting out of the maze at all is an outcome.
And in 1954 K. C. Montgomery showed that rats will learn a discrimination for no reward except the chance to explore a novel maze, which makes exploration a drive in its own right rather than a null condition [4].
So the unfed condition was never a blank. It had a quieter set of consequences than the fed condition, which is a very different claim from none at all.
They Did Not Catch Up. They Passed.
The famous moment is the day after the food appears, and almost every retelling underdescribes it.
The usual phrasing is that the delayed group did as well as, or became comparable to, the always rewarded group. What happened is stronger.
Jensen's reading of the original data is that on the very next trial these newly fed rats surpassed the performance of the rats fed from the first day [3]. They did not draw level. They went past.
That overshoot is interesting for a reason nobody writing for a general audience mentions. It has a name in the animal learning literature, and it is motivational rather than cognitive.
The year before Blodgett, in the same University of California series, M. H. Elliott switched rats from one food to a less preferred one partway through a maze study, and their performance dropped below the level of controls that had never had the better food.
In 1942 Leo Crespi measured both directions of the effect systematically and named them elation and depression, in a paper that is still the reference point for incentive contrast [5].
The effect held up over decades. David Zeaman's 1949 runway study found response latency changing sharply with the amount of reinforcement, in both directions [6]. Charles Flaherty's 1982 review pulled together roughly thirty years of shifts in reward magnitude and the performance swings that followed them [7].
Here the inference is mine, not the literature's. No published source calls the Tolman and Honzik overshoot a contrast effect. What can be said is that a sudden upward shift in incentive produces exactly this shape, and that Elliott and Crespi documented it in the same laboratory in the same decade. Once you know that, you read the famous graph differently.
The Group Whose Food Was Taken Away
The 1930 paper is called "Introduction and removal of reward, and maze performance in rats". Almost every page that tells this story quotes that title, then reports half the experiment.
There was a fourth condition. Those rats were fed in the goal box from the first day, exactly like the standard rewarded group, and then on day eleven the food stopped.
Jensen labels them HR-NR and reproduces their curve, on which the errors rise back to match the level of the unfed rats [3].
This arm is the mirror image of the famous one and fits the tidy moral much less well. If the animals had built a map of the maze, taking the food away should not have unbuilt it. Their knowledge did not go anywhere. Their reason to use it did. That is the learning versus performance point running in the opposite direction, and the strongest single demonstration in the paper that reward does motivational work rather than instructional work.
Ten seconds of searching finds the introduction half. Almost nothing will tell you the removal half exists.
The Reply From the Other Side
Every explainer says the result broke behaviourism. It is the single most repeated claim about latent learning, and it is not what happened.
Clark Hull built the most elaborate stimulus response theory of the period, in which the strength of a response was a multiplicative function of drive, stimulus intensity, incentive and habit strength. Hull worked directly on these curves.
Using what he called the fractional antedating goal reaction, combined with the shift in incentive magnitude when food appeared on day eleven, he derived the Tolman and Honzik pattern from his own system, in a treatment his 1932 goal gradient paper had already set up [8]. Not explained it away afterwards. Derived it.
Guthrie is the more awkward case for the standard story. His 1930 statement of contiguity theory held that learning happens when a response occurs in the presence of a stimulus, full stop, with reinforcement playing no part in forming the association [9].
On that account latent learning is not a surprise. It is the prediction. If you are told the finding refuted behaviourism, notice that a leading behaviourist theory expected it before anyone ran it.
Guthrie is also the source of the best insult in the debate. Tolman had argued that what a rat does at a choice point, hesitating and looking back and forth, reveals expectation rather than habit, and in 1938 he named it vicarious trial and error [10].
In the 1952 revised edition of The Psychology of Learning, on page 143, Guthrie wrote his reply. "So far as the theory is concerned the rat is left buried in thought; if it gets to the food box at the end that is its concern, not the concern of the theory."
The joke aged badly. A. David Redish reviewed the recording work in 2016 and reported that the pause at a choice point coincides with hippocampal activity sweeping down each candidate path in turn, one after another [11].
The rat is deliberating. It just does it faster than you can watch.
A Failure to Find the Blodgett Effect
The 1940s and early 1950s produced an experimental war over this that has left almost no trace in the public account, and the results were nothing like unanimous.
K. W. Spence and R. Lippitt ran a Y-maze study in 1946 at Iowa, Hull's own institution, with rats satiated for both food and water during training, and found that the animals had nonetheless learned which arm held which [12]. Tolman thought it was the best latent learning experiment anyone had done, which is a strange compliment to receive from the opposing camp. Howard Kendler tested latent learning in a T-maze the following year and read his results as unfavourable to Tolman's expectancy account [13]. John Seward's 1949 analysis found latent learning after just 30 minutes of free exploration, a result Soderstrom and Bjork still cite seventy years later [14]. Those three point in three different directions, all within four years.
Then, in 1951, Paul Meehl and Kenneth MacCorquodale published a paper whose title says everything: "A failure to find the Blodgett effect, and some secondary observations on drive conditioning" [15]. The two of them had published a further T-maze study of latent learning three years earlier [16], so this was not a hostile outsider taking a swing.
It was researchers inside the problem reporting that the effect did not appear.
Almost nothing written for a general reader mentions that a titled replication failure exists.
James Deese added another awkward result the same year, showing that a discrimination could be extinguished without the animal ever performing the choice response [17].
Thirty Years and No Verdict
In 1951 Donald Thistlethwaite published a review in Psychological Bulletin covering the whole of the latent learning literature to that point [18].
It is not a triumphant document. It catalogues two decades of experiments that kept giving different answers depending on the maze, the drive state and the control condition.
The debate did not end with a winner. Jensen is blunt about how it ended: what had begun as a lively empirically based debate over fundamental issues in learning ultimately ended in a stalemate, and by the mid 1960s many psychologists considered the matter dead [3].
That is the honest summary, and more interesting than the version where one side wins.
Was Anything Really Unrewarded?
This is the crux of the whole controversy, and almost nobody writing about latent learning says so.
The standard line is that the rats got nothing for their trouble. Look closely and the unrewarded condition turns out to be full of consequences. Blind alleys punish. Being lifted out of the goal box and returned to a home cage is an event that reliably follows a run. The animals were hungry throughout and were fed later. Exploration itself functions as a drive [4].
Jensen's argument in 2006 is that you can rebuild the entire 1930 result out of punishment, stimulus control and ordinary environmental contingency, without invoking a map at all [3].
You do not have to accept it to notice that it exists and that nobody writing for a general audience engages with it.
Here is the reasonable position. Reinforcement in the narrow sense, food at the goal, was absent. Consequences were not. Calling the condition reward free overstates it, and calling it proof that learning needs no consequences overstates it much further.
The Cognitive Map Came From a Different Maze
Ask anyone where the idea of the cognitive map comes from and they will point you at these rats. It is the wrong maze and the wrong decade.
The concept arrived in 1946, sixteen years later, in a different apparatus.
E. C. Tolman, B. F. Ritchie and D. Kalish trained rats on an indirect path to a goal, then blocked it and offered a fan of eighteen alternative alleys radiating out in different directions [19].
Rats that chose the alley pointing at where the food had been, the argument went, must be working from a map rather than a chain of learned turns. That is the sunburst maze, and that is where your textbook's cognitive map comes from.
Jensen counted how often textbooks get this right. Across 10 introductory psychology textbooks published between 2000 and 2003, 7 of the 10 referred only to the Tolman and Honzik study when explaining the cognitive map [3]. His wider survey covered 48 introductory textbooks published from 1948 to 2004, 21 of them since 1999 [3].
David Olton had already flagged the deeper version of this in a 1979 review in American Psychologist, pointing out that the T-maze studies were experiments about behavioural stereotypy while the map concept was derived from experiments about flexibility [20]. Different apparatus, different question, different evidence.
For the modern account of what a cognitive map is and where it lives, the neighbouring article on how the hippocampus decides what to remember covers the machinery.
Eighty Years On, the Sunburst Barely Replicates
This is the part of the story that only became tellable in January 2026.
Éléonore Duvelle and Roddy Grieves published a meta-analysis of every attempt to repeat Tolman's sunburst experiment.
Counting the original, they found 13 studies reporting 47 separate experiments, across rats, squirrel monkeys and humans, and shortcutting exceeded chance in 17 percent of them [21].
What the animals did instead is the interesting part. In 32 percent of experiments they preferred a path immediately adjacent to their training route, and in 26 percent they showed no preference at all.
In 13 percent they favoured paths with no remarkable features, and 6 percent apiece went to a visual cue or to one of the outermost alleys [21].
The historical detail is worse than the headline number. None of the six studies published in the 20 years following Tolman's 1946 paper replicated the shortcutting preference, and the authors call Tolman's own result a clear outlier against the literature that followed it [21].
Humans do not do much better. In 2018 Stuart Wilson and Paul Wilson ran the experiment in virtual reality with 60 participants, 30 men and 30 women with a mean age of 20.33 years, using a head mounted display [22].
Everyone completed seven training trials along the same corridor route, then found it blocked and 18 alternatives open. Ask yourself which way you would turn.
With 20 people in each condition, only 20 percent of those in the best performing condition chose the path leading directly to the goal, and in the other two conditions the median choice sat about 80 degrees off target [22].
One caution, and it is the authors' own. Duvelle and Grieves say it plainly, so this article repeats it. Their reinterpretation does not challenge the broader cognitive map framework, which stands on behavioural and neurophysiological evidence gathered elsewhere [21].
What fails to replicate is one canonical demonstration, not the theory it was recruited to support. That distinction is easy to lose, and worth holding onto as you read the rest.
What the Brain Is Doing
The theory the sunburst maze was supposed to support got much better evidence later, from electrodes rather than alleys.
In 1971 John O'Keefe and Jonathan Dostrovsky recorded from the hippocampus of freely moving rats and found cells that fired when the animal was in one particular place and nowhere else [23]. Place cells turned the cognitive map from an inference into something measurable. By 2008 Edvard Moser and colleagues could describe a full spatial system of place cells, grid cells and head direction cells [24].
Latent learning has now been recorded at that level. In 2024 Wei Guo, Jie Zhang, Jonathan Newman and Matthew Wilson tracked hippocampal CA1 in 16 male mice exploring a maze without reward for 30 to 40 minutes a day across five to seven days [25].
The population code reorganised into a low dimensional structure resembling the physical environment. What moved it was sleep between sessions, not time in the maze.
More than 30 percent of the recorded cells in those 16 mice never formed a place field at all, and they behaved differently from the ones that did [25].
Two other results push in the same direction. In 2015 Freyja Ólafsdóttir and colleagues let 4 rats see food in an arm they were never allowed to enter, and found the hippocampus running sequences through that unvisited space during rest, at 7.37 percent of events for the baited arm against 4.41 percent for the unbaited one [26].
Note what that asymmetry means. The preplay happened only where a reward had been seen, which complicates the story rather than confirming it.
In 2025 Lennart Wittkuhn and colleagues found replay in the human visual cortex during brief pauses in a task, and it predicted implicit learning of the sequence structure in participants who could not report that any structure existed [27].
Even the reward system turns out to be less binary than the textbook version.
In 2017 Melissa Sharpe and colleagues showed that a pulse of dopamine is enough on its own to bind two neutral cues together, and that blocking it prevents the binding, with no reward involved [28].
If dopamine builds structure between things that were never rewarded, the clean line between rewarded and unrewarded learning was always going to be hard to hold. The neighbouring article on dopamine and learning goes into what that molecule is actually signalling. The sleep dependence in the mouse work connects to the same story told from another angle in sleep and memory.
Latent Learning in People
The honest answer has two sides, and giving only one is how this topic usually gets written.
On the positive side, Layla Unger and Vladimir Sloutsky ran five experiments with 438 adults in 2022, exposing people to categories incidentally while they worked on an unrelated cover task [29].
The finding is precise and often misquoted. Incidental exposure produced a ready to learn effect even when participants showed no evidence of category learning during the exposure itself. Exposure created readiness to learn. It did not create the learning.
Milena Rmus and colleagues tested 77 participants in 2022 on an abstract graph they had explored with no goal, and found them judging distances between nodes they had never seen paired at 67 percent accuracy against a chance level of 50 [30].
That is latent learning of relational structure rather than of physical space.
Now the other side. Miguel Vadillo and colleagues ran four preregistered experiments in 2020 to replicate a well known demonstration of latent learning from ignored visual context, with final samples of 49, 55, 47 and 106 participants against an original study of 20 participants [31].
Every relevant interaction came out nonsignificant, and the Bayes factors ran from 3.55 to 11.28 in favour of the null. As they put it, far from uncovering latent learning, reversing the colour of distractors actually disrupts the expression of contextual cueing.
That is a live disagreement about a specific paradigm, published five years ago, and almost nothing written for a general reader mentions it.
The oldest human result also contradicts a popular assumption.
Harold Stevenson tested latent learning in children in 1954 and found it in children as young as three, increasing with age rather than decreasing [32].
Young children absorb less from unrewarded exposure, not more.
There is a much larger literature that is latent learning under other names.
In 1996 Jenny Saffran, Richard Aslin and Elissa Newport gave 24 infants aged eight months two minutes of continuous nonsense speech and found them afterwards listening reliably longer to strings that crossed a word boundary [33].
Nothing in those two minutes was rewarded, and this is roughly how you learned your own first language.
Arthur Reber showed the same thing with artificial grammars in 1967, with people sorting strings correctly and unable to say what rule they used [34]. Contextual cueing, described in 1998, shows repeated visual layouts speeding search while the searcher never notices the repetition at all [35].
It Is Not Only Mammals
Here is the finding that quietly demolishes the lesson most explainers draw.
In 2023 Laure-Anne Poissonnier, Yannick Hartmann and Tomer Czaczkes published a study in which ants learned an object's affordance, specifically whether a feeder could hold one ant or several, without that attribute ever being rewarded [36]. The ants later used that memory to avoid crowding they had predicted but never experienced.
Every explainer treats latent learning as evidence of higher cognition, something map like and mammalian sitting above simple association. An insect doing it suggests the opposite of the moral you were given, that acquiring structure from unrewarded exposure is a basic property of learning systems rather than a badge of sophistication.
Four Kinds of Learning That Get Confused
If you have been trying to tell these four apart you are in large company, and the pages that explain them mostly offer a definition and a shrug.
The pair worth pulling apart is the first two. Albert Bandura's 1965 study is the cleanest demonstration, using 66 children aged between 42 and 71 months at the Stanford nursery school, split into three groups of 22 [37]. Children who watched a model punished imitated less than children who watched a model rewarded. Then everyone was offered an incentive to reproduce what they had seen, and the differences between groups vanished, with the punished group showing the largest jump at t = 5.00 [37].
The children had all learned the same amount. Only their willingness to show it differed.
That is the identical logic to the maze, running through a social channel instead of a spatial one, and the fuller account is in the article on what the Bobo doll experiment really showed. Insight learning is a different animal again, and it gets its own treatment in insight and aha moments.
Why the Effect Kept Not Replicating
A 2026 computational paper offers something the 1950s never had, which is a reason why the same experiment kept giving different answers.
Matheus Menezes, Xiangshuai Zeng and Sen Cheng modelled latent learning using successor representations, a formalism in which an agent learns which states tend to follow which others rather than which actions pay [38].
Their result is that the size of the latent learning benefit depends on the statistics of how the animal explored during the unrewarded phase. Exposure targeted toward the region that will later matter helps most. Undirected exposure helps less. Mistargeted exposure helps least.
The connection to the historical record is mine, not theirs. The 1940s and 1950s protocols varied the maze, the drive state and the exploration schedule without controlling any of them tightly, which is exactly what this model says the effect is sensitive to. That predicts an inconsistent literature. It is not a claim the paper makes.
The idea holds up outside the model. In 2017 human choices were predicted better by an agent learning which states follow which than by one learning which actions pay [39], and a companion paper recast hippocampal place fields as a predictive map, firing for where an animal is about to be [40].
In 2020 a model unifying spatial and relational memory was named the Tolman-Eichenbaum Machine, which tells you whose problem the field still thinks it is working on [41].
What This Means If You Are Studying
There is a piece of advice circulating on education sites that cites latent learning to argue that students should be exposed to material without the pressure of constant testing. It is a reasonable sounding inference and it is backwards.
The evidence on retrieval practice is among the most replicated in educational psychology.
In 2006 Henry Roediger and Jeffrey Karpicke gave 120 undergraduates a passage to either restudy or test themselves on, and after a week the tested group recalled 56 percent of the material against 42 percent for the restudy group, an effect size of 0.83 [42]. Their second experiment, with 180 undergraduates, was sharper still: repeated testing produced 61 percent recall after a week against 40 percent for repeated study, even though the repeated study group had read the passage 14.2 times and the tested group only 3.4 [42].
John Dunlosky and colleagues reviewed ten study techniques in 2013 and rated rereading and highlighting as low utility, with practice testing among the two highest [43]. Roediger and Andrew Butler put the same conclusion in one line in 2011: retrieval practice is more effective for long term retention than restudying [44]. The effect even shows up in animals [45].
So what does latent learning actually give you? One thing, and it is worth more than a study hack.
It tells you that your performance right now is a bad estimate of what you have learned. That is why a revision session can feel productive and change nothing, and why material you cannot produce today can come back the moment a real demand appears.
Soderstrom and Bjork describe storage strength as acting as a latent variable, something you cannot read off performance directly [1].
The practical version of this is covered in the article on desirable difficulties, and the habit side of the same coin is in how habits become automatic.
One popular example is worth retiring while we are here. London taxi drivers are usually offered as the human case of a cognitive map built by exposure.
The Knowledge is three to four years of deliberate, effortful study. In 2011 Katherine Woollett and Eleanor Maguire followed 79 male trainees and 31 controls through it, and found grey matter growth in the posterior hippocampus only in the 39 who qualified [46].
The ones who trained and failed showed no such change. What separated them was effort: the qualifiers put in 34.56 hours a week against 16.70. That builds on the 2000 finding, in 16 right-handed men who drove taxis against 50 controls, where posterior hippocampal volume tracked time driving at r = 0.6 [47].
Taxi drivers are the contrast case for latent learning, not an example of it.
Machines Have the Same Gap
The oldest argument in this article turned up in a machine learning paper last year. Nobody would have predicted that in 1965.
In September 2025 a group at Google DeepMind and Stanford, led by Andrew Lampinen with James McClelland among the authors, named a specific weakness of machine learning systems. They do not exhibit latent learning [48]. Anything a model met during training but did not need at the time is not available to it later.
Their proposal is that episodic memory can complement parametric learning by making that stored experience reusable.
Four months later, in January 2026, a nine author group published an experiment giving large language models an unrewarded exploration phase before reinforcement learning began [49].
Models that explored without reward first ended up more capable than models trained with rewards throughout. The paper names Blodgett and Tolman directly.
That deserves to be read at the right size. It is one preprint, seven months old, and a language model is not a rat. But the shape of the finding is the shape of the 1929 result, and engineers arriving at the same distinction by a different road is a reasonable sign that it is real.
What Is Settled and What Is Not
Several parts of this story are beyond argument now. Blodgett ran the first experiment, in 1929, and named the effect. The HNR-R rats improved sharply when food appeared on day eleven. The reward removal group exists and its errors rose. The cognitive map dates from 1946 and a different maze. Learning and performance come apart, and an incentive can move performance without changing what was learned.
Much of it, though, remains genuinely open. Whether latent learning demonstrates that reinforcement is unnecessary for learning was disputed by Hull and Guthrie on theoretical grounds and by Meehl and MacCorquodale on empirical ones, and Thistlethwaite's review left it open in 1951. Whether the unrewarded condition was ever really consequence free is still arguable. Whether latent learning of ignored visual context happens in humans now looks doubtful, given four preregistered experiments in one paper that between them found nothing, against a small original.
The last distinction here is worth marking clearly, before you go. The sunburst result is contested. The cognitive map is not. Confusing those two is the error the meta-analysis authors went out of their way to prevent.
Conclusion
The reason to get this story right is not pedantry about who published in 1929.
The version everyone repeats is a story about a hero and a broken paradigm. The real one is better. A graduate student ran a careful experiment, a more famous colleague replicated it and spent years giving him the credit, the leading theorists explained the result from their own systems, a generation of experiments produced answers that would not line up, and the field put the question down without settling it. Eighty years later the canonical demonstration turned out to work about one time in six, and the idea underneath it turned up in ant colonies and in language models.
What holds up through all of that is the distinction Blodgett's rats made visible. What you can do at this moment and what you have actually learned are two different quantities, and only one of them is available for inspection. That was worth a thirty year argument. It is still worth knowing.
Frequently Asked Questions
What is latent learning in simple terms?
Latent learning is learning that takes place without reinforcement and stays hidden until you have a reason to use it. The standard demonstration is a rat that explores a maze for days with no food at the end, shows little sign of improvement, and then runs the maze almost perfectly the day after food appears. The knowledge was acquired during the unrewarded period. What the reward supplied was a motive to display it, which is why your own performance is a poor guide to what you actually know. The concept was named by H. C. Blodgett in 1929 and it is the foundation of the modern distinction between learning and performance.
Who discovered latent learning?
H. C. Blodgett, at the University of California, Berkeley, in 1929. He ran three groups of rats through a six unit alley maze, one trial per day, with one group fed at the end of every run, one first fed on the seventh day and one first fed on the third. He also coined the term. If you have read that Tolman discovered it, you have read the version everyone repeats. Edward Tolman and Charles Honzik replicated the design at larger scale in 1930 using a 14 unit T-maze, and their version is the one that ended up in the textbooks. Tolman was explicit about the priority in his 1948 paper, writing that Blodgett not only performed the experiments but also originated the concept.
Does latent learning require reinforcement?
No reinforcement is delivered at the goal during the latent phase, which is the whole point of the design. Whether the animal receives no consequences at all is a different and much more contested question, and you should treat the two as separate. Blind alleys in a maze stop forward movement, removal from the goal box is itself an event, the rats in the 1930 study were food deprived throughout, and exploration was later shown to function as a drive in its own right. Clark Hull derived the classic results from stimulus response theory, and Guthrie's contiguity theory predicted latent learning in advance because it required no reinforcement for association at all.
What did Tolman's experiments with rats in mazes demonstrate?
They demonstrated that a rat's performance in a maze can change overnight when an incentive appears, without any corresponding change in what it has had the opportunity to learn. The 1930 study used four conditions, not the three usually reported: rewarded throughout, never rewarded, rewarded only from day eleven, and rewarded until day eleven and then not. The delayed group's errors dropped so sharply that they went below the always rewarded group on the very next trial. The removal group's errors climbed back up. What moved in both cases was performance, not what the animals knew, and the same gap sits between what you can recall today and what you have actually learned. What the study did not demonstrate is that behaviourism could not account for the result, since the leading behaviourist theories of the period both did.
How does latent learning relate to cognitive maps?
Less directly than almost every account suggests. The cognitive map was not proposed on the basis of the 1930 maze study. It came from the sunburst maze of 1946, in which rats were trained on an indirect route and then offered eighteen alternative paths after the original was blocked, and from Tolman's 1948 paper. A 2026 meta-analysis of 13 studies and 47 experiments found that shortcutting exceeded chance in only 17 percent of them, and that none of the six studies published in the twenty years after 1946 replicated the effect. If you take one thing from that, take the authors' own care in noting that it challenges one canonical demonstration rather than the cognitive map framework itself, which rests on separate behavioural and neurophysiological evidence including the discovery of place cells in 1971.




