Does the Marshmallow Test Predict Anything?

TL;DR

  • Between 1968 and 1974, 653 children passed through a waiting experiment at one Stanford nursery school. The famous SAT link rests on 35 of them.
  • The correlation appeared in one of four waiting conditions and was negative in the other three. The authors gave a confidence interval of .10 to .66 for the verbal score and warned the numbers could exaggerate the true association.
  • A 2018 study asked the same question of 918 children. Waiting predicted achievement at 15 at 0.24 standard deviations, 0.08 once family background was accounted for, and a non-significant 0.05 once you add what the child could already do at four.
  • Nearly all of that came from clearing twenty seconds: children who held out the full seven minutes scored no better than children who managed half a minute.
  • At 26 the same children were surveyed on nine adult outcomes. Two showed a raw association (education .17, body mass index -.17) and both vanished under controls.
  • Two teams reading the identical data disagree about what this means. A small association with achievement is real. The story about destiny is not.

A child of four is led into a small room at the Bing Nursery School on the Stanford University campus. There is a table, a chair, a bell and a plate holding one marshmallow. An adult explains the arrangement: eat this one now, or wait until I come back and have two. Then the adult leaves, and out of sight a stopwatch runs. Among the children later followed up, the average wait was 512.8 seconds, over eight minutes.

You know what comes next, because everyone does. The children who held out grew up to score higher on college entrance exams, stay slimmer, earn more and cope better. One act of patience at four, and the arc of a life is drawn.

That story has circulated for thirty years. Underneath it sits a table of correlations built on thirty-five children.

What the room was originally for

The waiting task was never designed as a test of character. Walter Mischel, then a professor of psychology at Stanford, was studying attention: what a child does with their mind while enduring something unpleasant. In one central paper, with Ebbe Ebbesen and Antonette Raskoff Zeiss, the manipulation was where the treats sat and what the child was told to think about (Mischel, Ebbesen & Zeiss, 1972).

Children told to think about how delicious the marshmallow would taste caved fast. Children given something else to think about lasted far longer. Children with the treats hidden under a cake tin lasted longest, and suggesting a distraction to them added nothing, because the hiding had done the work already. Waiting was a skill of looking away, and the room could be arranged to make it easy or hard. Those versions of the room are what the follow-up study is built on.

Thirty-five children, four rooms

In 1990 Yuichi Shoda and Walter Mischel, both then at Columbia University, with Philip Peake of Smith College, published the follow-up that made the task famous (Shoda, Mischel & Peake, 1990). The arithmetic of its sample is worth walking through, because almost nobody repeats it.

  • 653 children took part at the Bing School between 1968 and 1974, of whom 550 sat the standard waiting version.
  • Two rounds of mailings to parents, in 1981 and 1984, produced any follow-up data at all for 185.
  • Parents of 94 of those wrote down their child’s SAT scores.
  • Split across the four versions of the room, the group carrying the headline result numbered 35.

Those 35 had waited with the treats in plain sight and no strategy suggested, the condition the authors predicted in advance would be revealing. There, delay time correlated with SAT verbal at .42 and quantitative at .57. In the three other conditions, with samples of 33, 14 and 12, the correlations were negative and not statistically significant.

A correlation of .57 in a sample of 35 is a wobbly thing, and the authors said so. Even the largest coefficients, they wrote, accounted for about a quarter of the variance, the small sample meant the observed numbers could exaggerate the true association, and the 95% confidence interval for the verbal correlation ran from .10 to .66. An interval that wide fits an effect you could barely detect and one that would reshape education policy.

What the popular version saysWhat the source saysGrade
Waiting at four predicts SAT scoresr = .42 verbal and .57 quantitative in one of four conditions, n = 35, 95% CI for verbal .10 to .66; negative and non-significant in the other three conditionsP
It holds in ordinary children, not just Stanford onesHalves in a sample of 918: 0.24 standard deviations at 15, then 0.08 with background controls and 0.05 with early-ability controlsP
Waiting longer means doing better, all the way upA threshold, not a gradient: past twenty seconds the extra minutes add nothing measurable at 15R
It predicts adult behaviour and healthNine outcomes at age 26, two with raw associations, none surviving controls for background and homeR
It is a measure of self-controlNot straightforwardly: it correlates more strongly with an applied-maths subtest (.37) than with attention, impulsivity or rated self-control (.22 to .30)P
Grades use the registry scale: C confirmed, P partial, H untested, R refuted.

How a table of correlations became a parable

A finding in Developmental Psychology reaches a few thousand readers. What carried the marshmallow test into staff rooms and airport bookshops was Daniel Goleman’s Emotional Intelligence, published in 1995 and sold by the million. The historian Michael Staub, tracing where the idea of self-control as a master skill came from, notes that Goleman leaned almost entirely on Mischel’s delay experiments and added a claim the data could not carry: that anyone could acquire the skill (Staub, 2016).

From there the number travelled without its footnotes. Search today and you will find pages announcing that the waiters scored 210 points higher on the SAT. No such figure appears in the 1990 paper, which reports correlations and nothing else. Of 113 widely read pages opened while checking this, 13 still quote the 210 points as a finding, and only 6 give the sample size behind the correlation.

918 children, and what happened to the number

In 2018 Tyler Watts, then at New York University, with Greg Duncan and Haonan Quan of the University of California, Irvine, went looking for the same association in a data set nobody had used for the purpose (Watts, Duncan & Quan, 2018). A national study of early child care had run a version of the task at 54 months and followed the children for years. The analysis sample was 918, ten times the 1990 follow-up and far closer to the American population in race, income and parental education. Much of the work concentrated on the 552 children whose mothers had not finished college, the group least like Stanford faculty families.

Three numbers tell the story, all for achievement at fifteen, in standard deviations:

  • 0.24 with nothing accounted for: statistically solid, and already about half the 1990 figures.
  • 0.08 after adding family background, income, race, mother’s education and an observed measure of the home.
  • 0.05, no longer distinguishable from zero, after also accounting for what the same child could already do at four.

To translate: 0.24 of a standard deviation is roughly the gap between a pupil in the middle of a class and one a little above it. Visible across hundreds of children, invisible in any one of them. And 0.05 is a pencil line.

The behaviour side collapsed harder. Where the 1990 work had parents rating long-waiters as calmer and better at concentrating, the 2018 team found next to no association with mother-reported problems at fifteen, before any controls at all.

The twenty-second cliff

The finding that unsettles the folk version most has nothing to do with controls. Watts and colleagues split the children by how long they lasted: under twenty seconds, twenty seconds to two minutes, two to seven minutes, and the full seven at which the task stopped.

With background accounted for, the three groups who cleared twenty seconds landed between 0.23 and 0.30 standard deviations above the children who did not, and were nowhere near statistically different from one another. At fifteen, a child who lasted twenty-five seconds and a child who sat out all seven minutes had the same expected achievement.

If patience were a dial that turns life outcomes, more of it would buy more. Instead the whole measurable difference sits at the bottom of the distribution, among children who could not manage twenty seconds. Less a scale of willpower than a screen catching something else.

What else is going on in the room

Two small experiments make the alternative concrete. At the University of Rochester, Celeste Kidd, Holly Palmeri and Richard Aslin gave 28 children an art session before the marshmallow arrived (Kidd, Palmeri & Aslin, 2013). Fourteen were promised better crayons and better stickers, and the researcher came back with them. For the other fourteen the researcher returned apologising, empty-handed, twice. Then came the identical waiting task, capped at fifteen minutes.

Children who had been let down waited an average of 3 minutes 2 seconds. Children whose researcher had delivered waited 12 minutes 2 seconds. One of the fourteen let-down children lasted the full fifteen minutes, against nine of fourteen in the other group. Two broken promises moved the score by nine minutes.

Laura Michaelson and Yuko Munakata, then at the University of Colorado Boulder, sharpened it (Michaelson & Munakata, 2016). Their children were not let down personally. They watched an adult behave well or badly towards a third person, and were then tested by that adult. Children assigned the untrustworthy adult were less likely to hold out and waited less time overall. Watching someone else be short-changed was enough.

Neither study says self-control is fictional. Both say a child who eats the marshmallow early may be making a sound bet about the room they are in.

The argument that followed, which is still live

The 2018 paper was read in the press as a debunking, and researchers who work on delay of gratification objected on two fronts. The economists Armin Falk, Fabian Kosse and Pia Pinger, then at the briq Institute in Bonn and the universities of Bonn, Munich and Cologne, published a preregistered reanalysis arguing the comparison was set up unfairly, partly because the newer task stopped at seven minutes while the Stanford version ran to fifteen or twenty (Falk, Kosse & Pinger, 2020). Watts and Duncan replied in the same issue, conceding a point that cuts both ways: given how wide the confidence intervals around the 1990 correlations were, almost any non-zero result would have fallen inside them (Watts & Duncan, 2020). A study most possible outcomes cannot contradict has not told you much.

The second objection used the identical data. Michaelson, by then at the American Institutes for Research, and Munakata, by then at the University of California, Davis, ran their own preregistered analysis and reached different conclusions (Michaelson & Munakata, 2020). Preschool waiting predicted three of the five adolescent outcomes they tested: achievement at 0.27, fewer problem behaviours at -0.22, better social skills at 0.18. The problem-behaviour link survived their multivariate and multilevel models, where the 2018 analysis had found nothing, largely because the two teams built the behaviour measure differently.

Then they asked what explained it. A composite of the child’s self-control across eleven years accounted for 17.1% of the variance in adolescent problem behaviour. A composite of social support instead accounted for 30.2%. Their reading: the task predicts because it registers something about the child’s surroundings. So two preregistered analyses of one data set disagree about whether the test predicts adolescent behaviour, and agree that self-control explains it poorly. Still unresolved.

The same children at twenty-six

In 2024 the cohort was old enough to check against adult life. Jessica Sperber and Tyler Watts of Teachers College, Columbia University, with Deborah Lowe Vandell and Greg Duncan of the University of California, Irvine, surveyed 702 of them at 26 across nine outcomes: education, earnings, debt, body mass index, depression, substance use, criminal involvement, risk-taking and impulsive behaviour (Sperber, Vandell, Duncan & Watts, 2024). The analysis was preregistered before anyone looked.

Two of the nine showed a raw association: years of education at .17 and body mass index at -.17. Adding basic demographics took education to .04, and adding the home environment took body mass index to -.07. Seven outcomes showed nothing even before controls. With background accounted for, children who waited the full seven minutes at four differed on no adult outcome from children who lasted under twenty seconds.

Their summary: the task derives much of its predictive power from links to other features of a child’s early life. A careful way of saying the marshmallow was reporting the weather rather than making it.

What Mischel said about all this

He did not defend the destiny version. He spent years arguing against it. In his 2014 book The Marshmallow Test: Mastering Self-Control, four years before the big replication, he wrote that «correlations that are meaningful, consistent, and significant statistically can allow broad generalizations for a population» while not licensing confident predictions about any individual, and noted that some children start out poor at waiting and improve while others go the other way (quoted in Ferlazzo, 2014). Asked about strategies rather than scores, he talked about what children do with their attention, his subject from the beginning.

The 1990 paper is equally careful: its authors called for replication in other cohorts and warned against generalising from thirty-five children at one university nursery school.

The shift: a headline correlation is a claim about a group, computed on one sample under one set of conditions. Before repeating it, find the sample size and the condition it came from.

First move: when a study is famous for a number, check whether the number is in the paper. The 210-point SAT gap is not in the 1990 study.

Second move: look for a threshold. An association that goes flat above some low cut-off describes the bottom of the distribution, not a virtue that scales.

The claim travels in company: emotional intelligence as the thing that accounts for most of career success came out of the same 1995 bestseller, growth mindset shrank the same way once large preregistered trials arrived, and willpower as a resource that depletes is the other pillar of the self-control story. All three sit in the registry of checked claims.

The boring bottom line

Something small and real survives. Across 918 children, waiting longer at four went with scoring about a fifth of a standard deviation higher at fifteen, and an independent preregistered reanalysis of the same cohort found links to behaviour and social skills too. Nobody serious claims the association is zero.

What does not survive is everything that made the story worth telling. The prediction halves outside Stanford, loses two thirds of the rest once family background and early ability are accounted for, sits almost entirely in the first twenty seconds, and cannot be found in nine adult outcomes at twenty-six. The two teams who disagree about the adolescent numbers agree that self-control explains them poorly: environment and social support do more of the work.

A four-year-old eating a marshmallow is telling you about the world they have learned to expect. Reading it as a forecast of their life was always more than thirty-five data points could bear.

Sources

  • Mischel, W., Ebbesen, E. B., & Raskoff Zeiss, A (1972). Cognitive and attentional mechanisms in delay of gratification. Journal of Personality and Social Psychology. 21(2), 204-218; the method paper behind the task: delay depends on where the rewards are and what the child is told to think about, and hiding the rewards removes the benefit of a suggested distraction doi:10.1037/h0032198
  • Shoda, Y., Mischel, W., & Peake, P. K (1990). Predicting adolescent cognitive and self-regulatory competencies from preschool delay of gratification: Identifying diagnostic conditions. Developmental Psychology. 26(6), 978-986; 653 children tested 1968-1974 at the Bing School, 185 with any follow-up measure, 94 parent-reported SAT scores, mean preschool delay 512.8 s (SD 368.7); Table 4 gives SAT verbal r = .42 and quantitative r = .57 for n = 35 in the exposed-rewards spontaneous-ideation condition, and negative nonsignificant values for n = 33, 14 and 12 in the other three; authors report 95% CI .10 to .66 (verbal) and .29 to .76 (quantitative) and warn the coefficients could exaggerate the true association doi:10.1037/0012-1649.26.6.978
  • Staub, M. E (2016). Controlling Ourselves: Emotional Intelligence, the Marshmallow Test, and the Inheritance of Race. American Studies. 55(1), 59-80; documents that Goleman's Emotional Intelligence (1995, ~5 million copies) relied almost entirely on Mischel's delay experiments to argue self-control is the key component of emotional intelligence, and added the claim that the skill can be learned doi:10.1353/ams.2016.0058
  • Watts, T. W., Duncan, G. J., & Quan, H (2018). Revisiting the Marshmallow Test: A Conceptual Replication Investigating Links Between Early Delay of Gratification and Later Outcomes. Psychological Science. 29(7), 1159-1177; analysis sample n = 918, of whom 552 had mothers without a college degree; for that subsample achievement at age 15 beta = 0.24 (SE 0.04, p < .001) bivariate, 0.08 (SE 0.03, p = .016) with background and HOME controls, 0.05 (SE 0.03, p = .140) with 54-month controls; Grade 1 equivalents 0.28, 0.10, 0.05; no association with behaviour composites at either age; with controls the three groups waiting over 20 s fell between 0.23 and 0.30 and did not differ (p = .752); marshmallow performance correlated r(916) = .37 with the WJ-R Applied Problems subtest against rs = .22-.30 for attention, impulsivity and self-control doi:10.1177/0956797618761661
  • Kidd, C., Palmeri, H., & Aslin, R. N (2013). Rational snacking: Young children's decision-making on the marshmallow task is moderated by beliefs about environmental reliability. Cognition. 126(1), 109-114; 28 children aged 3;6-5;10, 14 per condition; after two broken or kept promises, mean wait 181.57 s (unreliable) vs 722.43 s (reliable), W = 22.5, p < 0.0005; 1 of 14 vs 9 of 14 waited the full 15 min doi:10.1016/j.cognition.2012.08.004
  • Michaelson, L. E., & Munakata, Y (2016). Trust matters: Seeing how an adult treats another person influences preschoolers' willingness to delay gratification. Developmental Science. 19(6), 1011-1019; children who merely watched an adult behave untrustworthily towards a third person were less likely to wait the full delay, and waited less time overall, when tested by that adult doi:10.1111/desc.12388
  • Falk, A., Kosse, F., & Pinger, P (2020). Re-Revisiting the Marshmallow Test: A Direct Comparison of Studies by Shoda, Mischel, and Peake (1990) and Watts, Duncan, and Quan (2018). Psychological Science. 31(1), 100-104; preregistered commentary (osf.io/5jpt4) arguing the 2018 comparison over-controlled and was affected by the 7-minute censoring of the newer task; the article itself is paywalled, and its Monte Carlo simulations are described in the authors' reply below doi:10.1177/0956797619861720
  • Watts, T. W., & Duncan, G. J (2020). Controlling, Confounding, and Construct Clarity: Responding to Criticisms of "Revisiting the Marshmallow Test" by Doebel, Michaelson, and Munakata (2020) and Falk, Kosse, and Pinger (2020). Psychological Science. 31(1), 105-108; read in the accepted preprint (10.31234/osf.io/hj26z); concedes that the confidence intervals around the Shoda et al. correlations were wide enough that virtually any non-zero estimate would fall inside them, and that the censoring did not change the dummy-variable conclusion because the 7-minute group's return equalled the 20-second group's doi:10.1177/0956797619893606
  • Michaelson, L. E., & Munakata, Y (2020). Same Data Set, Different Conclusions: Preschool Delay of Gratification Predicts Later Behavioral Outcomes in a Preregistered Study. Psychological Science. 31(2), 193-201; independent preregistered analysis of the same SECCYD cohort; bivariate associations for 3 of 5 adolescent outcomes (achievement beta = 0.27, problem behaviour -0.22, social skills 0.18), problem behaviour holding at -0.18 in multivariate models; in multilevel models a self-control composite explained 17.1% of Level 2 variance in adolescent problem behaviour against 30.2% for a social-support composite doi:10.1177/0956797619896270
  • Sperber, J. F., Vandell, D. L., Duncan, G. J., & Watts, T. W (2024). Delay of gratification and adult outcomes: The Marshmallow Test does not reliably predict adult functioning. Child Development. 95(6), 2015-2029; preregistered (osf.io/67XFN) age-26 follow-up of 702 SECCYD participants across nine outcomes; bivariate educational attainment beta = .17 (p < .001) and BMI beta = -.17 (p < .001), falling to .04 (p = .23) and -.07 (p = .11) with controls; no significant associations for the other seven outcomes; in fully adjusted models no adult outcome distinguished children who waited the full 7 min from those who waited under 20 s doi:10.1111/cdev.14129
  • Ferlazzo, L (2014). 'The Marshmallow Test': An Interview With Walter Mischel. Education Week. Published 21 September 2014; quotes Mischel's own book The Marshmallow Test: Mastering Self-Control (2014) saying that statistically significant correlations allow broad generalizations for a population but not confident predictions for an individual, and that some children start low in delay ability and improve while others decline edweek.org