Getting Things Done: Book Notes and the Real Evidence

TL;DR

  • «Getting Things Done» claims that unfinished commitments held in your head create a persistent drain on attention, and that writing every one of them into a trusted external system removes the drain. The first half of that claim has real experimental support. The second half is where it splits.
  • The paper most often cited to prove Allen right also contains the finding that damages his subtitle. In Masicampo and Baumeister’s Studies 5A and 5B (N = 174 and N = 80), making a plan removed intrusive thoughts about the unfinished task but left self-reported anxiety unchanged.
  • The Zeigarnik effect, the 1927 result that «open loops» leans on, is the weakest brick in the wall. A 2025 meta-analysis of 38 publications found a weighted recall ratio of 0.99 — interrupted tasks are remembered about as well as finished ones.
  • Externalizing intentions genuinely works. Across four online experiments with 1,196 participants, setting a reminder improved follow-through, and people set reminders adaptively based on memory load.
  • The GTD package itself has never been tested in a controlled study. The two academics who wrote the only serious scientific defence of it said so in 2008. Searches of Crossref and Europe PMC on 10 August 2026 turned up nothing since.
Getting Things Done by David Allen — cover

Verdict

Read the notes, not the book — but the mechanism is stronger than most critics allow. Confidence: moderate.

Two of Allen’s operational moves — write it down, and decide the next physical action — sit on top of some of the better-replicated findings in applied cognition. That deserves plain credit. The problem is not that GTD is bunk. The problem is that the 350-page apparatus around those two moves is untested, and the one thing the book promises in its subtitle — stress-free — is the specific outcome that failed to appear in the flagship experiment.

Our earlier verdict on this page said the theory underneath GTD is weaker than the confident prose suggests. That was half right. The theory Allen cites is weak. The theory that actually explains why capture and next-action work is better than he knew, and he never mentions it.

The claim on trial

Stated so it can be checked: incomplete commitments held only in memory consume attention continuously, and moving all of them into a trusted external system releases that attention and reduces stress.

Allen’s compressed version is «your head is for having ideas — not for holding them». The book treats this as settled. It is three separate claims stacked: that unfinished commitments impose a cost, that externalizing them removes the cost, and that removing the cost produces calm. They have different amounts of evidence behind them, and the third has the least.

The cost of unfinished commitments

This part checks out, and the evidence is more specific than «you feel scattered».

Masicampo and Baumeister ran five studies on this in Masicampo and Baumeister, Journal of Experimental Social Psychology, 2011, with samples of 87, 83, 47, 38 and 54 undergraduates. Participants primed with an unfulfilled goal solved fewer anagrams, ate more cookies when dieting, and did worse on logic problems. They did not do worse on general-knowledge questions. That dissociation is the useful part: the cost lands on tasks needing executive control, not on retrieval. Study 5 showed the interference disappeared once the frustrated goal was completed.

Sophie Leroy’s two experiments in Leroy, Organizational Behavior and Human Decision Processes, 2009 found the same shape at work and named it attention residue. People who switched away from an unfinished task carried part of their attention with them and did worse on the next one. Her twist is one Allen would not like: finishing the first task was not enough on its own. Time pressure while finishing it was what let people disengage.

The diagnosis in «Getting Things Done» is sound. Open commitments are not free.

Where the Zeigarnik foundation collapses

Allen’s fans, and much of the productivity internet, credit this to the Zeigarnik effect: Bluma Zeigarnik’s 1927 finding that interrupted tasks are recalled better than completed ones. That specific claim does not survive.

Colin MacLeod’s history in MacLeod, Memory & Cognition, 2020 reconstructs the original work — 32 adults given 22 tasks in the first experiment, with an interrupted-to-completed recall ratio near 2.0 — and then documents seventy years of trouble. Alper reported in 1952 that «few investigators could unequivocally reproduce Zeigarnik’s findings». Van Bergen’s 1968 review found fewer than a third of 44 papers reported the effect at all, and none of her own seven new experiments verified it.

The 2025 meta-analysis settles it. Ghibellini and Meier, Humanities and Social Sciences Communications, 2025 pooled 59 publications. For the 38 that used the standard recall-ratio measure, the weighted ratio was 0.99 — interrupted and finished tasks recalled about equally. Where effect sizes could be computed from 8 publications, the average was dz = 0.15. Their conclusion: «the Zeigarnik effect lacks universal validity».

The same meta-analysis found the companion Ovsiankina effect alive and well: across 21 publications, people resumed interrupted tasks 67% of the time. That is the finding that survives. Unfinished things pull you back toward doing them. They do not necessarily sit in memory glowing.

This matters for GTD because the mechanism Allen implies — that your mind keeps re-presenting undone items — is not the mechanism the data supports. What the data supports is that unfinished goals compete for executive control. Same practical advice, different physiology, and a book that confidently describes the wrong one.

What externalizing actually buys you

This is the strongest part of the case, and Allen gets full credit for it.

Sam Gilbert ran four online experiments with 1,196 participants in Gilbert, Quarterly Journal of Experimental Psychology, 2015, using a task where people could either hold a delayed intention in mind or externalize it by setting a reminder. Setting the reminder improved performance. People set reminders adaptively — more often under high memory load, more often when distraction was likely — and performance on the lab task predicted whether they fulfilled a real intention embedded in their lives up to a week later.

Benjamin Storm and Sean Stone found something adjacent in Storm and Stone, Psychological Science, 2015: across three experiments, saving one file before studying a new one improved memory for the new file. The benefit vanished when the saving process was described as unreliable. That last detail is the interesting one for GTD readers. Allen insists the system must be trusted or it does not work. He is right, and this is the experiment that shows why: the offloading benefit is conditional on believing the external store will hold.

The broader review, Risko and Gilbert, Trends in Cognitive Sciences, 2016, sets out cognitive offloading as a general strategy with real gains and real costs.

And here is the cost GTD never mentions. In Gilbert and colleagues, Journal of Experimental Psychology: General, 2020, participants chose between a larger reward for remembering unaided and a smaller reward per item for using reminders. They were significantly biased toward reminders, even with money on the line for choosing correctly. The bias was stable across time and predicted by underconfidence in their own memory. It disappeared when participants were given accurate metacognitive advice. People do not need persuading to externalize. They already over-externalize, and the reason is that they underrate their own heads.

GTD as a package has never been tested

The only serious scientific treatment is Heylighen and Vidal, Long Range Planning, 2008, which is a theoretical defence, not an evaluation. It contains no participants and no data. It argues from situated cognition, distributed cognition, stigmergy and flow that GTD ought to work. The authors were candid about the gap. They were explicit about it in the paper: «In spite of the many testimonials that GTD works in practice, however, as yet no academic papers have investigated this method.» And: «While it would be interesting to test GTD empirically, e.g. by comparing the productivity of people using GTD with the one of people using different methods, this is intrinsically difficult.» Their stated reason is that GTD refuses explicit priorities, so there is no obvious yardstick.

Eighteen years on, that sentence still stands. Searching Crossref by title and by bibliographic query, and Europe PMC, on 10 August 2026, produced no controlled evaluation of the method. Europe PMC returns five records mentioning «Getting Things Done» alongside David Allen; all five are advice columns, interviews or career essays. Neither index returns a single study of a two-minute completion threshold or of a weekly review ritual. Those two are folklore — plausible, widely practised, unmeasured. Untested is not refuted. It is also not evidence.

Two tested cousins do exist. Implementation intentions — if-then plans specifying when, where and how — produced a medium-to-large effect on goal attainment across 94 studies, d = 0.65, in Gollwitzer and Sheeran, Advances in Experimental Social Psychology, 2006. That meta-analysis predates preregistration norms, so treat 0.65 as an upper bound. Still, GTD’s «next action» plus its context lists (@calls, @errands) is a rough hand-built version of the same if-then structure, and it is the part of the book with the best pedigree.

The other cousin is time management as a whole. Aeon, Faber and Panaccio, PLOS ONE, 2021 pooled 158 studies, 490 effect sizes and 53,957 participants. Time management correlated with job performance at r = .259 (k = 21, N = 3,990), wellbeing at r = .313 (k = 30, N = 9,905) and life satisfaction at r = .426 (k = 9, N = 2,855). Its link to distress was negative but modest, r = −.222. The pattern is worth noticing: time management tracks feeling better more strongly than producing more. Only three studies in the whole pool evaluated actual training programmes, r = .173, N = 846.

The part of the subtitle that failed

«The Art of Stress-Free Productivity» is the promise on the cover, and it is the claim that broke in the lab.

Masicampo and Baumeister, Journal of Personality and Social Psychology, 2011 is the paper everyone reaches for. It earns its reputation. In Study 1, 73 undergraduates were split into unfulfilled-task, plan and control conditions. Those who wrote out how, when and where they would do the task reported fewer intrusive thoughts (M = 1.77 versus 3.00, F(1,66) = 6.41, p = .014) and read better afterwards (M = 6.94 versus 6.13, F(1,66) = 5.29, p = .025). Mind-wandering during reading fell from 65.2% to 33.3%. Intrusive thoughts fully mediated the comprehension gain. Study 4, with 97 participants, showed the same on anagrams (M = 9.55 versus 6.55, F(1,94) = 6.60, p = .012).

Then come Studies 5A and 5B, with 174 and 80 participants. They tested whether the benefit runs through emotion. It does not. Plans did not reduce anxiety or negative affect. Participants reported the same anxiety whether or not they had planned — and the intrusive thoughts still went away.

Read that against the cover. Planning cleared the mental interruptions and improved cognitive performance. It did not make anyone calmer. «Mind like water» is not what the evidence delivered. What it delivered was a working memory that stopped being nagged, which is a smaller and more useful thing.

Who should actually read it

  • People whose work arrives from many directions with no single queue — managers, freelancers, clinicians with administrative loads. The capture-and-clarify loop is built for exactly that mess.
  • People who already keep lists and still miss commitments. The book’s diagnosis — that a list of vague nouns is not a system — is correct and worth the read on its own.
  • Not people looking for relief from anxiety. The most-cited experiment behind GTD found no effect on anxiety. If that is the problem, this is the wrong intervention.
  • Not people whose difficulty is deciding what matters. GTD is deliberately agnostic about priority. Heylighen and Vidal named that as the reason it resists measurement, and it is also why it can leave you efficiently busy on the wrong things.

One thing to try

Take the five commitments currently occupying the most head-space. For each, write one sentence in if-then form: the trigger, the place, and the single next physical action. «When I open my laptop tomorrow morning, I will call the clinic and book the appointment» — not «sort out clinic». That is Allen’s next action fused with the implementation-intention format that carries d = 0.65 across 94 studies. Then put the five sentences somewhere you will reliably see them, because the offloading benefit in Storm and Stone’s experiments disappeared when the store was not trusted.

Expect fewer interruptions in your thinking. Do not expect to feel calmer. That is not what the data promised.

Get the book

Find «Getting Things Done» on Amazon — as an Amazon Associate, The Boring Work earns from qualifying purchases (disclosure).

When to see a professional

This is general information about a productivity book and the research around it, not medical or mental-health advice. Persistent intrusive thoughts, sustained anxiety, or an inability to start or finish ordinary tasks are not organizational problems and will not be fixed by a better list system. Talk to a GP, a clinical psychologist, or a psychiatrist. The 2011 finding that planning removed intrusive thoughts without touching anxiety is a reminder that these are separable, and only one of them is a filing question.

The boring bottom line

«Getting Things Done» diagnosed a real cost, prescribed two moves that later research vindicated, attributed them to a 1927 effect that mostly does not replicate, and promised an emotional payoff that its own flagship study did not find. The system as a whole has never been evaluated by anyone, including the two academics who wrote its defence. Take the capture habit and the next-action habit, write them as if-then sentences, and skip the taxonomy. More of our book reviews work the same way.

Sources

Leave a Reply