Predictably Irrational, Rerun: What Happened When Other Labs Tried

TL;DR

  • Three of the book’s showpiece demonstrations have since been rerun by other people. The decoy shrank to number grids, the social-security-number anchor on willingness to pay did not hold up, and the Ten Commandments experiment came back at d = −0.04 across 19 laboratories.
  • Ordinary anchoring, where a number you have just seen drags your estimate, replicated consistently in a coordinated project spanning 36 samples. The flashy version and the textbook version had different fates.
  • In 91 attempts across 23 product classes, the decoy effect appeared 11 times, and dropped to chance when products were shown with pictures or words instead of numbers.
  • Defaults, the least theatrical example in the book, are still standing.
  • A 2012 paper from the same research programme, used in a later book, has been retracted; a replication of it found nothing.
Predictably Irrational by Dan Ariely, cover

Verdict

Skip it, and take the effects that survived from somewhere that also tells you their limits. A book of advice can be read with a pinch of salt; a book of facts about your own mind cannot, because nothing on the page tells you which of those facts still hold.

Verdict card for Predictably Irrational reading Skip: three signature experiments, three failed tests
Behavioural economics is not the problem here. This particular set of demonstrations is.

A book that sold proof

Dan Ariely, a behavioural economist then at MIT, published the book in 2008 and did something most popular psychology writers avoid. He did not ask you to trust his opinions. He showed you experiments, one per chapter, each built so that you would recognise yourself in the people who got fooled. A free chocolate beats a cheap one by more than the price gap should allow, and a reminder of the Ten Commandments stops students from cheating.

That format is the book’s strength and its exposure. If the experiments hold, the chapters hold. If an experiment fails when someone else runs it, there is no fallback position of «well, it’s only advice», because the experiment was the argument. So the fair way to review this book is to collect what happened when other laboratories reran its showpieces, which is what this page does, chapter by chapter.

Eight claims from Predictably Irrational with verdicts from confirmed to refuted
Two confirmed, two partial, two untested, two refuted; the refuted pair are among the book’s most quoted pages.

The magazine, the decoy and 91 attempts

The best-known example comes from Ariely’s own talks. A magazine ad offered three subscriptions: «an online subscription for 59 dollars, a print subscription for 125 dollars, or you could get both for 125» (Ariely, 2008). Nobody wants print alone at that price. It sits on the menu to make the bundle look like a bargain, and when Ariely gave the menu to 100 MIT students with and without that option, the most popular choice and the least popular one swapped places.

The effect is real. Its habitat turned out to be small. One team ran a series of studies and concluded that the attraction effect «may be restricted to stylized product representations in which every product dimension is represented by a number», and that it does not typically appear «when consumers experience the product … or when even one of the product attributes is represented perceptually» (Frederick et al., 2014).

A second, independent group counted. «Ninety-one attempts to produce an attraction effect (involving a total of 23 product classes and 73 different decoyed choice sets) produced only 11 reliable effects», which is fewer than the studies’ own statistical power predicted, and with verbal descriptions or pictures the effect fell to chance (Yang & Lynn, 2014).

Two bars: 91 attempts to produce the attraction effect and 11 reliable effects
Yang & Lynn, 2014. Fewer than the studies’ own power would predict, and near zero with realistic stimuli.

The researchers who first described the attraction effect in 1982 replied in the same journal issue (Huber et al., 2014), so the argument is live and both sides have had their say. What nobody on either side disputes is the boundary. A price table with three columns of numbers: yes, a decoy can tilt it. A shelf of wine you can see and hold: mostly no.

That boundary is more useful than the chapter. It tells you to look twice at subscription tiers, phone plans and insurance grids, which is where the effect lives, and to stop worrying about it in a shop. The magazine example was well chosen, since a pricing table is the most decoy-friendly format there is. The step from that table to «this is how people choose» is the step the later evidence does not take.

Social security numbers and a bottle of wine

The anchoring chapter rests on an experiment in which students wrote down the last two digits of their social security number and then bid on wine, chocolate and computer accessories. Students with higher digits bid more, and the original paper read this as evidence of «coherent arbitrariness»: preferences that are internally consistent and anchored on nothing in particular (Ariely et al., 2003).

Two teams of economists later ran that kind of design again, larger. One examined «the strength of certain anchoring results» and found them weaker than the originals, using the case to argue for caution about first findings in general (Maniadis et al., 2014). The other tested anchoring on willingness to pay and willingness to accept and reported that the effects were not robust (Fudenberg et al., 2012). A working paper by three other researchers argued that the larger replication is compatible with the original once read properly (Simonsohn et al., 2013); it was not peer-reviewed, and it belongs in the record anyway.

Set that next to a different kind of anchoring. A coordinated project reran 13 classic effects in 36 independent samples with 6,344 participants, and «in the aggregate, 10 effects replicated consistently»; the clear failures were flag priming and currency priming, with imagined contact weak (Klein et al., 2014). Anchoring was one of the ten. Those studies asked people to estimate things such as a mountain’s height after seeing a high or low number, which is the version in every psychology textbook.

Grid contrasting judgment anchoring, which replicated everywhere, with the willingness-to-pay anchor, which did not
Klein et al. 2014 against Maniadis et al. 2014. One of these is in every psychology textbook for good reason.

Hold those two sentences apart. «A number you just saw pulls your next estimate toward it» has been tested in dozens of laboratories at once and stands. «The last digits of your social security number set what you will pay for a bottle of wine» is the book’s version, the one people retell at dinner, and it is the one that wobbled when economists reran it. The book prints both in the same confident voice, and a reader has no way to see the seam.

Nineteen laboratories and the Ten Commandments

The honesty chapters describe the experiment readers remember best. Students recalled either the Ten Commandments or ten books they had read in school, then got a chance to over-report how many puzzles they had solved for cash. In the original, the moral reminder group claimed 1.45 fewer solved matrices, Cohen’s d = 0.48 (Mazar et al., 2008). A d of 0.48 is a lot for a five-minute recall task: it means roughly two out of three students in the reminder group cheated less than the average student in the comparison group.

Years later, a registered replication ran the same protocol 25 times, total N = 5,786. Every laboratory agreed on the method before collecting a single data point. In the primary analysis of 19 replications with 4,674 participants, «participants who were given an opportunity to cheat reported solving 0.11 more matrices if they were given a moral reminder than if they were given a neutral reminder (95% confidence interval = [−0.09, 0.31])», a small effect pointing the opposite way from the original, d = −0.04 (Verschuere et al., 2018). The original authors published a response alongside it (Amir et al., 2018).

Two bars: original effect size 0.48 and replication effect size minus 0.04
Mazar et al. 2008 against Verschuere et al. 2018, 19 laboratories and 4,674 participants.

Two features make this heavier than one failed study. The protocol was fixed and published before data collection, so nobody could keep analysing until something appeared. And 19 separate laboratories ran it, so a single unlucky room cannot explain the result. When a design like that returns an interval sitting on zero, the honest reading is that the effect, in the form the book describes, is not there.

The broader idea in those chapters, that most people cheat a little and that cues can nudge them toward honesty, is not settled by this one replication. The specific demonstration the book uses to prove it is.

What is still standing

The quietest page in the book has aged best. European countries with similar cultures register very different shares of potential organ donors, and the difference follows whether the form asks you to opt in or to opt out: in one well-known comparison, about 12 per cent effective consent in Germany against 99.98 per cent in Austria (Johnson & Goldstein, 2003). Defaults move what people end up signed up for, and that finding has been built on ever since rather than knocked down.

Textbook anchoring survived as well, as above. Two more chapters sit in a third category. The claim that a price of zero behaves as a special psychological category (Shampanier et al., 2007) and the day-care study behind the chapter on social and market norms, where a fine for late pick-up increased lateness (Gneezy & Rustichini, 2000), have no independent direct replication that I could find. Untested is a different verdict from refuted, and both should be stated precisely.

Line the four survivors up against the casualties and a pattern appears that is worth more than any single chapter. What survived was old, replicated elsewhere and dull: the default box on a form, a number seen just before a guess. What shrank or vanished was what the book existed to deliver: the surprising result from a single laboratory that you repeat at dinner.

That link between memorable and fragile has a mechanism. A study gets famous because its effect is large. Effects come out large most easily when samples are small, the setting is artificial and the analysis had room to move, and those are the conditions under which a result is least likely to survive a rerun. Carry that rule out of this book; it works on most of the shelf.

A note on the research programme

One more fact belongs here, stated narrowly. A 2012 paper from the same research programme, used in a later book rather than this one, reported that signing an honesty pledge at the top of a form instead of the bottom reduced cheating. A large replication found no such effect (Kristal et al., 2020), and the 2012 paper has since been retracted.

That is a fact about a different paper in a different book, and it is not evidence about the experiments in this one. It is here because a reader deciding how much to trust a 2008 book is entitled to know the later record of the programme it came from, and because it is the reason this site does not use that line of research as support for anything.

Nothing on this page claims anything about how any particular result came to be what it was. Every verdict above rests on published replications run by people other than the original authors, with the original authors given space to reply in the same journals.

Who should read it

Read it if you want to see popular behavioural economics at its peak, and what happens to a genre built on demonstrations once the demonstrations are rerun. As a period piece it is excellent. Ariely writes an experiment the way a good teacher tells a story, with you placed inside it, and few science books manage that.

Do not use it as a reference. The two effects worth keeping, defaults and textbook anchoring, are better learned from sources that state their limits, and the decoy needs the boundary the chapter leaves out. If you keep one page from the book for the next decade, keep the organ-donation forms. It is the least dramatic thing in it, and nobody has had to defend it since.

The book’s priming cousins met the same fate in another bestseller: Thinking, Fast and Slow, checked chapter by chapter, where the author retracted a chapter himself. For the same pattern in the willpower literature, see does willpower run out like a muscle. The full shelf of checked books is under book reviews.

The boring bottom line

Ninety-one attempts at the decoy produced 11 reliable effects, and those faded once products stopped being grids of numbers. The social-security-number anchor did not survive the economists who reran it, while textbook anchoring replicated across 36 samples. The Ten Commandments experiment returned d = −0.04 in 19 laboratories against an original 0.48.

What holds is the unglamorous half: defaults change what people sign up for, and a number you have just seen changes the number you are about to say. Both were known before the book, and neither needed it.

Sources

  • Ariely, D (2008). Are we in control of our own decisions? — TED talk transcript. TED. Read on 22 August 2026, so the claims under test are the author's own wording rather than a summary of the book. ted.com
  • Frederick, S., Lee, L., & Baskin, E (2014). The limits of attraction. Journal of Marketing Research. 51(4), 487-507. Where the decoy effect lives, and where it stops. doi:10.1509/jmr.12.0061
  • Yang, S., & Lynn, M (2014). More evidence challenging the robustness and usefulness of the attraction effect. Journal of Marketing Research. 51(4), 508-513. An independent series reaching the same conclusion as the study above. doi:10.1509/jmr.14.0020
  • Huber, J., Payne, J. W., & Puto, C. P (2014). Let's be honest about the attraction effect. Journal of Marketing Research. 51(4), 520-525. Included so the exchange appears as an exchange rather than a verdict. doi:10.1509/jmr.14.0208
  • Ariely, D., Loewenstein, G., & Prelec, D (2003). «Coherent arbitrariness»: Stable demand curves without stable preferences. Quarterly Journal of Economics. 118(1), 73-106. The primary source of the book's anchoring chapter, included strictly as the object of examination. doi:10.1162/00335530360535153
  • Maniadis, Z., Tufano, F., & List, J. A (2014). One swallow doesn't make a summer: New evidence on anchoring effects. American Economic Review. 104(1), 277-290. The direct test of the book's most striking experiment. doi:10.1257/aer.104.1.277
  • Fudenberg, D., Levine, D. K., & Maniadis, Z (2012). On the robustness of anchoring effects in WTP and WTA experiments. American Economic Journal: Microeconomics. 4(2), 131-145. An independent check on the same paradigm, in the same direction as the replication above. doi:10.1257/mic.4.2.131
  • Simonsohn, U., Simmons, J. P., & Nelson, L. D (2013). Anchoring is not a false-positive: Maniadis, Tufano, and List's (2014) «failure-to-replicate» is actually entirely consistent with the original. SSRN working paper. A working paper rather than a peer-reviewed article, and labelled as such here. Included because the dispute is live and a reader deserves both sides. doi:10.2139/ssrn.2351926
  • Klein, R. A., Ratliff, K. A., Vianello, M., Adams, R. B., Jr., Bahník, Š., et al (2014). Investigating variation in replicability: A «many labs» replication project. Social Psychology. 45(3), 142-152. Read in an open repository copy on 22 August 2026 to confirm that anchoring was among the effects tested and among those that held. doi:10.1027/1864-9335/a000178
  • Mazar, N., Amir, O., & Ariely, D (2008). The dishonesty of honest people: A theory of self-concept maintenance. Journal of Marketing Research. 45(6), 633-644. Included only as the object of the replication below. This site does not cite this line of work as evidence for anything. doi:10.1509/jmkr.45.6.633
  • Verschuere, B., Meijer, E. H., Jim, A., Hoogesteyn, K., Orthey, R., et al (2018). Registered Replication Report on Mazar, Amir, and Ariely (2008). Advances in Methods and Practices in Psychological Science. 1(3), 299-317. A direct test of one of the book's most repeated demonstrations, with the original effect at d = 0.48 and the replication at d = −0.04. doi:10.1177/2515245918781032
  • Amir, O., Mazar, N., & Ariely, D (2018). Replicating the effect of the accessibility of moral standards on dishonesty: Authors' response to the replication attempt. Advances in Methods and Practices in Psychological Science. 1(3), 318-320. Listed so the reply is visible next to the replication. doi:10.1177/2515245918769062
  • Johnson, E. J., & Goldstein, D (2003). Do defaults save lives?. Science. 302(5649), 1338-1339. Effective consent to organ donation across European countries by opt-in versus opt-out default; the Germany and Austria figures (12% and 99.98%) are from this paper's comparison. doi:10.1126/science.1091721
  • Shampanier, K., Mazar, N., & Ariely, D (2007). Zero as a special price: The true value of free products. Marketing Science. 26(6), 742-757. The source of the book's free-products chapter. No independent direct replication was located, which is why that claim is graded untested here. doi:10.1287/mksc.1060.0254
  • Gneezy, U., & Rustichini, A (2000). A fine is a price. Journal of Legal Studies. 29(1), 1-17. The source of the social-versus-market-norms chapter. No independent direct replication was located. doi:10.1086/468061
  • Kristal, A. S., Whillans, A. V., Bazerman, M. H., Gino, F., Shu, L. L., Mazar, N., & Ariely, D (2020). Signing at the beginning versus at the end does not decrease dishonesty. Proceedings of the National Academy of Sciences. 117(13), 7103-7107. Cited once, in a paragraph about the wider research programme rather than about this book. The 2012 paper it examined has since been retracted. doi:10.1073/pnas.1911695117