Do Hungry Judges Really Deny Parole?

TL;DR

  • The numbers are quoted correctly. Three researchers logged 1,112 parole rulings by eight Israeli judges over 50 days, and favourable rulings do fall from about 65% at the start of a session to nearly zero at the end, then jump back after a food break.
  • The size of the drop is the first warning: a standardised difference of about 1.96, roughly the height gap between Dutch men and Dutch women. Hunger does not move people that far.
  • Case order was not random. Prisoners without a lawyer are heard last in each session and were released 15% of the time, against 35% for prisoners with counsel.
  • A simulated judge who never gets hungry draws almost the same curve. Approvals take longer to write, so a judge who avoids starting a long case before a break ends most sessions on a refusal.
  • Larger tests find much less. Across nearly 100,000 American arraignment decisions, release rates drift about six points in two of four daily sessions, and do not change after a meal.

On a weekday morning in an Israeli prison, a man waits in a corridor with his lawyer for a door to open. Behind it are three people at a table: a judge, a criminologist and a social worker. They will spend about six minutes on his request for early release, write a short verdict, and call the next name. He does not know what number he is in the queue, or that his place in it is about to become the most quoted variable in the psychology of judging.

The graph that made it famous appeared in April 2011. Along the bottom, the order in which cases were heard. Up the side, the share of prisoners who got what they asked for. The line starts around 65%, slides to the floor, then leaps back the moment the judge returns from a snack, three times a day, like a pulse. Fifteen years of argument later, the pattern is still in the data. The sentence people staple underneath it has come apart.

ClaimSourceWhat the source saysGrade
Favourable rulings fall from ≈65% to near zero within a session, then return after a breakDanziger, Levav & Avnaim-Pesso, 2011Reported for 1,112 rulings, 8 judges, 50 days across ten monthsC
The drop is caused by mental depletion from repeated decidingSame paper, plus every later testOffered as speculation; the paper has no direct measure of the judges’ mental resourcesR
Case order was random with respect to the breaksWeinshall-Margel & Shapard, 2011Unrepresented prisoners go last and win 15% of the time, against 35% with counselR
The same swing appears in other courtsShroff & Vamvourellis, 2022; Hemrajani & Hobert, 2024A six-point drift in some sessions, none in others, no jump after a mealP
Grades: C confirmed, P partial, H untested, R refuted.

What is inside the 1,112 rulings

The study came from Shai Danziger of the Department of Management at Ben Gurion University of the Negev, Jonathan Levav of Columbia Business School, and Liora Avnaim-Pesso, also of Ben Gurion (Danziger, Levav & Avnaim-Pesso, 2011). It is careful fieldwork, worth describing before taking apart.

Eight judges, two parole boards, four prisons, 50 hearing days across ten months. The boards handle roughly 40% of all parole requests in the country, and the judges averaged 22.5 years on the bench. Each day brought 14 to 35 cases at about six minutes each, split into three sessions by a late morning snack and lunch. Across the sample, 64.2% of requests were turned down. The paper’s summary is that favourable rulings «drop gradually» within each session, from about 65% to nearly zero, then return «abruptly» to about 65% once the judge has eaten. Regressions controlling for offence severity, months served, previous incarcerations, rehabilitation programme, sex and nationality left queue position standing.

Two details matter later. Approvals were slower than refusals, 7.37 minutes on average against 5.21. And the graph is not the whole dataset: sessions have different lengths, so the authors dropped the last 5% of each one to keep the thin tail from wobbling.

The authors were careful about the cause. They wrote that they could not tell whether resting or eating did the work, that mood was never measured, and that «we do not have a direct measure of the judges’ mental resources». Depletion was a closing speculation. It became the finding on the way out into the world.

The number that should have stopped everyone

Going from 65% approval to 5% is an odds ratio of about 35. In the currency psychologists use for comparing effects, that is a standardised difference of about 1.96. Such numbers are hard to picture, so use a physical one: the average height gap between young men and young women in the Netherlands is around 13 centimetres, a standardised difference of about 2. You can see that gap across a room without measuring anybody.

Daniël Lakens, an experimental psychologist at Eindhoven University of Technology, made this his whole objection (Lakens, 2017). If a psychological mechanism moved human judgement that far, we would have built our lives around it already: «If hunger had an effect on our mental resources of this magnitude, our society would fall into minor chaos every day at 11:45.» He adds a hole that is hard to unsee: the post-lunch slump is an ordinary, universally felt thing, and this graph shows judges at their most generous right after lunch.

Set that against the theory the 2011 authors reached for: ego depletion, self-control drawing on a shared reserve that repeated effort runs down. In 2016, twenty-three laboratories ran one pre-registered protocol on 2,141 people to test it, led by Martin Hagger, then at Curtin University in Perth. The pooled result was a standardised difference of 0.04, interval −0.07 to 0.15 (Hagger et al., 2016). An interval straddling zero means the effect could not be told apart from nothing.

So the proposed cause measures about 0.04 and its supposed consequence about 1.96. Something in that chain is wrong.

Who walks in first

The argument rests on one assumption: that case order has nothing to do with the breaks. If harder cases cluster before a break for administrative reasons, the graph draws itself.

Six months after publication, that assumption was challenged in the same journal by Keren Weinshall-Margel of the Israeli Courts Research Division at the Supreme Court in Jerusalem and John Shapard, formerly of the Federal Judicial Center in Washington (Weinshall-Margel & Shapard, 2011). Rather than build a model, they collected 227 decisions from 12 further hearing days and asked how the room works: three attorneys, a panel judge, five staff from the prison service and court management.

What they were told is mundane and devastating. The board tries to finish all the cases from one prison before it breaks, then starts the next prison afterwards. Within each session, prisoners with no lawyer are usually heard last. In their data, unrepresented prisoners were about a third of all cases and prevailed 15% of the time against 35% for prisoners with counsel. Drop the deferrals and the split is 39% against 67%.

One coding choice deserves its own line. The 2011 paper counted a refusal and a postponement as the same outcome, and postponements made up 48.4% of the refusal column. A prisoner told to come back in a month has not been refused parole in any sense a lawyer recognises.

Their conclusion: «The phenomenon of favorable decisions peaking after a meal break is likely an artifact of the order of case presentation.» They add a detail that reads oddly against the study’s framing: the decision is not made by a judge at all, but by a panel of three.

The reply, and the number it left out

The original authors answered in the same issue (Danziger, Levav & Avnaim-Pesso, 2011b). They recoded their data to include legal representation, reran every model, and reported that «the original results replicate in every analysis». Dropping deferrals gave the same. On prison batching, they said that for the five days where they had prison of origin, prisoners from one prison appeared both before and after a break.

The reply is strong on one axis and silent on another. Whether the effect survives adjustment for representation, and how large it then is, are separate questions, and the reply answers the first without touching the second. Konstantin Chatziathanasiou of the University of Münster makes the objection plainly (Chatziathanasiou, 2022): five days of prison data is not fifty, and statistical control over a variable measured in rough categories does not remove that variable’s full influence.

A judge who never gets hungry, and the same graph

The most elegant objection came five years later and needed no interviews. Andreas Glöckner, then at the University of Hagen and the Max Planck Institute for Research on Collective Goods in Bonn, built a fake judge in a computer (Glöckner, 2016). His judge is perfect: no hunger, no fatigue, no bias. She has a rough time limit for each session, and a rough sense from the thickness of a file of whether the next case will be quick or slow. When the next case would run past the limit she stops and eats, and that case opens the following session. Feed her cases in random order, with approvals taking 7.37 minutes and refusals 5.21, and she draws a downward slope of nearly the same shape and size.

The reason is arithmetic rather than psychology. Long cases are more often approvals, and as the clock runs down inside a session, fewer long cases still fit. With fifteen minutes left almost anything fits and you see the ordinary mix. With five minutes left, 12% of the approvals still fit but 46% of the refusals do. The tail of every session therefore fills with refusals, for the same reason that the last passengers squeezed into a full lift are the small ones. And since sessions tend to end just before an approval, the postponed approval opens the next session, which is where the spike comes from.

Then he turns to the dropped 5%. Trimming the last cases of a session removes the long ones, which are disproportionately approvals, so the censoring deepens the slope. Even with a judge who has no foresight and merely stops once the limit is passed, that censoring conjures a slope of about 15 points out of nothing.

His verdict is measured: the drop from 65% to almost zero «does not conclusively indicate bias or error in judicial decision making». He is equally honest about his ceiling, since time management and selective dropout «cannot account for all aspects of the data». His simulated curve starts around 45% rather than 65%, and the day’s very first case resists a postponement story. His estimate is that the mechanism buys a drop of 15 to 45 points, leaving part of the gap unexplained.

What happened in other courtrooms

Arguments about one dataset run forever. The useful move is to look elsewhere, with more cases. Ravi Shroff of New York University and Konstantinos Vamvourellis of the London School of Economics did that with arraignment hearings from a large American city between 2010 and 2015 (Shroff & Vamvourellis, 2022): nearly 100,000 decisions by 41 judges, each with at least 500 hearings. That is ninety times the Israeli sample.

They found a drift, and it is small and lopsided. Adjusted for the charge and the defendant’s record, release rates fall about six percentage points across the morning session and again across the evening one, while the post-lunch and post-dinner sessions are flat. The piece that should have shown up most clearly is missing: «release rates remain unchanged after a meal break».

A smaller study points the same way. Rahul Hemrajani and Tony Hobert examined every traffic case heard in Pulaski County, Arkansas in 2019 and 2020, and report that charges are less likely to be dismissed at the end of a session in arraignments, though the pattern does not hold for trials (Hemrajani & Hobert, 2024). Their full text sits behind a paywall, so this rests on the published abstract.

The oddest follow-up is not about judges. A team at Memorial Sloan Kettering tested whether the hour of day predicted how suspicious a report was, across 35,004 prostate scans read by 26 radiologists (Becker et al., 2024). The signal was tiny, 0.005 per hour on a five-point scale, and it points the wrong way for a depletion story: later in the day they assigned more suspicious scores, the effortful call rather than the quiet default.

What survives

The graph is honest. Those 1,112 rulings do line up that way, nobody has accused anyone of inventing anything, and the 65% is something a peer-reviewed paper reported about a real courtroom.

The clause after the comma is where it fails. «Because the judge is depleted» needs queue position to be independent of the break, and prisoners without lawyers are heard last. It needs an effect around fifty times larger than the mechanism it names has ever produced under controlled conditions. And it has to beat an explanation using no psychology at all, in which a judge with a clock draws the same curve while behaving rationally.

Chatziathanasiou concludes there is no hungry judge effect. Glöckner says his account covers large parts of the pattern rather than all of it. The honest position sits between them: the causal claim is unsupported, and how much of the graph survives once the scheduling artefacts come out is unknown, because the original data have never been released.

Verdict: P, partially supported. The 65% and the near-zero are real numbers from a real dataset, and the pattern in it is undisputed. The explanation attached to them is refuted: case order was not random, the mechanism invoked failed a 23-laboratory replication, a fatigue-free simulation reproduces most of the curve, and the largest test elsewhere finds no change in release rates after a meal.

The shift: stop reading this graph as proof that deciding drains a fuel tank, and start reading it as a lesson about queues. When outcomes vary by position in a list, ask who built the list before asking what was happening inside anyone’s head.

First move: when a finding arrives with a number that would be visible without a study, distrust the number. A gap the size of the male-female height difference in human judgement is a signal to check the design.

Second move: if you use breaks to structure your day, keep doing it. The case for pausing never rested on this study, so stop citing this study for it.

The mechanism the 2011 authors invoked has its own record, traced in what happened when self-control as a limited resource was tested by 23 laboratories at once. The glucose half of the story, where a sugary drink tops the tank back up, is covered in the evidence on sugar and self-control.

This one joins the other graded claims in our registry of checked claims, alongside the eight-second attention span and the website it came from.

The boring bottom line

Eight judges, 1,112 rulings, a line that drops from 65% to nothing three times a day. The line is there. The reason given for it is not.

Prisoners without lawyers were heard last. Approvals took two minutes longer to write than refusals, which alone pushes them towards the front of every session, and trimming the last 5% pushed them further. In a court system with ninety times as many cases, release rates do not budge after lunch.

Take your breaks. Do not tell anyone a parole board proved you should.

Sources

  • Danziger, S., Levav, J., & Avnaim-Pesso, L (2011). Extraneous factors in judicial decisions. Proceedings of the National Academy of Sciences. 108(17), 6889-6892. 1,112 rulings by eight Israeli judges over 50 days in ten months; favourable rulings drop from about 65% to nearly zero within each session and return after a food break; approvals averaged 7.37 minutes against 5.21 for refusals; the last 5% of each session was dropped from the figure; the paper states it has no direct measure of the judges' mental resources doi:10.1073/pnas.1018033108
  • Lakens, D (2017). Impossibly hungry judges. The 20% Statistician. Blog post, 3 July 2017; argues the effect is too large for any psychological mechanism, cites Glöckner's conversion of the drop to an odds ratio of 35 or d = 1.96, compares it with a Cohen's d of about 2 for the male-female height gap in the Netherlands, and notes the graph conflicts with the ordinary post-lunch dip daniellakens.blogspot.com
  • Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., et al (2016). A Multilab Preregistered Replication of the Ego-Depletion Effect. Perspectives on Psychological Science. 11(4), 546-573. Twenty-three laboratories, total N = 2,141, standardised protocol; pooled effect d = 0.04, 95% CI -0.07 to 0.15. Cited here only as the test of the mechanism the 2011 authors invoked, not as support for any claim doi:10.1177/1745691616652873
  • Weinshall-Margel, K., & Shapard, J (2011). Overlooked factors in the analysis of parole decisions. Proceedings of the National Academy of Sciences. 108(42), E833. Letter; 227 decisions from 12 further hearing days plus interviews with three attorneys, a parole panel judge and five prison service and court management staff; the board clears one prison before a break, unrepresented prisoners go last and prevail 15% of the time against 35% with counsel (39% against 67% excluding deferrals); deferrals were 48.4% of the original refusal category; decisions are made by a panel of three doi:10.1073/pnas.1110910108
  • Danziger, S., Levav, J., & Avnaim-Pesso, L (2011). Reply to Weinshall-Margel and Shapard: Extraneous factors in judicial decisions persist. Proceedings of the National Academy of Sciences. 108(42), E834. Recoded the data with legal representation and reran every model, reporting that the original results replicate in every analysis; prison of origin was available for five days only; shared counsel occurred on 39 days with two clients and 4 with three; the effect size after adjustment is not reported doi:10.1073/pnas.1112190108
  • Chatziathanasiou, K (2022). Beware the Lure of Narratives: Hungry Judges Should Not Motivate the Use of Artificial Intelligence in Law. German Law Journal. 23(4), 452-464. Open access review of the whole episode; objects that five days of prison-of-origin data is not fifty, that the reply never reports the effect size after adjusting for representation, and that statistical control over roughly measured variables does not remove their full influence; concludes there is no hungry judge effect doi:10.1017/glj.2022.32
  • Glöckner, A (2016). The irrational hungry judge effect revisited: Simulations reveal that the magnitude of the effect is overestimated. Judgment and Decision Making. 11(6), 601-610. Simulates a bias-free judge with a session time limit who avoids starting a case that will not fit; because approvals run longer, the tail of each session fills with refusals and the postponed approval opens the next session; censoring the last 5% deepens the slope and produces about 15 points even without foresight; the author states the account covers large parts but not all of the data, with his simulated curve starting near 45% rather than 65% doi:10.1017/s1930297500004812
  • Shroff, R., & Vamvourellis, K (2022). Pretrial release judgments and decision fatigue. Judgment and Decision Making. 17(6), 1176-1207. Arraignments 2010-2015 in a large urban US court system, nearly 100,000 decisions by 41 judges with at least 500 hearings each; covariate-adjusted release rates fall about 6 percentage points in the pre-lunch and pre-dinner sessions, are flat in the post-meal sessions, and remain unchanged after a meal break; agreement with prosecutor bail requests is unaffected by time of day doi:10.1017/s1930297500009384
  • Hemrajani, R., & Hobert, T (2024). The Effects of Decision Fatigue on Judicial Behavior: A Study of Arkansas Traffic Court Outcomes. Journal of Law and Courts. 12(2), 435-443. All traffic cases in Pulaski County, Arkansas in 2019-2020; charges less likely to be dismissed at the end of a session in arraignment hearings but not in trial hearings, which the authors read as context-specific. Full text is paywalled; used at abstract level only doi:10.1017/jlc.2023.21
  • Becker, A. S., Woo, S., Leithner, D., Tong, A., Mayerhoefer, M. E., & Vargas, H. A (2024). The Hungry Judge effect on prostate MRI reporting: Chronobiological trends from 35'004 radiologist interpretations. European Journal of Radiology. 179, 111665. 35,004 prostate MRI reports by 26 radiologists over eight years; a significant but very small association between hour of day and mean PI-RADS score (beta = 0.005, p < 0.001), with more suspicious scores assigned later in the day doi:10.1016/j.ejrad.2024.111665