Pattern/Game Theory/No. 0773
Prisoner’s Dilemma
The prisoner’s dilemma is a game in which each of two players benefits from defection regardless of the other’s choice, yet both gain more from mutual cooperation. Developed by Merrill Flood and Melvin Dresher, it makes mutual defection a Nash equilibrium.
- Evidence
- Well established
- Read
- 6 min
- Links
- 11 connections
01You've seen this when…
- in life
You and a friend agree to put your phones away at dinner. Each sneaks a look while the other talks; soon neither feels listened to.
- at work
Two rival stores keep extending their opening hours to capture each other’s customers. Neither wants to close first. Both pay for longer shifts without much increase in total sales.
- out in the world
Two countries expand their arsenals because neither wants to fall behind. Each increase gives the other a reason to spend more. Both budgets grow without either country feeling safer.
02The idea
People can see the better outcome clearly while each person’s incentives favor the worse shared outcome. That gap makes the dilemma frustrating.
The prisoner’s dilemma has two players, each choosing between cooperation and defection. Defection means taking the self-interested option. A player can take that option while keeping promises and avoiding betrayal.
Consider these illustrative rewards, where more points are better:
- Both cooperate. Each gets 3 points.
- Only one defects. The defector gets 4 points; the cooperator gets 0.
- Both defect. Each gets 1 point.
If the other person cooperates, you get more by defecting: 4 rather than 3. If the other defects, you still get more by defecting: 1 rather than 0. Defection is a dominant strategy, meaning it pays better regardless of the other’s choice.
Yet following that strategy leaves both with 1 point instead of the 3 each could have received. Mutual defection is a Nash equilibrium: neither benefits by changing their choice alone. That makes it stable, not good.
03Why it happens
- The individual comparison favors defection. Each person compares their own two options while holding the other’s choice fixed. In both comparisons, defection wins. The jointly better outcome requires both to change.
- Cooperation leaves you exposed. If you contribute and the other person doesn’t, you receive the worst outcome. The risk persists even with good intentions.
- Assurances can leave the rewards intact. Even if you believe the other person will cooperate, the original payoff structure still rewards you for defecting. In the strict, one-round game, defection still pays better when players trust each other.
- The future may carry too little weight. A later loss of trust, business or support can make defection expensive. But that consequence matters only if another encounter is likely, behavior is observable, and future benefits are worth enough now.
The pattern can arise among well-informed people who communicate clearly and act with goodwill. Even perfect mutual understanding can coexist with incentives that pull them toward the worse shared outcome.
04A worked example
Consider an invented case. Two software teams can each build a separate reusable testing tool. Each tool costs its builder 10 hours and saves each team 8 hours during the quarter. Either team can use either tool. Assume the savings add together and there are no other benefits or penalties.
What it looks like Each team sensibly protects its limited engineering time. Building a tool costs 10 hours but returns only 8 hours to its own team.
What’s actually going on A team that builds alone ends up 2 hours behind, while building both tools leaves each team 6 hours ahead after saving 16 hours and spending 10. A team that builds nothing gains 8 hours for free if the other builds; when neither builds, both gain nothing.
For either team, refusing to build is individually better. If the other builds, refusing yields 8 hours rather than 6. If the other doesn’t, refusing yields 0 rather than minus 2. Both refuse, even though both building would save them 12 hours in total.
What would have helped A shared budget could credit builders for the time they save the other team. An enforceable reciprocal agreement could also work. Either approach changes the rewards. Simply asking the teams to be more generous or saying that everyone benefits leaves the incentive problem in place when each team’s decision is judged only on its own hours.
05How to spot it
Check all four conditions before diagnosing a prisoner’s dilemma. Two people can end up unhappy for other reasons.
06What to do about it
- Change what each person gains or loses. Reward contributions, share their costs, or attach a credible consequence to defection. Mechanism design asks how rules can make individually attractive choices serve the shared goal.
- Make reciprocity enforceable where possible. Contracts, conditional exchanges and deposits can keep one side from taking the benefit without supplying its part. A strategic commitment requires an actual constraint on a later choice; announcing good intentions leaves that choice open.
- Give the relationship a future. In repeated games, today’s defection can cost tomorrow’s cooperation. Make continued dealings worthwhile and avoid rewards that encourage a final grab before departure.
- Make behavior visible enough to respond to. Reciprocity fails when nobody can tell whether someone contributed. Use observable commitments, while allowing for mistakes and events outside a person’s control.
- Keep responses proportionate and repairable. Tit for tat illustrates conditional cooperation: begin cooperatively, then respond to the other’s previous move. But copying every apparent defection can turn misunderstandings into lasting retaliation. Leave a route back.
If the incentives are beyond your control, recognize the exposure. Cooperation may still express your values, even when defection gives you the better individual payoff in the original game.
07When it isn’t a prisoner’s dilemma
In a coordination game, your best choice depends on what the other chooses. You may mainly need to agree on a common standard. In a prisoner’s dilemma, agreement without enforcement leaves the incentive to defect intact.
A public goods game often captures a related contribution problem with many participants. The term covers more than this particular two-player payoff structure.
Repetition makes cooperation possible. Whether players cooperate depends on the game’s conditions. If a fixed final round is known, reasoning backward can unravel cooperation under standard assumptions about rationality and knowledge. The analysis can change when the ending is uncertain, players care about their reputations or their preferences extend beyond immediate rewards.
Finally, cooperation between players can harm outsiders. Competing firms might both profit by coordinating prices while harming customers and breaking the law. The model identifies an incentive problem; deciding whose interests deserve priority requires a separate judgment.
08Roots
At RAND in 1950, Merrill Flood and Melvin Dresher were studying situations where individually sensible decisions produced disappointing joint results. One experiment had two participants make choices over 100 rounds. The payoff structure favored defection, but actual play included cooperation. From the start, the puzzle concerned both the incentives on paper and what people did when facing them repeatedly.
Albert Tucker supplied the memorable story: two suspects questioned separately, each offered a better personal outcome for informing on the other. Both would receive lighter punishment if both stayed silent than if both informed. Tucker invented the prisoner scenario as an explanatory fiction. The story gave the mathematical pattern its name.
The idea later became a laboratory for studying cooperation. Robert Axelrod invited researchers to submit strategies for computer tournaments of the repeated game. Simple reciprocal strategies performed strikingly well in that setting. The lesson traveled into economics, political science and organizational life: the same immediate temptation can lead to different behavior when today’s partner is also tomorrow’s.
09How solid is this?
The one-round incentive result follows mathematically from the specified payoffs. Experiments document both cooperation and defection; their frequency depends on repetition, information, framing and participants’ preferences. Applying the label to a real conflict requires checking its actual incentives.
10Connections
- Often confused withCoordination Game
- Countered byMechanism Design, Strategic Commitment, Shadow of the Future, Tit for Tat
- Part ofPublic Goods Game, Zero-Sum vs. Non-Zero-Sum
- See alsoDominant Strategy, Nash Equilibrium, Repeated Games, Backward Induction
+ 1 more in the list
11Origin and sources
Merrill Flood and Melvin Dresher developed the game at RAND in 1950. Albert Tucker supplied the prisoner interpretation and name that year.
- [1]Flood, M. M. (1958). Some Experimental Games. Management Science, 5(1), 5–26.
- [2]Rapoport, A., & Chammah, A. M. (1965). Prisoner's Dilemma: A Study in Conflict and Cooperation. University of Michigan Press.
- [3]Axelrod, R. (1980). Effective Choice in the Prisoner's Dilemma. Journal of Conflict Resolution, 24(1), 3–25.
Suggest an edit· Updated 2026-10-02