Tool/Mental Model/No. 0820
Red Teaming
Red teaming is a process in which an independent group tests a plan or system by acting as an opponent and challenging its assumptions. Rooted in military war-gaming, it is used in security, intelligence analysis, and organizational planning to expose weaknesses before real threats do.
- Evidence
- Useful, modest evidence
- Read
- 6 min
- Links
- 10 connections
01You've seen this when…
- in life
You plan to buy a rental property. Before making an offer, you ask a friend with no stake in the deal to attack your cash-flow estimate using vacancies, repairs, and lower rents.
- at work
Your team approves a new subscription price. A separate group plays demanding customers and competing sellers, then shows how each could make the offer unprofitable.
- out in the world
A city rehearses its heatwave response. Reviewers outside the planning team play residents who lack cars, smartphones, or English fluency, and discover that several routes to help don’t work.
02The idea
A plan can survive months of review because everyone reviewing it shares the same assumptions. Red teaming changes the assignment: give a separate person or group permission to challenge the plan by behaving like an opponent or testing how it fails.
Independent challengers need a clear target and a realistic challenge. They can come from inside or outside the organization. Independence means they can report unwelcome findings without needing the original plan to succeed or needing to please its author.
A red team might simulate a competitor’s response, examine an intelligence judgment from another country’s perspective, or attempt an authorized breach of a security system. These activities differ, but all bring resistance into the planning process before it arrives for real.
This is more than inviting criticism. A reviewer might notice a questionable assumption. A red team tries to demonstrate what happens when that assumption is wrong. That makes it a practical counterweight to confirmation bias and groupthink.
A pre-mortem asks people to imagine failure and explain it. Red teaming assigns challengers to investigate or simulate the challenge. Stress testing pushes a system under demanding conditions; red teaming also asks which conditions an intelligent opponent would choose.
03How to use it
- Define the decision and the boundaries. State what is being tested: a launch, a security control, a forecast, or a strategy. Agree on permitted methods, protected information, and stopping conditions. Never treat the label as permission to deceive uninvolved people or probe systems without authorization.
- Choose challengers with relevant knowledge and freedom to disagree. Include people who understand the system but didn’t build the case for it. Add a different perspective where useful. Outsiders can spot blind spots, but unfamiliarity alone isn’t expertise.
- Give them a specific assignment. Specify who the opponent is and what they aim to achieve under the constraints they face. For a non-adversarial review, name the assumption or failure condition to investigate. A vague request to find problems produces vague objections.
- Let them work before the joint discussion. Provide documents, access, and enough time to investigate. Ask them to distinguish observed failures, plausible scenarios, and unsupported guesses. Where feasible, turn claims into small, authorized tests using a falsification test.
- Review findings without scoring personalities. Have the plan’s authors explain their reasoning and the challengers show their evidence. Decide which findings require a change, further investigation, or explicit acceptance. The aim is to help colleagues reach a better decision together.
- Assign fixes and retest. Give important findings an owner and a deadline. Repeat the relevant challenge after changes. Carry the report’s findings through to resolution.
04A worked example
In 2015, the US Department of Homeland Security’s inspector general reported on covert testing of Transportation Security Administration airport checkpoints. Authorized testers attempted to bring prohibited items through passenger screening. The public report described vulnerabilities in equipment and procedures and identified problems with officers’ performance. Sensitive details were withheld.
What it looks like At a working checkpoint, the line keeps moving while trained officers follow screening procedures and use scanners. Those visible activities can make the security system appear effective.
What’s actually going on The screening procedure’s success depends on detecting the prohibited items it is meant to catch. Testers pursued the outcome the checkpoint was designed to prevent. Their attempts supplied evidence about detection beyond what an inspection of written procedures alone could provide. The report recommended corrective action.
What made it work The challengers had authorization for hands-on tests of the actual defenses. The result was a set of concrete problems to address. Claims that every airport performed identically or that subsequent fixes worked would require further testing.
The transferable lesson is to test the effectiveness of security procedures. A defense needs to be evaluated against realistic attempts to defeat it, not just checked for compliance with its own instructions.
05When to reach for it
06When it misleads
- The team becomes a ceremonial opposition. Leaders commission criticism but punish anyone who changes the decision. Without psychological safety and a route to action, the exercise mostly produces paperwork.
- Winning replaces learning. Challengers seek clever gotchas; defenders conceal weaknesses. Agree in advance on what evidence matters. An adversarial collaboration can help both sides design a test they will accept.
- The opponent is unrealistic. A team with unlimited resources can defeat almost anything. Model the threat’s capabilities in a team that faces matching incentives and constraints, and label more extreme scenarios separately.
- Possibility gets mistaken for probability. A demonstration shows how failure can occur. Estimating how often it will occur requires separate evidence. Combine findings with estimates of exposure and consequences that account for uncertainty before deciding what to fix first.
- A clean result creates false confidence. Finding nothing means this team found nothing using these methods. That result leaves the broader question of safety unresolved. Rotate perspectives and test again when the system or its opponents change.
07Roots
In 1824, Georg Heinrich Rudolf von Reisswitz introduced a war game to the Prussian military using terrain maps, troop blocks, and an umpire. Officers could issue orders and see how events unfolded against another side. The problem was practical: a plan on paper doesn’t tell you what happens when an enemy reacts. These games were ancestors of modern red teaming, not its single moment of invention.
Military exercises developed opposing forces, conventionally labeled red against blue. At their best, the red side could act like a capable enemy and challenge the planners’ expectations. Red teaming also expanded beyond battlefield simulation, with teams challenging assumptions and intelligence assessments and testing technical defenses. There is no universally recognized inventor of that broader practice.
A 2003 Defense Science Board report examined red teaming across the US Department of Defense, treating it as an activity worth strengthening rather than an occasional exercise. Security organizations adopted adversarial testing, while intelligence and business teams used independent challenges to examine judgments and strategies. The common problem remained the same: people who built a plan are rarely the best people to discover every way it could break.
08How solid is this?
Documented cases show that authorized adversarial tests can expose concrete weaknesses. Evidence that red teaming consistently improves final decisions or reduces real-world failures is less conclusive; results depend heavily on realism, independence, and follow-through.
09Connections
- Often confused with Stress Testing
- Helps counter Groupthink, Overconfidence Effect, Confirmation Bias
- Includes Consider-the-Opposite Strategy
- See also Pre-Mortem, Falsification Test, Psychological Safety, Adversarial Collaboration, Streisand Effect
10Origin and sources
Developed from military war-gaming and opposing-force traditions, with no single recognized inventor. Later extended into intelligence analysis, security testing, and organizational decision-making.
- [1]Perla, P. P. (1990). The Art of Wargaming: A Guide for Professionals and Hobbyists. Naval Institute Press.
- [2]Defense Science Board. (2003). The Role and Status of DoD Red Teaming Activities. Office of the Under Secretary of Defense for Acquisition, Technology, and Logistics.
- [3]U.S. Department of Homeland Security, Office of Inspector General. (2015). TSA Can Improve Aviation Passenger Screening with Enhanced Processes and Technologies. OIG-15-150.
- [4]Zenko, M. (2015). Red Team: How to Succeed By Thinking Like the Enemy. Basic Books.
Suggest an edit· Updated 2026-10-02