Skip to the content
five peers
The method

What we can show you, and what we cannot.

Five Peers makes claims about how decisions get better. Below is the evidence each one rests on, with the paper, and a section at the end for the claims that are not settled. We removed three sources this site used to cite, because we checked them and they did not say what we were using them to say.

Checked at source, August 2026.
Figure in deep shadow, hand at the chin
Falsification, applied to ourselves
Six findings

Each rule the product follows, and where it comes from.

Structured disagreement

Argument beats agreement, and it costs you something

Two groups given the same strategic problem: one told to reach consensus, one told to argue it out through devil's advocacy or dialectical inquiry. The arguing groups produced better recommendations and surfaced better assumptions. They also came out of it less satisfied, less attached to the decision, and less keen to work together again.

This is why the argument runs inside the session and not at you. The quality gain is real and so is the cost, so Five Peers keeps the disagreement between the methods and hands you the output, at a bluntness you set yourself.

Schweiger, Sandberg & Ragan, Academy of Management Journal, 1986. Meta-analysis: Schwenk, Organizational Behavior and Human Decision Processes, 1990. Source

Trainable methods

What improves judgement is method, not personality

In the IARPA forecasting tournament, about an hour of training in probabilistic reasoning and scenario work improved accuracy by roughly 10% on Brier scores, and the effect held across four years. The training was concrete: state coherent probabilities, challenge your assumptions, name the causal drivers, build the best and worst case, avoid over-correcting.

That is the list the five methods are built from. It is also why Five Peers does not sell you a simulated Steve Jobs: no study shows that borrowing a famous person's voice improves a decision.

Mellers, Tetlock et al., Good Judgment Project. Training effects reported in Judgment and Decision Making and related tournament papers. Source

Why we never ask

You cannot self-report your own blind spots

People readily see bias in others and deny it in themselves, and they keep denying it when shown the evidence that they have just exhibited it. They are not performing: they genuinely believe their own read is the unbiased one. The finding has since been replicated in a preregistered study.

So the onboarding never asks which biases you have. It watches which trade-offs you actually make, on a decision that is real, and infers from there.

Pronin, Lin & Ross, Personality and Social Psychology Bulletin, 2002. Preregistered replication in Collabra: Psychology, 2024. Source

Prospective hindsight

Assume it already failed and you find more of the reasons

Asking people to explain a future event as though it had already happened, rather than to assess whether it might, increased correct identification of reasons by around 30%, and produced more concrete, action-shaped causes rather than abstract ones.

That is the pre-mortem method, one of the twelve. It is in the library because of this result, not because it makes a good slide.

Mitchell, Russo & Pennington, Journal of Behavioral Decision Making, 1989. Popularised as the pre-mortem by Klein, Harvard Business Review, 2007. Source

Changing your mind

The one disposition worth measuring

Actively open-minded thinking, the habit of going looking for the argument against your own position, is measurable with a short scale and predicts real belief updating and forecasting accuracy. It is the best evidence-to-length ratio in the whole field.

It is the one dimension the onboarding measures head-on, because it decides how hard your board should push before you stop listening.

Baron; Stanovich & West. Short form validated by Haran, Ritov & Mellers, Judgment and Decision Making, 2013. Source

Risk is not a trait

There is no such thing as your risk appetite

Risk taking is domain-specific. The same person is cautious with money and reckless with their reputation, or the reverse. A single risk score collapses five uncorrelated things into one number and tells you nothing useful.

So there is no risk slider in the onboarding. Risk is asked inside the domain of the decision you actually brought.

Blais & Weber, DOSPERT scale, Judgment and Decision Making, 2006. Source

The library

Twelve reasoning methods. You get the five you are missing.

Every method is a question asked relentlessly and a thing it refuses to let pass. None of them is a person, a job title, or a personality type. Which five you get is decided by the onboarding, and it changes as the sessions accumulate evidence about how you actually decide.

Falsification
What would have to be true for this to be wrong?

Goes looking for the evidence that would kill the plan, before the market finds it for you.

Covers · confirmation, moving straight from a hunch to a commitment
Base rates and calibration
How often does this actually work out, for people in your position?

Puts a number on the claim, then a confidence interval around the number, so the argument stops being a feeling.

Covers · overconfidence, arguing from a vivid example instead of a rate
Second-order consequences
And then what happens?

Runs the decision two and three moves out, where the cost usually lives.

Covers · short horizon, optimising the first move
Pre-mortem
It's a year from now and this failed. What killed it?

Treats the failure as already true and works backwards, which surfaces causes a risk review does not.

Covers · planning optimism, unnamed risk
Reversibility
Is this a door you can walk back through?

Separates the decisions worth agonising over from the ones worth making today and correcting later.

Covers · deliberating on everything at the same weight
Execution and constraint
Who does this, with what, by when?

Checks that the decision survives contact with the people and the calendar you actually have.

Covers · deciding in the abstract, ignoring the constraint
Interests and incentives
Who wins, who loses, and who quietly resists?

Maps the decision onto the people it moves, including the ones who will never say no out loud.

Covers · social naivety, mistaking silence for agreement
Cost of delay
What does waiting cost, in euros and in position?

Prices the option of not deciding yet, which is a decision that rarely gets costed.

Covers · avoidance dressed up as prudence
Non-consensus
What makes this different rather than marginally better?

Refuses the incremental version of the plan and asks what you believe that your market does not.

Covers · benchmarking yourself into the middle
Precedent
Who has done this before, and what happened to them?

Finds the closest real case rather than the most flattering analogy.

Covers · treating your situation as unprecedented
Externalities
Who bears a cost that isn't on your P&L?

Names the customers, staff and third parties who pay for the decision without appearing in the model.

Covers · narrow framing, reputational blind spots
Compression
Can you state this in one sentence someone can act on?

Forces the decision into a form your team can execute without a translation layer.

Covers · deciding in prose, leaving the room without a call
The composition rule

How five get picked, and how you can check it is not arbitrary.

01

A profile vector, not a personality type

The onboarding produces a small set of scores: how much you go looking for the counter-argument, how you trade speed against information, how you handle risk inside the specific domain of your decision. No type, no letter code, no colour.

Measured on the 64 profiles the journey can produce: 12 of them share no seat at all with their exact opposite, and the rest still differ by at least one.

02

A coverage matrix

Each of the twelve methods covers a specific failure mode. Falsification covers running from a hunch to a commitment. Base rates covers arguing from one vivid example. Cost of delay covers avoidance dressed up as prudence. The matrix is written down and versioned, not inferred by a model on the day.

03

Deficits first, then a diversity floor

We score where your coverage is thinnest, rank the methods that close those gaps, and impose a minimum spread so you never get five variations on the same move. Same rules for everyone, different output for everyone.

04

Tests we hold ourselves to

Retake the onboarding and the board should be broadly the same. Nudge one score slightly and the board must not flip entirely. And a personalised board has to beat a generic one on the same decisions, or the personalisation is decoration.

Status. The stability and sensitivity tests run today. The ablation against a generic board is being built with the pilot cohort. We will publish what it says, including if it says the personalisation adds little.
The record

Why the record matters more than the methods.

Five methods can be reproduced by anyone with an afternoon and a decent model. What cannot be reproduced is your own decision history with the assumptions attached: what you were betting on, whether the bet held, the pattern you keep repeating. Each session writes a structured entry, the entry gets a review date, and the review asks whether the assumption survived contact with reality.

That record is stored and structured by HOLCO, hosted in France. Five Peers orchestrates and reads it at the moment of deciding. It exports in an open format, on request, without asking us. A record you cannot take with you is not yours.

What we do not claim

The honest limits.

A method page that only lists supporting evidence is marketing. Here is what is unresolved, and what we removed.

Cognitive complementarity is our product hypothesis, not a settled theorem. The best known formal argument that diverse groups beat able ones (Hong and Page) drew a peer-reviewed mathematical rebuttal in the Notices of the American Mathematical Society in 2014, and rebuttals of the rebuttal since. It is a live debate. We build on it because the reasoning is good, not because it is proven.

Most of this research is about human teams. Schweiger studied people in rooms. We take composition principles from it. We do not claim it proves that five prompted methods deliberate as well as five executives, and nobody should tell you otherwise.

Self-report has limits we cannot fully engineer away. Short scales trade validity for completion. We use the shortest instruments with the best published record, we measure what we can behaviourally instead, and we keep correcting the profile from real sessions.

The twelve-month claim is a target we measure, not a result we have. "Better after a year of memory" only means something against a held-out set of past decisions, replayed with and without the record. We are building that test. Until it reports, treat the claim as an intention.

Three sources we removed in August 2026. This page used to cite a 1970 study on group size to justify "five and no more": we checked it, and it measures member satisfaction while explicitly reporting that size had negligible effects on performance. It cited a well-known social network study whose causal claims are contested on confounding grounds. And it cited a 2017 magazine article on cognitive diversity, which is editorially reviewed rather than peer reviewed and rests on a commercial assessment tool. None of them were load-bearing for the product. They were load-bearing for the page, which is worse.

A method you can check, including the parts that do not work yet.

Compose my board