What we can show you, and what we cannot.
Five Peers makes claims about how decisions get better. Below is the evidence each one rests on, with the paper, and a section at the end for the claims that are not settled. We removed three sources this site used to cite, because we checked them and they did not say what we were using them to say.
Checked at source, August 2026.
Each rule the product follows, and where it comes from.
Argument beats agreement, and it costs you something
Two groups given the same strategic problem: one told to reach consensus, one told to argue it out through devil's advocacy or dialectical inquiry. The arguing groups produced better recommendations and surfaced better assumptions. They also came out of it less satisfied, less attached to the decision, and less keen to work together again.
This is why the argument runs inside the session and not at you. The quality gain is real and so is the cost, so Five Peers keeps the disagreement between the methods and hands you the output, at a bluntness you set yourself.
Schweiger, Sandberg & Ragan, Academy of Management Journal, 1986. Meta-analysis: Schwenk, Organizational Behavior and Human Decision Processes, 1990. Source
What improves judgement is method, not personality
In the IARPA forecasting tournament, about an hour of training in probabilistic reasoning and scenario work improved accuracy by roughly 10% on Brier scores, and the effect held across four years. The training was concrete: state coherent probabilities, challenge your assumptions, name the causal drivers, build the best and worst case, avoid over-correcting.
That is the list the five methods are built from. It is also why Five Peers does not sell you a simulated Steve Jobs: no study shows that borrowing a famous person's voice improves a decision.
Mellers, Tetlock et al., Good Judgment Project. Training effects reported in Judgment and Decision Making and related tournament papers. Source
You cannot self-report your own blind spots
People readily see bias in others and deny it in themselves, and they keep denying it when shown the evidence that they have just exhibited it. They are not performing: they genuinely believe their own read is the unbiased one. The finding has since been replicated in a preregistered study.
So the onboarding never asks which biases you have. It watches which trade-offs you actually make, on a decision that is real, and infers from there.
Pronin, Lin & Ross, Personality and Social Psychology Bulletin, 2002. Preregistered replication in Collabra: Psychology, 2024. Source
Assume it already failed and you find more of the reasons
Asking people to explain a future event as though it had already happened, rather than to assess whether it might, increased correct identification of reasons by around 30%, and produced more concrete, action-shaped causes rather than abstract ones.
That is the pre-mortem method, one of the twelve. It is in the library because of this result, not because it makes a good slide.
Mitchell, Russo & Pennington, Journal of Behavioral Decision Making, 1989. Popularised as the pre-mortem by Klein, Harvard Business Review, 2007. Source
The one disposition worth measuring
Actively open-minded thinking, the habit of going looking for the argument against your own position, is measurable with a short scale and predicts real belief updating and forecasting accuracy. It is the best evidence-to-length ratio in the whole field.
It is the one dimension the onboarding measures head-on, because it decides how hard your board should push before you stop listening.
Baron; Stanovich & West. Short form validated by Haran, Ritov & Mellers, Judgment and Decision Making, 2013. Source
There is no such thing as your risk appetite
Risk taking is domain-specific. The same person is cautious with money and reckless with their reputation, or the reverse. A single risk score collapses five uncorrelated things into one number and tells you nothing useful.
So there is no risk slider in the onboarding. Risk is asked inside the domain of the decision you actually brought.
Blais & Weber, DOSPERT scale, Judgment and Decision Making, 2006. Source
Twelve reasoning methods. You get the five you are missing.
Every method is a question asked relentlessly and a thing it refuses to let pass. None of them is a person, a job title, or a personality type. Which five you get is decided by the onboarding, and it changes as the sessions accumulate evidence about how you actually decide.
Goes looking for the evidence that would kill the plan, before the market finds it for you.
Puts a number on the claim, then a confidence interval around the number, so the argument stops being a feeling.
Runs the decision two and three moves out, where the cost usually lives.
Treats the failure as already true and works backwards, which surfaces causes a risk review does not.
Separates the decisions worth agonising over from the ones worth making today and correcting later.
Checks that the decision survives contact with the people and the calendar you actually have.
Maps the decision onto the people it moves, including the ones who will never say no out loud.
Prices the option of not deciding yet, which is a decision that rarely gets costed.
Refuses the incremental version of the plan and asks what you believe that your market does not.
Finds the closest real case rather than the most flattering analogy.
Names the customers, staff and third parties who pay for the decision without appearing in the model.
Forces the decision into a form your team can execute without a translation layer.
How five get picked, and how you can check it is not arbitrary.
A profile vector, not a personality type
The onboarding produces a small set of scores: how much you go looking for the counter-argument, how you trade speed against information, how you handle risk inside the specific domain of your decision. No type, no letter code, no colour.
Measured on the 64 profiles the journey can produce: 12 of them share no seat at all with their exact opposite, and the rest still differ by at least one.
A coverage matrix
Each of the twelve methods covers a specific failure mode. Falsification covers running from a hunch to a commitment. Base rates covers arguing from one vivid example. Cost of delay covers avoidance dressed up as prudence. The matrix is written down and versioned, not inferred by a model on the day.
Deficits first, then a diversity floor
We score where your coverage is thinnest, rank the methods that close those gaps, and impose a minimum spread so you never get five variations on the same move. Same rules for everyone, different output for everyone.
Tests we hold ourselves to
Retake the onboarding and the board should be broadly the same. Nudge one score slightly and the board must not flip entirely. And a personalised board has to beat a generic one on the same decisions, or the personalisation is decoration.
Why the record matters more than the methods.
Five methods can be reproduced by anyone with an afternoon and a decent model. What cannot be reproduced is your own decision history with the assumptions attached: what you were betting on, whether the bet held, the pattern you keep repeating. Each session writes a structured entry, the entry gets a review date, and the review asks whether the assumption survived contact with reality.
That record is stored and structured by HOLCO, hosted in France. Five Peers orchestrates and reads it at the moment of deciding. It exports in an open format, on request, without asking us. A record you cannot take with you is not yours.
The honest limits.
A method page that only lists supporting evidence is marketing. Here is what is unresolved, and what we removed.
Cognitive complementarity is our product hypothesis, not a settled theorem. The best known formal argument that diverse groups beat able ones (Hong and Page) drew a peer-reviewed mathematical rebuttal in the Notices of the American Mathematical Society in 2014, and rebuttals of the rebuttal since. It is a live debate. We build on it because the reasoning is good, not because it is proven.
Most of this research is about human teams. Schweiger studied people in rooms. We take composition principles from it. We do not claim it proves that five prompted methods deliberate as well as five executives, and nobody should tell you otherwise.
Self-report has limits we cannot fully engineer away. Short scales trade validity for completion. We use the shortest instruments with the best published record, we measure what we can behaviourally instead, and we keep correcting the profile from real sessions.
The twelve-month claim is a target we measure, not a result we have. "Better after a year of memory" only means something against a held-out set of past decisions, replayed with and without the record. We are building that test. Until it reports, treat the claim as an intention.
Three sources we removed in August 2026. This page used to cite a 1970 study on group size to justify "five and no more": we checked it, and it measures member satisfaction while explicitly reporting that size had negligible effects on performance. It cited a well-known social network study whose causal claims are contested on confounding grounds. And it cited a 2017 magazine article on cognitive diversity, which is editorially reviewed rather than peer reviewed and rests on a commercial assessment tool. None of them were load-bearing for the product. They were load-bearing for the page, which is worse.
A method you can check, including the parts that do not work yet.
Compose my board