The biotech industry runs clinical trials that can cost hundreds of millions of euros, require years of follow-up, and are rejected by regulators if the statistical methodology cannot survive independent scrutiny. The digital startup world launches features based on a single A/B test run over a weekend, adopts management frameworks from bestselling memoirs, and calls it data-driven.
This gap is not merely cultural. It has consequences. And the best conceptual toolkit for closing it does not come from management science — it comes from a debate inside medicine.
Evidence-Based Medicine Was a Revolution. It Was Also Not Enough.
Evidence-Based Medicine (EBM), formalised in the early 1990s, was a genuine intellectual breakthrough. The core idea: clinical decisions should be grounded in the best available research evidence, not in the authority of experts or the traditions of a speciality. Systematic reviews, randomised controlled trials, meta-analyses — these became the gold standard.
The problem is that EBM was designed to evaluate plausible treatments. It was not designed to handle the systematic exploitation of its own methodology by advocates of implausible ones.
When homeopathy, acupuncture, or certain supplements are subjected to EBM-style meta-analysis, they sometimes appear to “work” — not because they do, but because the EBM toolbox has known failure modes that practitioners of pseudoscience have learned to exploit.
Science-Based Medicine (SBM), associated with Steven Novella, David Gorski, and others, adds one critical element: prior probability. How plausible is this intervention before we look at the data? The answer to that question changes how we interpret statistical results — radically.
The Mammogram Problem, and Why It Matters for Management
Here is a number that surprises most people: a diagnostic test with 90% sensitivity and 90% specificity, when used on a population with a 1% baseline prevalence of the disease, will generate roughly nine false positives for every true positive.
Not because the test is bad. Because when the base rate is low, even excellent tests produce a lot of noise.
This is not a mathematical curiosity. It is the correct way to think about any claim that a given management intervention works. Before asking “what does the study say?”, ask: how likely was this to be true before the study ran?
A well-designed RCT showing that a personality test predicts job performance, conducted by the researchers who developed the test and published in a journal with no pre-registration requirement, tells you less than its p-value suggests — because the prior probability that personality tests have strong predictive validity is, based on decades of replication attempts, low.
The Failure Modes EBM Cannot Correct on Its Own
Publication bias and p-hacking
Positive results are published. Negative results are filed. The consequence is that the literature systematically overstates effect sizes. P-hacking — adjusting a study’s methodology until p < 0.05 — is not a rare occurrence. In fields where incentives are misaligned (academic prestige, consulting fees, speaker engagements), it is structural.
The correction is not to distrust all research. It is to weight evidence from pre-registered studies more heavily, to look for meta-analyses that distinguish registered from non-registered trials, and to treat initial effect sizes as upper bounds, not point estimates.
The decline effect
The first study of an intervention consistently reports larger effects than subsequent replications. This is the decline effect: as independent teams with varying methodologies test the same hypothesis, the measured effect size trends asymptotically toward its true value — often much lower than the initial reading, sometimes indistinguishable from zero.
The practical implication is not cynicism about research. It is patience before adoption. An intervention with a single impressive study behind it is a hypothesis, not a policy. An intervention that has been replicated across populations, settings, and independent research teams with consistent (if smaller) effect sizes is closer to something you can act on.
Statistical significance versus clinical significance
Relative risk language conceals absolute magnitudes. “Coffee doubles your risk of disease X” sounds alarming. If the baseline risk is 0.1%, the absolute increase is 0.1 percentage point. Not alarming at all.
Management research makes this error constantly. “Diverse teams outperform homogeneous ones by 35%” (a figure from a widely-cited McKinsey report) is meaningless without knowing: 35% of what? Measured how? Over what time horizon? In what type of organisation? With what effect size in pre-registered replications?
The correct question is never “is the effect statistically significant?” It is “is the effect large enough to matter, and robust enough to replicate?”
Heterogeneity and the limits of meta-analysis
Meta-analyses pool results from multiple studies. When those studies are sufficiently similar in methodology, population, and measurement, pooling is informative. When they are not — when you are averaging apples, oranges, and estimates of how many pieces of fruit the concept of “fruit” contains — the result is statistical noise with a confidence interval.
Much of the management literature suffers from this problem. Studies of “leadership effectiveness” or “organisational culture” measure different constructs with different instruments in different contexts. Aggregating them does not resolve the ambiguity; it launders it.
Applying This to Entrepreneurship
The obvious objection: business decisions cannot wait for the clinical trial infrastructure of pharmaceutical development. True. But the conclusion that follows is not “therefore, anything goes.” It is “therefore, we must be explicit about our epistemic position.”
What this looks like in practice:
State your priors. Before implementing a new HR practice, a new performance management framework, or a new incentive structure, write down what you believe the effect will be and why. This is not about prediction — it is about making assumptions visible so they can be updated.
Distinguish exploration from confirmation. Piloting a new onboarding process in one team is exploratory. Concluding that it works and rolling it out company-wide on the basis of that pilot is treating exploration as confirmation. Run a proper comparison if the decision is consequential.
Look for effect sizes, not testimonials. The biography of a successful founder is not evidence that their management practices caused their success. Selection bias, survivorship bias, and narrative reconstruction make biography a particularly unreliable source of generalisable lessons.
Weight replication. A practice supported by one well-cited study is a conjecture. A practice supported by five independent replications with consistent (if smaller) effect sizes across different organisations is worth serious consideration.
Know the base rates. Most startups fail. Most management interventions have small effects. Most team dynamics problems are structural, not motivational. Keeping these priors in mind does not make you pessimistic — it makes you calibrated.
What This Is Not
Science-based management is not an argument for inaction. The demand for replicated evidence does not mean waiting indefinitely before making decisions. It means being honest about the confidence level attached to each decision.
It is also not an argument against innovation in management practice. New ideas are generated all the time, and occasionally they turn out to be robustly effective. The point of SBM’s methodological rigour is to identify those innovations — not to prevent them from being tested.
The green zone, the grey zone, and the black zone exist in management as in medicine. Some practices have enough evidence behind them to recommend with confidence. Others are genuinely uncertain, and the decision can reasonably depend on context and preference. Others are well-evidenced to be ineffective or harmful — they belong in the bin, regardless of how compelling the TED talk was.
The goal is not to turn scepticism into paralysis. It is to stop mistaking confidence for evidence.
Lionel Arnaud is the founder of Dinnizer SAS, a fractional Deputy CEO practice for deep-tech and high-growth ventures. He has spent a decade applying the methodological standards of drug development to operational management. He can be reached at contact@dinnizer.com.