Skip to content

The Texas Sharpshooter Fallacy: How You Paint a Target Around Your Own Bullet Holes

C. Pearson C. Pearson
/ / 4 min read

Picture a man who fires fifty rounds at a barn wall, walks up to where the bullets clustered, and draws a bullseye around the tightest group. Then he tells everyone he's a sharpshooter.

Colorful vintage cars covered in graffiti art at the famous Cadillac Ranch, Texas. Photo by Allie on Pexels.

That's the fallacy. And it's running wild in data science.

The Texas Sharpshooter Fallacy happens when you let the data define the hypothesis instead of the other way around. You look at results, find a cluster that looks interesting, and then treat that cluster as if you'd predicted it in advance. The target gets drawn after the shot. The pattern feels real because it's sitting right there in front of you, but you manufactured it by choosing which data to highlight after seeing all of it.

This isn't a fringe problem. It's baked into how exploratory analysis gets misreported, how medical studies get published, and how business analysts build narratives that sound rigorous but aren't.

The Mechanism Hiding in Plain Sight

Here's the core problem: random data always contains clusters. Always. Given enough observations, any sufficiently flexible pattern-finder will find something that looks non-random.

Suppose you track 50 health outcomes across 10 towns. With 500 data points and no real signal, you will almost certainly find a town-outcome pair that looks alarming just from random variation. If you then write the paper about that specific pairing, you've committed the fallacy. You've fired into the barn first.

The math on this is unforgiving. If you run 20 independent tests with a p-value threshold of 0.05, you expect one false positive just from chance. Do 100 tests and you expect five false positives. When the hypothesis is selected because it passed, the p-value means nothing. The significance threshold was designed for pre-specified tests. Applied after data-dredging, it's decoration.

Why It's So Hard to Catch

Part of what makes this fallacy pernicious is that the person doing it rarely knows they're doing it. Exploratory data analysis is legitimate and valuable. You should look at your data. The problem arrives when exploration slides silently into confirmation without anyone marking the transition.

A researcher looks at cancer rates across 200 counties. One county in the sample has a rate three standard deviations above the mean. They investigate: there's a chemical plant nearby. They write up the finding. The story is compelling. The statistics look clean. But they never asked: how many counties would we expect to be three standard deviations out just by chance in a sample this size? The answer, roughly, is one.

The bullseye was drawn around that one county because the bullet landed there.

A Diagram of the Trap

graph TD
    A[Collect Data] --> B[Observe All Results]
    B --> C{Find Interesting Cluster}
    C --> D[Define Hypothesis Around Cluster]
    D --> E[Run Statistical Test on Same Data]
    E --> F[Report Significant Finding]
    F --> G((False Discovery))

Notice that hypothesis definition happens after observation. That's the trap. Confirmatory analysis requires the arrow between hypothesis and data collection to go in the other direction.

Real Consequences, Not Just Academic Ones

This fallacy has a body count. A string of published studies in the early 2000s linked specific genes to depression, schizophrenia, and other complex conditions. Sample sizes were small, hypotheses were often post-hoc, and clusters were identified from the data itself. When larger pre-registered replication studies ran, most findings evaporated. The Candidate Gene Era, as it's now called, produced thousands of papers and very few durable results.

Business intelligence has the same problem. An analyst slices a revenue dataset by region, product, time period, customer segment, and channel. Some combination will show a remarkable Q3 spike. Present that spike to leadership without disclosing how many slices were examined, and you've just drawn a target around a bullet hole.

How to Actually Protect Yourself

Separate your data into exploration and confirmation sets before you start looking. Whatever you find in the exploration phase gets tested on held-out data with a pre-specified hypothesis. This isn't just a best practice; it's the only way the statistics remain interpretable.

Pre-registration takes this further. You write down your hypothesis, your analysis plan, and your sample size before collecting data. Then you run exactly that analysis. Pre-registered studies replicate at dramatically higher rates than unregistered ones.

When you can't pre-register (legacy data, observational studies), be explicit about what you're doing. Label exploratory findings as hypothesis-generating. Don't report p-values from data-dredged results as if they mean what p-values mean in confirmatory tests. They don't.

And when you see a striking pattern in someone else's analysis, ask the question they didn't: how many other patterns were examined before this one was selected? If they can't answer that, the bullseye might be painted on.

Get Mean Methods in your inbox

New posts delivered directly. No spam.

No spam. Unsubscribe anytime.

Related Reading