What deserves analytical effort — and what does not. The discipline that makes the skeleton useful.
Three cards on what deserves analytical effort — and what does not.
Module 02 built the skeleton: reasoning mode, MECE structure, the conversion from concepts to variables to hypotheses. Module 03 adds the constraint that makes that skeleton useful. Not every legitimate hypothesis deserves to be tested. Not every measurable variable deserves to be measured. Not every real driver deserves to be investigated.
Orient the hypothesis, Select worth testing, Narrow to the dominant drivers
Hypotheses as direction, not bias. Why hypothesis-free analysis hides bias rather than removing it — and the three properties that make a working claim genuinely useful.
What deserves to be tested — and what does not. The decision-linkage discipline that separates worthwhile hypotheses from decorative ones.
Focus over volume. How the three-question test surfaces the few drivers that dominate outcomes — and the discipline of leaving real things unexamined.
The tool legitimacy ladder continues to narrow: not every legitimate hypothesis deserves the inferential machinery.
"Staying neutral" by exploring without a hypothesis does not avoid bias — it hides it.
Without a working claim to focus the investigation, analysts explore widely, examine many things shallowly, and produce descriptions rather than conclusions.
The data tour replaces the data investigation.
The deck grows longer; the findings get vaguer.
And then — because human minds need narrative — a story is constructed at the end to explain what was seen.
The work feels objective because it began without assumptions — but the absence of stated assumptions does not mean assumptions were absent.
What creates bias is protecting a hypotheses from being challenged.
States what relationship or difference is expected, directional and specific.
Vague hypotheses cannot be tested because they do not commit to anything specific enough to be wrong.
Defines what evidence would disprove it.
If no conceivable data could contradict the hypothesis, it is a belief, not an analytical instrument.
The most important property and the one most often missing.
Held conditionally, not absolutely.
Framed as an if then statement that allows revision.
The analyst is committed to investigating it, not defending it.
Before beginning analysis, the analyst should be able to complete this sentence:
"I would conclude this hypothesis is wrong if the data showed _____."
If the blank cannot be filled with a specific, observable finding, the hypothesis is not falsifiable and must be reformulated.
Defining the kill condition explicitly, before looking at data, is the single most reliable defence against unconscious confirmation bias.
Mark if the hypothesis passes each property, choose the maybe kill condition(s).
The hypothesis is on the record. Now: is the test designed to investigate it
| Approach | What it produces | Counts as analysis? |
|---|---|---|
| Designed to confirm | Selective comparisons, narrow controls, specifications chosen because they support the hypothesis. Confirmation is the most likely outcome by construction. | No. Advocacy with statistical decoration. the hypothesis was protected, not investigated. |
| Designed to challenge | Comparisons that could plausibly reject the hypothesis. Controls for the most likely confounders. Specifications persuasive even to a skeptic. | Yes. The hypothesis survives genuine challenge or is revised — either outcome advances understanding. |
| Designed to explore | Multiple specifications, comparisons across many cuts, hunting for patterns. No specific hypothesis at stake. | Only as inductive work. Findings are provisional and must become hypotheses for actual testing. |
Even legitimate hypotheses must be worth testing.
A subtle but costly failure: treating all hypotheses as equally worthy of testing. Once data exists, analysts feel obligated to use it. Once a hypothesis can be tested, it tends to be tested. The result is long lists of statistically "significant" findings that do not meaningfully influence any decision.
Modern infrastructure rewards analytical overanalysis?
The discipline runs against that current. It demands the analyst stop and ask, before any test is run: even if I find something here, will it change anything?
Modern data infrastructure makes hypothesis testing cheap.
Modelling tools make complex tests easy.
Stakeholders often equate the volume of analysis with its rigor.
Together, these forces push analysts toward testing whatever can be tested.
Hypotheses worth testing share three defining characteristics.
The outcome must clearly support one course of action over another. If neither confirmation nor rejection would alter what the business does, the hypothesis is analytically weak.
Articulates not just that a relationship exists, but how it operates and in which direction. Specifies variables, expected magnitude, relevant comparisons, and conditions.
Involves effect sizes that would matter operationally — not merely effect sizes that would be statistically detectable. Practical significance, not just statistical significance.
Four questions that filter hypotheses before any tool is applied
If the analyst cannot complete all four sentences, the hypothesis is not yet ready for testing — and the work to clarify these elements happens before the test, not after.
If the answers to questions 1 and 2 are the same — or if neither would produce action — the hypothesis does not deserve to be tested. Information without consequence is the most expensive kind of information to produce.
Find the Answers. Some are high-stakes; some are decorative;
The most important and least understood gap in applied analytics.
A hypothesis can pass every formal test of significance and still be operationally trivial. The relationship is real, the p-value is small, the confidence interval excludes zero — and yet the magnitude is too small to justify any change in action.
With a sufficiently large dataset, almost any non-zero relationship becomes statistically significant.
A 2-million-subscriber analysis found that app users were 0.4 percentage points more likely to renew. The effect is real, but too small to justify the cost of a promotion campaign.<
High-stakes hypotheses build practical significance into the framing from the start. Rather than asking "is there an effect?" the analyst asks "is the effect at least as large as X?" — where X is the minimum magnitude that would justify the action
Only high-stakes hypotheses justify inferential tools.
| Hypothesis status | Appropriate methods | Why the restriction matters |
|---|---|---|
| Decorative | None. Deprioritise before any tool is applied. | Applying inferential tools to decorative hypotheses produces statistically valid results with no decisional consequence. Looks rigorous, creates no value, consumes attention. |
| Vague but high-stakes | None yet. Must first be made directional and specific. | A vague high-stakes hypothesis cannot be properly tested — there is no clear claim to evaluate. Tools applied prematurely produce findings reshaped to fit whatever the data shows. |
| High-stakes & well-formed | Hypothesis tests, regression, classification, controlled experiments. | Precisely what inferential tools are designed for. The investment is justified because the results will reach a decision. |
Only high-stakes hypotheses justify inferential tools.
Statistical tests must specify, before the test, the threshold at which results would change action. Without this, p-values and confidence intervals become rituals rather than decision tools.
Power comes from focus, not from volume.
This card addresses a common misconception: that more analysis necessarily leads to better decisions. As datasets grow and tools become more powerful, analysts face pressure to analyse everything available. The infrastructure makes broad analysis cheap; stakeholders equate thoroughness with rigor.
Tests whether a driver is statistically dominant.
Filters out:Real but small dirvers. technically correct, but they do not move the answer enough to justify deep investigation.
Tests whether a small change in the driver produces a large change in the outcome.
Filters out:Existing drivers that rarely change in practice; theoretically important, operationally irrelevant.
Tests decision-relevance, not just analytical relevance.
Filters out drivers that are interesting but decision-inert.
The two categories look similar from the outside but produce very different analytical outcomes.
| Category | What it is | Risk of confusion |
|---|---|---|
| Important variables | Factors that dominate system behaviour — they explain a large share of variation, are sensitive to change, or would alter decisions if they behaved differently. | Often fewer than analysts expect. Many systems turn out to be driven by two or three factors rather than the ten or twelve on the initial issue tree. |
| Interesting variables | Factors that draw attention because they are intellectually appealing, easy to analyse, or unfamiliar. They produce striking patterns and stories that feel insightful. | Capture attention without earning it. Their effect on the actual decision is small. Disciplined analysts catch this pattern in themselves. |
The two categories look similar from the outside but produce very different analytical outcomes.
An analyst examining declining retention discovers that subscribers in Tier 3 (4% of revenue) show a striking pattern: their churn correlates with the day of week they activated. The finding is statistically robust and intellectually striking — the analyst spends three weeks on it.
Meanwhile Tier 1 subscribers (62% of revenue) show a steady, ordinary churn rise driven by a competitor's aggressive pricing — visible in the first hour of analysis. The interesting finding got the attention; the important one was never acted on. The work was rigorous in the wrong place.
Pick the 2–3 that form the critical path; the rest stay documented but deprioritised.
Pre-analysis tools, the work that determines where the deep work goes.
| Method | What it does | When to use |
|---|---|---|
| Driver ranking | Order candidate drivers by likely contribution, using existing data, prior knowledge, or rough decomposition. Goal: provisional ranking, not precision. | Many candidate drivers; need a top-2 or top-3 to commit to. |
| Variance decomposition | For numeric outcomes, decompose observed variation into components attributable to each driver. Typically reveals two or three drivers account for most variance. | Historical data exists and drivers can be measured directly. |
| Sensitivity checks (light) | Examine how the answer changes under different assumptions about each driver. Drivers producing large changes are high-priority. | The answer depends on uncertain assumptions. |
| Pareto-style analysis | Cumulative contribution charts that reveal whether a small set of drivers accounts for most of the outcome. | Many candidate drivers; need a quick read on concentration. |
Three questions across Cards G and H. Select the best answer — feedback appears immediately.
Three questions across Cards G and H. Select the best answer — feedback appears immediately.
Three questions across Cards G and H. Select the best answer — feedback appears immediately.
Two questions across Cards H and I — completing the module assessment.
Two questions across Cards H and I — completing the module assessment.
You've worked through the full Focus module — Card G's discipline of hypotheses-as-direction-not-bias, Card H's decision-linkage test, and Card I's critical path. Together these tighten the analytical funnel: from anything testable, to falsifiable and conditional, to decision-linked, to the few drivers that dominate outcomes. Module 04 takes the surviving hypotheses and teaches the modelling disciplines they justify.