What data can — and cannot — tell us about cause and effect.
Three cards on what data can — and cannot — tell us about cause and effect.
Most analytical work begins with patterns: things that move together, populations that differ, outcomes that follow events. Translating those patterns into causal claims — claims about what produces what — is where analysis becomes most useful, and most dangerous. The temptation to overstate what the data supports is constant.
The three cards work as a sequence. Card M draws the line between association and causation. Card N replaces single-cause thinking with causal structures. Card O sets the boundary that makes causal claims defensible in the first place.
Why two variables moving together doesn't mean one causes the other — and the three mechanisms that produce false relationships.
Why most problems do not have a single root cause but a causal structure — and how to map it.
Why defining what is outside the analysis is as important as defining what is inside.
Causality is a reasoning problem supported by data not a computational output.
Why "X drives Y" so often outruns the evidence.
Stakeholders ask causal questions ("why is this happening?"), so analysts naturally produce causal-sounding answers.
The pattern in the data feels like an explanation.
Once a story has been constructed, the boundary between association and causation blurs.
The analyst genuinely believes they have found a cause; they have not consciously decided to overclaim.
Determining whether an analysis can support action.
| Concept | What it tells us | What it supports |
|---|---|---|
| Correlation (association) | Two variables move together. | Questions worth investigating. |
| Causation | Changing one variable produces a change in another. | Decisions about what to do. |
Many compelling analytical stories are built on correlations that have no causal meaning.
Acting on such findings can waste resources, create unintended consequences, or mask the real drivers of outcomes.
A third variable Z influences both X and Y, creating a false relationship between them. Z is often invisible in the dataset — acting on the false X–Y link produces no effect on Y.
Y is causing X — the presumed direction of influence is wrong. Especially common in observational data where outcomes are seen after they have already played out.
The data only includes a non-random subset of cases, distorting observed relationships.
The problem is invisible because it lies in what is not in the dataset.
Recognising these three mechanisms is the first defence against causal overclaiming — they are the usual suspects whenever a correlation turns out not to mean what it appeared to.
The foundation of every defensible causal claim.
Counterfactual thinking forces the analyst to imagine the world in which the cause did not happen.
If the outcome would have been the same regardless, the cause is not actually causal.
If it would have been substantially different, the causal claim has support.
Without engaging this question, the analyst is comparing what happened to itself which can never establish causation.
The strongest causal claims rest on the most credible counterfactuals.
| Causal Claim | Implicit Counterfactual | What Would Need to Be True |
|---|---|---|
| The tariff change caused the subscriber-base decline. | Without the tariff change, the base would not have declined as much. | Comparable plans or segments without the tariff change show stable or rising subscribers over the same period. |
| The marketing campaign drove subscriber acquisition. | Without the campaign, fewer subscribers would have been acquired in this period. | Acquisition rates differ from a baseline period or comparable market segment without the campaign. |
| The process change improved provisioning time. | Without the process change, provisioning time would not have improved. | Comparable processes without the change show no improvement over the same period. |
Regression with controls produces conditional associations not causal estimates.
Regression with control variables is sometimes treated as if it produces causal estimates. It does not.
It produces conditional associations — the relationship between X and Y after holding the controls constant — a different and more limited claim.
Regression supports causal reasoning only when combined with strong logic, appropriate comparisons, and well-defined assumptions.
Claiming causation based solely on correlation coefficients, p-values, or model fit.
A significant coefficient on X is consistent with X→Y — and also Y→X, both driven by Z, and selection effects.
The statistical output cannot distinguish among these. Only the analyst's reasoning can.
Regression with controls produces conditional associations not causal estimates.
| Legitimate move | What it does |
|---|---|
| Compare treated and untreated groups | Approximates the counterfactual when groups are comparable on relevant dimensions. |
| Examine timing and sequence | Confirms the supposed cause preceded the effect; rules out reverse causality. |
| Test alternative explanations | Considers confounders, selection effects, and reverse causality before settling on a causal account. |
| State causal assumptions explicitly | Names what would have to be true for the claim to hold, so others can evaluate it. |
Real problems arise from multiple interacting drivers — not one root cause.
A churn spike might trace to a tariff change that made subscribers more sensitive to a network-quality issue that had been tolerable at the old price — three drivers interacting.
A productivity decline might combine a tooling change, a leadership transition, and a seasonal workload shift — with the transition getting the blame because it is most visible.
An underperforming store may suffer from poor location, understaffing, and a misjudged assortment at once. Fixing only one leaves the others intact.
Real problems arise from multiple interacting drivers — not one root cause.
Analysts are pressured to produce one clear explanation, so
→ responsibility can be assigned,
→ an action justified,
→ a narrative constructed.
Single explanations are easier to communicate, act on, and defend. But they are rarely analytically correct.
When the analyst forces a multi-cause reality into a single-cause story, the intervention addresses one piece and produces partial or no improvement.
The analyst was not wrong about the named cause; they were wrong about it being the only cause.
A chain is linear and converges on one root; a structure is networked and shows interacting drivers.
Asking "why" repeatedly can surface candidates — but applied uncritically it embeds the single-cause assumption.
Each "why" leads to one answer, then one further answer.
The technique produces a chain. It does not produce a structure.
Use it as a brainstorming tool, not a method that delivers final answers.
Three roles every driver plays in a causal structure — and what each implies for action.
| Role | What it means | Implication for action |
|---|---|---|
| Necessary cause | A factor that must be present for the outcome to occur. Without it, the outcome would not happen, even if other factors are present. | Removing it prevents the outcome entirely. These are leverage points — addressing them produces decisive change. |
| Sufficient cause | A factor that, on its own, can produce the outcome regardless of other factors. Each sufficient cause is enough. | Multiple sufficient causes may exist. Removing one leaves the others active. Solving the problem requires addressing each — or finding a deeper necessary condition. |
| Contributing factor | A factor that influences the magnitude or likelihood of the outcome but is neither necessary nor sufficient by itself. | Addressing it produces partial improvement. The outcome may still occur, but less severely or less frequently. |
An underperforming retail store has a sufficient cause: the worst foot traffic in the chain due to location. Management invests in marketing. Foot traffic improves modestly; performance does not. Closer investigation reveals additional sufficient causes — an understaffed sales floor and a poorly-curated assortment. Each is independently capable of causing underperformance. The campaign addressed one while the others continued to operate.
Map the structure first, then investigate — the output is a map of the system, not a single root cause.
| Tool | What it does | Best used when |
|---|---|---|
| Logic trees | Hierarchical decomposition of an outcome into the conditions that could produce it. | Surfacing candidates systematically; testing completeness. |
| Fishbone (Ishikawa) | Visual map of cause categories arranged around a central effect. | Causes span functional areas — people, process, technology, environment. |
| Causal mapping | Network diagram showing how drivers connect, reinforce, or counteract. | Interactions matter and need to be represented explicitly. |
Descriptive statistics and targeted analyses test whether proposed drivers are consistent with observed patterns — whether timing aligns, whether magnitudes are plausible. Data does not generate the structure; it tests and refines it.
For each link in the structure, decide the polarity — does the driver reinforce or counteract the next? Wrong picks get a specific explanation.
For each candidate sub-cause, pick the branch it belongs under — or mark it a decoy that doesn't fit the decomposition.
Organise each candidate cause by category. Decoys — causes with no plausible pathway — get marked as such.
Where analysis ends — and why that matters.
This card addresses a counterintuitive but essential skill: setting boundaries. Expanding scope is often seen as thoroughness — "consider everything," "look at the bigger picture." The cultural pull is toward inclusion. But analytical quality depends on a clear, defensible scope. Every analysis operates within a system boundary — factors inside the analysis, and factors treated as fixed background.
The impulse is strongest when the analyst hits a finding they can't explain. Rather than acknowledging the limits of the current analysis, the natural response is to broaden scope — pull in more variables, more periods, more context. Each addition seems reasonable in isolation. But the cumulative effect is that the analysis no longer addresses a clearly bounded system.
Professional analysts do not attempt to explain the entire world; they explain a bounded system well. The boundary is set consciously, based on the decision being supported and the time scale of the analysis.
Defining what is outside the analysis is just as important as defining what is inside it.
Factors inside are investigated; factors outside are acknowledged but treated as fixed background. The skill is not narrowness for its own sake — it is scope appropriate to the question being answered. The discipline of explicit exclusion protects analytical clarity; resisting boundary expansion under pressure protects credibility.
Each function is reason enough to set boundaries deliberately; together they make this one of the most leveraged decisions on a project.
Determines what can reasonably be assumed to remain stable during the analysis period. Done badly — boundaries too wide — external factors that should be background are themselves changing, distorting the analysis.
Controls which factors are candidates for causal explanation. Done badly — too many external factors included — causal pathways blur and the analysis resolves nothing.
Keeps explanations specific enough to act on. Done badly — "it's a complex interaction of many factors" — true but useless; recommendations dissolve into general advice.
A regional sales team's performance has declined. For each candidate factor, decide: inside scope (investigated) or outside (fixed background)?
The discipline that protects analytical credibility under stakeholder pressure.
An analyst is investigating why a regional sales team's performance declined. Defined scope: the team's pipeline metrics, rep activity, and lead-quality data over two quarters. Halfway through, a stakeholder asks: "What about the broader economy? Competitor pricing? Our regional marketing spend?" Each is reasonable — but each lies outside the deliberately-set boundary.
The disciplined response: "Good questions. They lie outside the current scope, which was set to focus on the team's pipeline. If our findings point to a need for that broader investigation, that's a follow-up analytical episode." The boundary holds; the current analysis produces a clear, decisional conclusion.
Quietly enlarging the scope so a new factor can be included. The conclusions no longer apply to a stable scope, and stakeholders cannot tell what the analysis actually examined.
Carefully bounding the analysis but then presenting findings as if they applied to the full unbounded system. "Within the scope of this analysis…" is not a hedge; it is accurate.
Three questions across Cards M and N. Select the best answer — feedback appears immediately.
Two questions across Cards O and M — completing the module assessment.
You've worked through the full Depth module — the discipline of separating association from causation (M), the move from single causes to causal structures (N), and the deliberate setting of system boundaries (O). These cards form the foundation for everything in Module 06 onward.