index Introduction The Mindset The Skeleton The Focus The Rigor The Depth The Output
Module 05 · The Depth

Causal Thinking

Cards M · N · O
Module 05 · The Depth

The Discipline of Causal Thinking

Three cards on what data can — and cannot — tell us about cause and effect.

Most analytical work begins with patterns: things that move together, populations that differ, outcomes that follow events. Translating those patterns into causal claims — claims about what produces what — is where analysis becomes most useful, and most dangerous. The temptation to overstate what the data supports is constant. Module 05 builds the discipline that resists it.

"Correlation can suggest questions; it cannot answer causal ones."

The three cards in this module work as a sequence. Card M draws the line between association and causation, and the mechanisms that blur it. Card N replaces single-cause thinking with causal structures. Card O sets the boundary that makes causal claims defensible in the first place.

M

Correlation vs. Causation

Why two variables moving together doesn't mean one causes the other — and the three mechanisms that produce false relationships.

N

Root Cause Analysis

Why most problems do not have a single root cause but a causal structure — and how to map it.

O

The System Boundary

Why defining what is outside the analysis is as important as defining what is inside.

How to use this module

Each card builds on the previous. The cards reference each other deliberately: the mechanisms in Card M become components of the structures in Card N, and both depend on the boundary judgment introduced in Card O. Work through them in order on first pass; revisit any section later as a reference.

Card M · Module 05

The Causal Overclaim

Why "X drives Y" so often outruns the evidence.

This card addresses one of the most persistent and dangerous analytical misconceptions: the assumption that observed relationships in data automatically imply cause and effect. The misconception is so widespread, and so reinforced by everyday language, that even experienced analysts slip into it without noticing. "X drives Y." "This causes that." "These customers churn because of price." The phrasing is causal, but the evidence is correlational. The gap between what was observed and what is being claimed is the gap this card is designed to close.

In many analytical contexts, correlations are treated as explanations. Statistical significance is mistaken for causal truth. Two variables move together, the relationship is robust, the model fits the data — and the conclusion is presented as causal. This error leads to decisions that are confident, precise, and wrong. The recommended action is taken. The expected effect does not materialise. The analyst is surprised, but they should not be: the data never supported the causal claim that drove the decision.

Why this happens

Stakeholders ask causal questions ("why is this happening?" "what's driving this?"), so analysts naturally produce causal-sounding answers. The pattern in the data feels like an explanation. Once a story has been constructed, the boundary between association and causation blurs. The analyst genuinely believes they have found a cause; they have not consciously decided to overclaim.

Association vs. causation — the working definitions

ConceptWhat it tells usWhat it supports
Correlation (association) Two variables move together. Questions worth investigating.
Causation Changing one variable produces a change in another. Decisions about what to do.

The difference is not semantic; it determines whether an analysis can support action. Many compelling analytical stories are built on correlations that have no causal meaning. Acting on such findings can waste resources, create unintended consequences, or mask the real drivers of outcomes. The remedy is not to avoid causal reasoning — business decisions require it — but to reason about causation responsibly: identifying the mechanisms that could be at work, considering alternatives, and stating conclusions with calibrated confidence.

Key Takeaway
Correlation can suggest questions; it cannot answer causal ones.
Card M · Module 05

Three Mechanisms That Produce False Correlation

Each is common in business data, each is easy to miss, and each produces decisions that fail when acted on.

Recognising the three mechanisms below is the first defence against causal overclaiming. They are the usual suspects whenever a correlation turns out not to mean what it appeared to mean.

Mechanism 01

Confounding

A third variable Z influences both X and Y, creating a false relationship between them.

Mechanism 02

Reverse Causality

Y is causing X — the presumed direction of influence is wrong.

Mechanism 03

Selection Effects

The data only includes a non-random subset of cases, distorting the observed relationship.

Confounding — apparent X→Y relationship. Click "Reveal the Hidden Structure" to see what's actually producing it.

1 · Confounding Variables

Confounding variables influence both the supposed cause and the supposed effect, creating a false relationship between them. The analyst observes that X and Y move together and concludes that X causes Y. In reality, a third variable Z is influencing both X and Y, producing the apparent relationship without any direct causal link.

Confounders are particularly dangerous because they are often invisible in the data being examined. The analyst sees the strong correlation between X and Y; they do not see Z, which lives outside the dataset or is not being controlled for. Acting on the false X–Y relationship produces no effect on Y, because changing X does not change the underlying driver. The decision was made on a real correlation that turned out not to be causal.

Business Example
The training program that didn't work

An HR analyst observes that employees who completed a leadership-training program were promoted 40% more frequently in the following year than those who did not. The conclusion seems obvious: the training drives promotion. The company expands the program.

A year later, no increase in promotion rates is observed. The actual driver was confounded: high-performing employees were both more likely to be selected for training and more likely to be promoted regardless. The training did not cause the promotions; underlying performance caused both. The original correlation was real but not causal, and the expanded program produced no effect because the actual mechanism was untouched.

2 · Reverse Causality

Reverse causality occurs when the presumed direction of influence is wrong — what appears to be a cause is actually an effect. The analyst observes that X and Y are related and concludes that X causes Y. In reality, Y is causing X, or both directions are operating simultaneously.

Reverse causality is especially common in observational business data, where the analyst sees outcomes after they have already played out. Without temporal precedence or experimental control, it is often unclear which variable preceded which. The story that the analyst constructs depends on which variable they decided to treat as the cause.

Business Example
The customer-success calls that didn't help

A SaaS company observes that customers who receive frequent customer-success calls have higher renewal rates than customers who do not. The conclusion: customer-success calls drive renewal. The team plans to expand the program.

Closer investigation reveals reverse causality: customer-success managers preferentially call customers who are showing signs of healthy engagement — regular logins, expanded usage, recent feature adoption. The healthier customers receive more calls because they are healthier; they were going to renew anyway. The relationship between calls and renewal is real, but the direction is opposite to what was assumed. Calling at-risk customers more frequently — the implied recommendation — may not produce the renewal lift the company expected.

3 · Selection Effects

Selection effects arise when the data only includes a non-random subset of cases, distorting observed relationships. The analyst sees patterns that hold within the visible data but do not hold in the broader population. Acting on these patterns produces results that do not match expectations, because the population being acted on is different from the sample that was studied.

Selection effects often look like robust findings within the available data. Every check the analyst can run inside the dataset confirms the relationship. The problem is invisible because it lies in what is not in the dataset — the cases that were filtered out before analysis began.

Business Example
The pricing experiment with hidden selection

A pricing analyst studies historical data to estimate how customers respond to discounts. The analysis shows a strong relationship: customers who received discounts had much higher purchase rates than those who did not. The recommendation is to expand discounting.

The expansion produces a fraction of the expected lift. The selection effect was hidden in plain sight: discounts had historically been offered selectively to customers who were already showing strong purchase signals, while customers showing weak signals received no discount. The data contained the wrong comparison — discounted-and-likely-to-buy versus undiscounted-and-unlikely-to-buy. When discounts were extended to the broader population, the original relationship did not generalise.

Card M · Module 05

Counterfactual Thinking

The foundation of every defensible causal claim.

Beyond recognising the mechanisms that produce false correlations, advanced analysts must develop a specific habit of thought: counterfactual reasoning. The counterfactual question is the foundation of all causal claims:

What would have happened if the supposed cause had not occurred?

Counterfactual thinking forces the analyst to imagine the world in which the cause did not happen and to reason about what the outcome would have been. If the outcome would have been the same regardless, then the cause is not actually causal — something else was driving the outcome. If the outcome would have been substantially different, then the causal claim has support. Without engaging this question, the analyst is comparing what happened to itself, which can never establish causation.

Three claims, three counterfactuals

Causal ClaimImplicit CounterfactualWhat Would Need to Be True
Pricing change caused the sales decline. Without the pricing change, sales would not have declined as much. Comparable products or segments without the pricing change show stable or rising sales over the same period.
Marketing campaign drove customer acquisition. Without the campaign, fewer customers would have been acquired in this period. Acquisition rates differ from a baseline period or comparable market segment without the campaign.
Process change improved cycle time. Without the process change, cycle time would not have improved. Comparable processes without the change show no improvement over the same period.
A note on certainty

Counterfactual reasoning rarely produces certainty. The analyst cannot literally observe what would have happened in the alternate world; they can only construct comparisons that approximate it — untreated groups, prior periods, comparable segments, similar markets. The quality of the causal claim depends on the quality of these comparisons. The strongest causal claims rest on the most credible counterfactuals.

What regression can — and cannot — do

At this stage, statistical tools such as regression models can help control for observable factors, but they cannot, by themselves, establish causality. This is one of the most important and least understood points in modern analytics. Regression with control variables is sometimes treated as if it produces causal estimates. It does not. It produces conditional associations — the relationship between X and Y after holding the controls constant — which is a different and more limited claim.

Regression supports causal reasoning only when combined with strong logic, appropriate comparisons, and well-defined assumptions. The math alone is insufficient. The analyst must articulate why the comparison being used (treated vs. untreated, before vs. after, this segment vs. that segment) approximates the counterfactual of interest, and why the controls included in the model address the most likely confounders. Without this articulation, the regression coefficient is just a number — real, statistically significant, and not necessarily causal.

Legitimate moves for strengthening causal claims

MoveWhat it does
Compare treated and untreated groupsApproximates the counterfactual when groups are comparable on relevant dimensions.
Examine timing and sequenceConfirms the supposed cause preceded the effect; rules out reverse causality.
Test alternative explanationsConsiders confounders, selection effects, and reverse causality before settling on a causal account.
State causal assumptions explicitlyNames what would have to be true for the claim to hold, so others can evaluate it.
What is not legitimate

Claiming causation based solely on correlation coefficients, p-values, or model fit statistics. A regression with a significant coefficient on X is consistent with X causing Y — and also with Y causing X, with both being driven by Z, and with selection effects producing a relationship that does not exist outside the sample. The statistical output cannot distinguish among these. Only the analyst's reasoning can.

Causality is a reasoning problem supported by data, not a computational output.
Card N · Module 05

From Single Cause to Causal Structure

Most real problems arise from multiple interacting drivers — not one root cause.

This card addresses a persistent simplification error in analytical work: the tendency to search for a single root cause. In many organisational settings, analysts are pressured to produce one clear explanation for a problem — usually so that responsibility can be assigned, a specific action can be justified, or a leadership narrative can be constructed. The desire for a single answer is understandable. Single explanations are easier to communicate, easier to act on, and easier to defend. But they are rarely analytically correct.

Advanced analytical thinking recognises that most real-world problems arise from multiple interacting drivers rather than a single cause. Performance issues, failures, and inefficiencies are typically the result of combinations of decisions, behaviours, constraints, and feedback loops operating together over time.

Illustration
Three problems, three structures

A churn spike might trace to a pricing change that made customers more sensitive to a service-quality issue that had been tolerable at the old price — three drivers interacting, none of which would have produced the observed outcome alone.

A productivity decline might combine a tooling change, a leadership transition, and a seasonal shift in workload — with the leadership transition getting the blame because it is the most visible.

An underperforming store may suffer simultaneously from poor location, understaffing, and a misjudged assortment. Fixing only one leaves the others intact.

Single-cause thinking maps poorly onto these situations. When the analyst forces a multi-cause reality into a single-cause story, several failures follow. The proposed intervention addresses one piece of the structure and produces partial or no improvement. Other contributors continue operating, sometimes amplifying the original problem. Stakeholders lose confidence in analysis when the recommended action does not produce the predicted result. The analyst was not wrong about the named cause; they were wrong about it being the only cause.

Chains vs. structures

A causal structure is fundamentally different from a causal chain. A chain is linear: A causes B, which causes C. The analyst follows the chain backward until reaching what feels like a root, then stops. A structure is networked: multiple drivers interact, some reinforcing each other, some counteracting, some operating at different time scales. The analyst maps the structure rather than walking the chain.

CAUSAL CHAIN One root — found by walking back Outcome Cause B Cause A Root cause CAUSAL STRUCTURE Multiple drivers, interacting Outcome Driver 1 Driver 2 Driver 3 Interaction Interaction reinforces
Linear chain converges to one answer; structure shows interacting drivers and reinforcing loops.
The limits of "5 Whys" used uncritically

The "5 Whys" technique — asking "why" repeatedly to drill down into causes — can be useful as a starting point, but applied uncritically it embeds the single-cause assumption into the analysis. Each "why" leads to one answer, then to one further answer, and so on. The technique produces a chain. It does not produce a structure. The analyst arrives at "the" root cause precisely because the method was designed to converge on one. Use it as a brainstorming tool to surface candidates — not as a method that delivers final answers.

A more powerful framing is to map the structure first and then investigate. The analyst lists the candidate drivers, considers which interact and how, then uses data to test the proposed structure. The output is not a single root cause but a map of the system that produced the outcome — with the dominant components identified and the interactions made explicit.

Tools for mapping structures

Click any row below to jump to the interactive practice for that tool.

ToolWhat it doesBest used when
Logic treesHierarchical decomposition of an outcome into the conditions that could produce it.Surfacing candidates systematically; testing completeness.
Fishbone (Ishikawa)Visual map of cause categories arranged around a central effect.Causes span functional areas — people, process, technology, environment.
Causal mappingNetwork diagram showing how drivers connect, reinforce, or counteract.Interactions matter and need to be represented explicitly.

Data plays a supporting role at this stage. Descriptive statistics, comparisons, and targeted analyses test whether proposed drivers are consistent with observed patterns. The analyst examines whether timing aligns, whether magnitudes are plausible, whether the structure explains both the central outcome and any related observations. Data does not generate the structure; it tests and refines it.

Key Takeaway
Most problems do not have one root cause; they have a causal structure.
Card N · Module 05

Necessary, Sufficient, Contributing

Three roles every driver plays in a causal structure — and what each implies for action.

Within causal structures, three roles deserve careful distinction. Confusing them produces interventions that look reasonable but fail in practice.

RoleWhat it MeansImplication for Action
Necessary cause A factor that must be present for the outcome to occur. Without it, the outcome would not happen, even if other factors are present. Removing it prevents the outcome entirely. These are leverage points — addressing them produces decisive change.
Sufficient cause A factor that, on its own, can produce the outcome regardless of other factors. Each sufficient cause is enough. Multiple sufficient causes may exist for the same outcome. Removing one leaves the others active. Solving the problem requires addressing each — or finding a deeper necessary condition.
Contributing factor A factor that influences the magnitude or likelihood of the outcome but is neither necessary nor sufficient by itself. Addressing it produces partial improvement. The outcome may still occur, but less severely or less frequently.

Most business problems involve a mix of these. The analyst's job is not to find "the" cause but to identify what role each contributor plays in the structure. A necessary cause that has been overlooked is the most leveraged finding — removing it solves the problem cleanly. Multiple sufficient causes require more careful intervention design, since addressing only one may leave the problem largely untouched. Contributing factors merit attention proportional to their estimated effect.

Business Example
Why one fix didn't work

A retailer notices that a particular store is consistently underperforming. An analyst identifies a sufficient cause: the store has the worst foot traffic in the chain due to its location. Management invests in a marketing campaign to bring more people through the door. Foot traffic improves modestly. Performance does not.

Closer investigation reveals additional sufficient causes: an understaffed sales floor and a poorly-curated assortment for the local demographic. Each is independently capable of causing underperformance. The marketing campaign addressed one sufficient cause while the others continued to operate. The analyst's original finding was correct — foot traffic was a real driver — but framing it as the cause led to an intervention that touched only one piece of a multi-driver structure.

What is — and isn't — legitimate at this stage

Legitimate

Mapping a causal structure with logic trees, Fishbone diagrams, or causal maps; classifying drivers by role; using data to test whether the proposed structure is consistent with what is observed; refining the structure when the data does not support it.

Not Legitimate

Declaring a "root cause" based solely on intuition, anecdotal evidence, or a single piece of data. Root cause analysis must remain grounded in both logic (the proposed structure must be coherent and complete) and data (the structure must be consistent with what is observed). Logic without data produces plausible-sounding stories that may not match reality. Data without logic produces patterns that have not been organised into a coherent explanation.

The most leveraged analytical finding is rarely "the" cause; it is the map of the system that produced the outcome.
Card N · Module 05

Mapping Practice

Three real-world scenarios for practising the three tools — starting with Causal Mapping.

The tools introduced on the Structure tab are most useful in your hands, not in your head. This is where you practise them. Each scenario below is a real-shaped business problem; the act of building the structure forces explicit choices about direction and polarity that loose prose can blur.

What's on this tab

Three practice widgets, one per tool from the Structure tab — Causal Map, Logic Tree, and Fishbone. Work through them in any order; the tools are complementary, not sequential. Each scenario is a real-shaped business problem where the act of building the structure forces choices that loose prose can blur.

Causal Map

The richest of the three tools — and the one that most directly demonstrates Card N's central claim that structures contain interactions and loops, not just chains. Pick a scenario, then build the map. Drag from one node to another to draw a causal arrow; toggle the polarity below for the next link you'll draw. Click Check Structure when you're done — wrong edges get a specific explanation, and the reinforcing loop turns gold when closed.

Next arrow:
Drag from any node to another to draw an arrow. Current polarity: + (reinforces).
How to read the feedback

Polarity wrong — the right edge in the wrong direction (+ vs −). The status panel will explain why that polarity is what it is. Arrow reversed — the relationship runs the other way; the explanation tells you what the correct direction means. Off-structure — the link doesn't appear in the map; the two nodes may both be downstream of other factors but don't directly cause each other.

Logic Tree

Logic trees decompose an outcome into the conditions that could produce it — surfacing candidates systematically while testing for completeness and overlap. The decomposition framework chosen at level 1 shapes everything downstream; the discipline is to commit to one frame and exhaust it before adding others.

For each scenario the level-1 framework is fixed. Drag each candidate sub-cause to the branch it belongs under. Some candidates are decoys — leave them in the pool if you think they don't fit the decomposition.

Drag each candidate to the branch it belongs under. Leave decoys in the pool.

Fishbone (Ishikawa)

Fishbone diagrams organise causes by category — most useful when causes span multiple functional areas. The visual structure forces the analyst to consider each area in turn, not just the one closest to hand. The classic manufacturing version uses six Ms; service and sales problems use whatever categorical decomposition fits.

Drag each candidate cause to the category bone it belongs to. Decoys belong in the pool.

Drag each cause to the category it belongs to. Decoys stay in the pool.
Card O · Module 05

The System Boundary Question

Where analysis ends — and why that matters.

This card addresses a counterintuitive but essential analytical skill: setting boundaries. In many analytical environments, expanding the scope of analysis is seen as a sign of thoroughness. Analysts are encouraged to "consider everything," "add more variables," or "look at the bigger picture." Stakeholders frequently reward broader analyses with more attention. The cultural pull is toward inclusion: when in doubt, consider one more factor.

This instinct is well intentioned but often counterproductive. Analytical quality depends on a clear, defensible scope. Every analysis operates within a system boundary — a defined region of factors that are inside the analysis and a complementary region of factors that are treated as fixed background. When boundaries are clear, the analysis can be coherent, the assumptions can be explicit, and the conclusions can be defended. When boundaries are unclear or expand mid-stream, the analysis becomes muddled, assumptions multiply silently, and the conclusions lose their grounding.

The expansion impulse

The expansion impulse is particularly strong when the analyst encounters a finding they are not sure how to explain. Rather than acknowledging the limits of the current analysis, the natural response is to broaden the scope: pull in more variables, examine more periods, consider more contextual factors. The analysis grows. Each addition seems reasonable in isolation. But the cumulative effect is that the analysis no longer addresses a clearly bounded system. It addresses a sprawling, ill-defined region in which causal pathways blur and conclusions weaken.

Advanced analytical thinking treats system boundaries as a deliberate design choice, not a default. Professional analysts do not attempt to explain the entire world; they explain a bounded system well. The boundary is set consciously based on the decision being supported, the time scale of the analysis, and the factors that can plausibly be held constant during the period being studied. Within that boundary, analysis is rigorous. Outside it, factors are treated as fixed background — acknowledged but not investigated.

THE SYSTEM BOUNDARY macro economy competitor moves regulatory change long-term trends held as fixed background INSIDE SCOPE Pipeline data Rep activity Lead quality Decision being supported
A boundary makes scope explicit. Factors inside are investigated; factors outside are acknowledged but treated as fixed background.

The purpose of this card is to teach that defining what is outside the analysis is just as important as defining what is inside it. The discipline of explicit exclusion protects analytical clarity. The discipline of resisting boundary expansion, especially under pressure, protects analytical credibility. The skill is not about narrowness for its own sake; it is about scope appropriate to the question being answered.

Card O · Module 05

Three Functions of a System Boundary

Each function is reason enough to set boundaries deliberately; together, they make this one of the most leveraged decisions on a project.

System boundaries serve three essential functions in analytical work. They reinforce each other: a boundary that produces stability also produces clearer causal attribution, and clearer attribution produces more specific recommendations. Conversely, when boundaries fail one function, they typically fail the others too.

Function 01

Stability

Determines what can reasonably be assumed to remain stable during the period of analysis. Factors outside the boundary are treated as fixed background conditions.

Done well
Background factors genuinely do hold steady over the analysis period, so the bounded analysis is interpretable.
Done badly
Boundaries too wide; the assumption of stability fails. External factors that should have been background are themselves changing, distorting the analysis.
Function 02

Causal Attribution

Controls which factors are candidates for causal explanation. Inside the boundary, causal pathways can be traced. Outside, factors are acknowledged but not attributed to.

Done well
Causal pathways are short and clear; direct drivers are distinguishable from distant influences.
Done badly
Too many external factors included; causal pathways blur. The analysis points in many directions at once and resolves none of them.
Function 03

Interpretability

Keeps explanations specific enough to be acted on. Bounded systems produce concrete recommendations because their components are clearly named and their interactions traced.

Done well
Recommendations are specific, actionable, and tied to named components of the system.
Done badly
"It's a complex interaction of many factors" — true but useless. Recommendations dissolve into general advice when the analysis has not committed to a manageable scope.
The analyst whose scope keeps expanding usually ends up with shaky stability assumptions, blurred causal pathways, and conclusions that no longer fit a single decision.

Boundaries are not permanent — but they are deliberate

A common objection to boundary-setting is that the world does not respect boundaries: factors interact across systems, today's background may be tomorrow's foreground, and the analyst risks missing something important by drawing lines too tightly. The objection is partly correct — but the response is not to abandon boundaries. The response is to set them deliberately, to reconsider them deliberately, and to expand them only when expansion is justified by what the current analysis has revealed.

Boundaries are properly thought of as the scope of the current analytical episode, not the final word on a topic. The first analysis sets a scope appropriate to the immediate decision and time horizon. If that analysis reveals that a factor previously treated as background is actually moving substantially, the next analytical episode can expand the boundary to include it. This is iteration, not expansion-on-the-fly. The boundary holds during a given analysis; it can be revised when a new analysis begins.

Card O · Module 05

Holding the Boundary

The discipline that protects analytical credibility under stakeholder pressure.

Boundary-setting tools are not statistical; they are conceptual. Their purpose is to make explicit what is inside the analysis, what is outside, and what is being assumed about the relationship between the two. This explicitness is itself a form of rigour.

ToolWhat it does
System mapsVisual representations of the components inside the boundary, showing their relationships. What is not on the map is, by design, not part of the analysis.
Scope diagramsDocuments that explicitly list factors inside scope and outside scope, with reasoning for each placement. The act of placing factors on one side or the other forces the analyst to confront ambiguity early.
Assumption registersLists of conditions being assumed to hold — particularly conditions related to factors outside the boundary. Documents what would have to remain true for the analysis to be valid.

Try it: a scope-building exercise

A regional sales team's performance has declined. The decision being supported is whether to restructure the team in the next quarter. Drag each candidate factor to Inside Scope (will be actively investigated) or Outside Scope (treated as fixed background for this analysis). Click Check My Scope when you're done.

Candidate factors — drag each one
Inside Scope · Investigated
Outside Scope · Background
Place each factor in one of the two zones, then check.
Business Example
Holding the boundary under pressure

An analyst is investigating why a regional sales team's performance has declined. The defined scope: the team's pipeline metrics, individual rep activity, and the lead-quality data over the past two quarters. Halfway through the analysis, a stakeholder asks: "What about the broader economic environment? What about competitor pricing? What about our marketing spend in this region?" Each is a reasonable question. But each lies outside the defined boundary, where it was deliberately treated as background.

Expanding the analysis now would dilute the work already done and likely produce conclusions about everything and nothing. The disciplined response: "Good questions. They lie outside the current scope, which was set to focus on the team's pipeline. If our findings point to a need for that broader investigation, that's a follow-up analytical episode. For now, let me complete what we set out to do." The boundary holds. The current analysis produces a clear, decisional conclusion. The broader questions are documented for next time.

Two failures to avoid

Silent expansion mid-analysis

When a finding suggests that something outside the scope might be operating, the disciplined response is to acknowledge the limit of the current analysis and document the question for follow-up — not to quietly enlarge the scope so that the new factor can be included. Mid-analysis expansion undermines coherence: the conclusions no longer apply to a stable scope, and stakeholders cannot tell what the analysis actually examined.

Communicating findings as unbounded

Equally problematic is the opposite failure: refusing to acknowledge boundary limits when they would temper conclusions. An analyst who has carefully bounded their analysis but then communicates findings as if they applied to the full unbounded system has produced misleading work. "Within the scope of this analysis…" is not a hedge; it is an accurate description of what the work supports.

Key Takeaway
Good analysis explains a bounded system clearly, not an unbounded system vaguely.
Module 05 · Assessment

Knowledge Check

Five questions to consolidate what you've learned across Cards M, N, and O. Select the best answer.

Q1 — A SaaS team finds that customers who receive customer-success calls renew at higher rates. They conclude that the calls drive renewals. The most likely mechanism producing this misleading correlation is:

A
A confounding variable that affects both calls and renewal
B
Reverse causality — healthy customers receive more calls because they look healthy
C
A selection effect that filters out unhealthy customers
D
Statistical noise — the relationship is not real
CSMs preferentially call customers showing healthy engagement signals. The healthier customers receive more calls because they are healthier — not the other way around.

Q2 — Which of the following is the foundational question of all causal claims?

A
Is the p-value below 0.05?
B
Does the model fit the data well?
C
What would have happened if the supposed cause had not occurred?
D
Is the correlation coefficient large enough?
Counterfactual reasoning is the foundation of every defensible causal claim. Without it, the analyst is comparing what happened to itself, which can never establish causation.

Q3 — A retailer's underperforming store has three independently sufficient causes: poor location, understaffing, and a misjudged assortment. A marketing campaign targets only the foot-traffic issue. The most likely outcome is:

A
Performance recovers fully — the analyst found the root cause
B
Limited improvement — the other sufficient causes continue producing the outcome
C
Performance worsens because the analysis was incorrect
D
Nothing changes — marketing campaigns never affect retail performance
Multiple sufficient causes each independently produce the outcome. Removing one leaves the others active. This is why single-cause framing maps poorly onto multi-driver realities.

Q4 — Mid-analysis, a stakeholder asks an analyst to "also look at competitor pricing and macro trends" beyond the originally defined scope. The disciplined response is to:

A
Expand the analysis to include them — broader is better
B
Refuse and ignore the questions
C
Document them as follow-up questions for a separate analytical episode
D
Silently widen the boundary to include them without acknowledging it
Mid-analysis expansion undermines coherence. Hold the boundary, complete the current analysis, and treat the broader questions as candidates for a follow-up episode.

Q5 — A regression model includes ten control variables and shows a statistically significant coefficient on X. Which is the most accurate statement?

A
X causes Y, because the controls have ruled out confounding
B
Y causes X, because regression cannot establish direction
C
The coefficient is a conditional association consistent with several causal structures; the analyst's reasoning, not the math alone, must support a causal claim
D
The coefficient is meaningless without an experiment
Regression produces conditional associations, not causal estimates. The same coefficient is consistent with X→Y, Y→X, hidden confounding, and selection effects. Only reasoning about counterfactuals and alternatives turns the number into a causal claim.