The Structure of Analytical Work
Three cards on what shape analysis takes — and why structure decides whether tools can help at all.
Module 01 established the mindset that has to come before tools — bias-aware judgment, decision-anchored framing, and the diagnostic instinct to ask whether the metric is the problem or merely a signal of one. Module 02 builds the skeleton: the shape analysis takes once the mindset is in place. Without that skeleton, even sophisticated methods produce confidently-wrong output. Tools amplify whatever structure they are applied to. Sound structure → sound results. Flawed structure → flawed results dressed up to look credible.
The three cards in this module work as a sequence. Card D establishes that not all analytical problems are the same kind of problem — each demands a different reasoning mode, and each mode justifies a different family of tools. Card E builds the structural scaffolding that prevents the most common silent errors: overlaps that double-count and gaps that hide what matters. Card F makes the structure analytically actionable by converting concepts into variables and variables into testable hypotheses.
Types of Analytical Problems
Three reasoning modes — deductive, inductive, abductive — and why using the wrong one for the problem produces sophisticated answers to questions that were never asked.
Structured Thinking (MECE)
Why most analytical errors are structural, not technical — and how the discipline of mutually exclusive, collectively exhaustive decomposition prevents the overlaps and gaps that downstream tools cannot fix.
From Issue Trees to Hypotheses
The bridge from thinking to measuring. How concepts become variables and variables become testable claims — and why structure alone, however clean, never produces findings on its own.
Module 01 forbade tools in Cards A and B, then permitted only descriptive tools in Card C. Module 02 extends the ladder: Card D introduces the mapping between reasoning mode and tool family. Card E still operates without quantitative tools — its work is purely structural. Card F is the threshold where inferential tools (regression, hypothesis tests, classification) become legitimate for the first time — but only because variables and hypotheses have been disciplined into shape first.
Three Reasoning Modes
Tool-first analysis fails not because the tool is wrong — but because the question asked of it was never the question the tool can answer.
This card addresses a hidden but highly consequential source of analytical failure: using the wrong type of reasoning for the problem at hand. In many organisations, analysts select tools first — regression, dashboards, machine learning — without clarifying what kind of reasoning the problem actually requires. The result is a strange mismatch: a tool designed to test known relationships gets used to discover unknown ones, or a tool designed to find patterns gets used to make causal claims. The output looks rigorous. The math checks out. But the conclusions cannot legitimately answer the question that was asked.
It does not look like a mistake. The analyst used a real tool. The tool produced real numbers. Stakeholders see results that appear sophisticated and believe the analysis is sound. But under the surface, the reasoning is broken. The wrong type of question was asked of the wrong type of tool, and the answer — however precise — cannot legitimately support the decision.
Advanced analytical thinking begins with recognising that not all analytical problems are the same kind of problem. Some require testing known rules or relationships. Others require discovering structure that is not yet known. Still others require explaining observed patterns when neither rules nor structure are fully understood. Each of these situations demands a different reasoning approach, and each approach justifies a different family of tools.
The three modes ahead
Test what you believe
Rules and relationships are known in advance. The analyst tests whether a specific hypothesis holds. General principle → specific case. Highest certainty.
Discover structure
No predefined categories. The analyst lets patterns emerge from data. Specific observations → general patterns. Provisional certainty.
Explain what happened
An outcome is observed; the cause is unclear. The analyst weighs competing explanations and chooses the most plausible. Observed effect → inferred cause. Lowest inherent certainty.
The next tab works through each mode in detail with a business example, then asks you to classify scenarios. The skill being built is conceptual, not technical — but it is upstream of every other analytical decision. Mistakes in reasoning mode cannot be corrected by better tools downstream.
Deductive · Inductive · Abductive
Each starts from a different position, asks a different kind of question, and produces a different kind of answer.
1 · Deductive Reasoning
Deductive reasoning is used when relationships, rules, or theories are known in advance, and the goal is to test whether they hold in the data. The analyst starts with a hypothesis derived from prior knowledge, theory, or business understanding, and checks whether the data supports or contradicts it. The reasoning moves from general principle to specific case.
Deductive reasoning is appropriate when expectations exist and the task is validation. It produces relatively high-certainty conclusions because the structure of the question is precise: a claim is made, evidence is examined, and the claim either holds or fails.
A retail company believes its 10% summer discount drives a meaningful increase in volume among price-sensitive segments. This is a deductive question. The hypothesis is explicit: discount → volume increase. The variables are defined: discount applied (yes/no), volume change, customer segment. The test is straightforward: compare volume in segments that received the discount against those that did not, controlling for the seasonal baseline. The result either supports the hypothesis or contradicts it.
The analyst is not discovering anything new about how the business works; they are testing whether a specific belief holds.
2 · Inductive Reasoning
Inductive reasoning is used when no predefined structure or categories are assumed, and the goal is to discover patterns directly from the data. The analyst does not begin with a specific hypothesis; instead, they allow structure to emerge. The reasoning moves from specific observations toward general patterns.
Inductive reasoning is exploratory by nature. It generates candidate explanations and hypotheses rather than confirmed truths. Because it starts without a fixed expectation, its conclusions are inherently provisional — they describe what the data appears to contain, not what is necessarily true about the underlying system.
A subscription business has not previously categorised its customers in any structured way. The analyst is asked: "Are there natural groupings of customers we should be aware of?" This is an inductive question. There is no hypothesis about how many segments exist, what defines them, or which behaviours matter most. The analyst applies clustering techniques to behavioural data and discovers, for example, that three meaningful groups emerge: high-frequency power users, intermittent occasional users, and a previously invisible segment of dormant accounts that still pay.
None of these segments was specified in advance. They emerged from the data. The output is a hypothesis about how customers naturally group — not a tested fact.
3 · Abductive Reasoning
Abductive reasoning is used when analysts observe patterns or outcomes and must infer the most plausible explanation among several possibilities. It combines evidence, logic, and judgment to arrive at a "best explanation" — while explicitly acknowledging that other explanations remain possible.
Abductive reasoning is essential in complex, real-world systems where controlled testing is difficult or impossible. Most strategic business questions are abductive in nature. The analyst cannot run a controlled experiment on the entire market. They must observe what happened, generate competing explanations, and reason about which one best fits the available evidence.
Of the three modes, abductive reasoning carries the lowest level of inherent certainty. It is also the mode most frequently misrepresented in business communication — where the most plausible explanation is often presented as the proven cause.
A SaaS company observes that 90-day retention dropped sharply for the cohort acquired in March. There is no controlled experiment to run — the cohort is already gone. The analyst must work backward from the outcome. Multiple plausible explanations exist: a March pricing change, a competitor launch, an onboarding flow update, an unusual marketing campaign that attracted the wrong customers, or seasonal effects.
The analyst examines the available evidence, weighs which explanation best fits the timing, magnitude, and pattern of the decline — and concludes that the marketing campaign is the most plausible cause. This is not proof. It is an inference. A disciplined analyst presents it as "the explanation most consistent with the evidence," not as "the cause."
Practice · Reasoning Mode Classifier
Each scenario below describes an analyst facing an actual analytical question. The reasoning mode required is determined by what is known, what is being asked, and what kind of answer would settle it. Classify each scenario, then check.
The Certainty Ladder
The three modes do not produce equally strong conclusions — and the language used to communicate them has to match.
One of the most important insights in this card is that the three reasoning modes do not produce equally strong conclusions. Certainty decreases as we move from deductive to inductive to abductive reasoning, and analysts must adjust their confidence and language accordingly.
This calibration matters because the same finding can be presented with very different levels of confidence depending on how it was reached. "The data shows that X causes Y" is a deductive claim and requires tested evidence. "The data suggests three customer types" is inductive and requires exploratory framing. "The most plausible explanation for the decline is competitive pressure" is abductive and requires acknowledgment that other explanations remain possible.
When analysts borrow the language of one mode while doing the work of another — for example, claiming deductive certainty for an abductive inference — they create a confidence gap. Stakeholders believe they have proven facts when they actually have informed guesses. Decisions made on this basis are more fragile than they appear.
Mapping reasoning mode to tools
Each reasoning mode corresponds to a different family of analytical methods. Confusing these mappings leads to serious analytical errors — often hidden under sophisticated tooling. The mapping is not a matter of preference; it is a matter of legitimacy. Tools designed for one mode of reasoning cannot answer questions from another.
| Mode | Legitimate methods | What they can — and cannot — do |
|---|---|---|
| Deductive | Hypothesis tests, regression models, statistical comparisons, A/B tests, controlled experiments | Can test specified relationships against data. Cannot discover relationships that were not specified in advance. |
| Inductive | Clustering techniques, exploratory decision trees, dimensionality reduction, association rule mining | Can surface patterns, segments, or rules from the data. Cannot prove causality or confirm patterns will persist. |
| Abductive | Hypothesis synthesis, causal logic, comparative case reasoning, system understanding — supported by data, not replaced by it | Can structure the reasoning behind a best-explanation inference. No single algorithm produces abductive conclusions; they emerge from disciplined judgment. |
A common confusion: classification ≠ discovery
Among the most frequent and damaging confusions in modern analytics is the misuse of classification and predictive tools as if they were discovery methods. These tools feel exploratory because they involve algorithms, training data, and machine learning — the visual and procedural language of discovery. But they are not discovery tools.
Methods such as K-Nearest Neighbors (KNN), classification trees, and supervised learning models rely on existing labelled knowledge. They take pre-defined categories and apply them to new observations. They are tools for prediction or classification — not for discovering new structure.
When an analyst uses a classifier to "identify" customer segments that were already labelled in training data, they have not discovered anything. They have demonstrated that the algorithm can apply known labels. The structure was input, not output. Treating the result as a discovery overstates what was actually found.
The distinction is subtle but consequential. A clustering algorithm starts without labels and lets structure emerge — inductive. A classification algorithm starts with labels and learns to apply them — not inductive, regardless of how complex the math becomes. When stakeholders ask "what kinds of customers do we have?" the answer requires an inductive method. When they ask "given these known types, which one does this customer belong to?" the answer requires a classification method. Mixing them produces conclusions that look like discoveries but are merely re-applications of pre-existing knowledge.
One question, three different analyses
The same business topic produces three very different analytical projects depending on what is actually being asked. Consider customer retention at a B2B software vendor:
| If the question is… | It is… | The analyst should… |
|---|---|---|
| "We believe customers who use feature X are less likely to churn. Is that correct?" | Deductive | Define the hypothesis precisely, identify comparison groups (users vs non-users of X), control for confounders, test via regression or A/B. Output: confirmed or rejected hypothesis. |
| "We don't really know our customer base anymore. Are there meaningful groupings we're missing?" | Inductive | Apply unsupervised clustering to behavioural and account data, allow segments to emerge, present them as candidate hypotheses. Output: provisional structure for further investigation. |
| "Our retention dropped 8 points last quarter and we don't know why. What's the most likely explanation?" | Abductive | Generate multiple plausible explanations (pricing, competitor, product, seasonality, sample composition), examine evidence for each, present the most plausible while acknowledging alternatives. Output: best-explanation inference, not a proven cause. |
The Structural Trap
Most analytical errors are structural, not technical — and the most damaging happen before any calculation begins.
This card addresses a deceptively simple but critical reality: most analytical errors are structural, not technical. Analysts often assume that once the right data and tools are used, correct answers will follow. The implicit belief is that errors arise from mistakes in calculation, sampling, or modelling — and that careful methodology will prevent them.
In practice, the most damaging errors occur far earlier — in the way the problem is decomposed before any calculation begins. Real-world problems are messy, multi-causal, and interconnected. When analysts decompose them poorly, two failure modes appear, both invisible to the analyst at the time:
The same driver double-counted
A single underlying driver shows up under multiple categories, getting counted twice and inflating its apparent importance. The analyst sees two large drivers when there is only one.
The driver that's silently missing
A driver is excluded from the structure, leaving an entire region of the problem unexamined. The blind spot is not visible from inside the analysis — the picture feels complete because it was complete from a limited frame.
Neither of these errors is corrected by sophisticated methods downstream. A regression model run on a structure with overlapping categories will report inflated effects. A clustering algorithm applied to a structure missing key dimensions will discover patterns that omit the actual drivers. The technical machinery is working perfectly; the inputs are quietly wrong.
If the structure is sound, sophisticated tools produce sound results. If the structure is flawed, those same tools produce confidently wrong results that look more credible than they deserve to. This is the core trap of Card E: structural errors are amplified by good tools, not corrected by them.
A flawed first attempt
To make the trap concrete, consider an analyst asked to decompose: "Why did revenue decline 12% last quarter?" A common first attempt produces a list that feels thorough — and is structurally broken:
Pricing changes · Marketing performance · Customer dissatisfaction · Competitor activity · Product issues · Sales team effectiveness · Economic factors
The list fails on both counts:
| Failure | Where it shows up in this list |
|---|---|
| Overlap (not mutually exclusive) | "Customer dissatisfaction" overlaps with "product issues," "pricing changes," and "sales team effectiveness" — any of these could cause dissatisfaction. "Marketing performance" overlaps with "competitor activity" since marketing performance is partly defined by competition. The same revenue effect could be assigned to multiple categories depending on framing. |
| Gaps (not collectively exhaustive) | Where do channel mix shifts go? Where does seasonality fit? What about changes in customer composition — the same number of customers but a different mix? None of these have a clear home in the structure, so they will likely be missed entirely. |
Worse, the list has the appearance of completeness. An analyst who builds this structure feels they have covered the problem. They have not. They have created a structure that will systematically distort the analysis from this point forward. The next tab introduces the discipline that prevents this — MECE.
The Two MECE Requirements
Mutually Exclusive · Collectively Exhaustive — both non-negotiable, both invisible failure modes when broken.
MECE stands for Mutually Exclusive, Collectively Exhaustive. The principle has two components, both non-negotiable. A structure that satisfies one but not the other is not MECE — and the analytical risks of partial MECE are nearly as severe as the risks of no structure at all.
Mutually Exclusive — no overlaps
Each category or branch represents a distinct driver, with no overlap between categories. A finding that falls into one category cannot also fall into another. The boundaries are clean.
When branches overlap, the same effect can be attributed to multiple causes. This inflates perceived importance and leads to incorrect prioritisation. If "marketing problems" and "customer dissatisfaction" are both branches of a structure, every customer who left because of a poor advertising experience appears under both branches. The analyst sees two large drivers when there is only one.
Collectively Exhaustive — no gaps
All plausible drivers are included within the structure. Nothing relevant is missing. Together, the categories account for the full space of the problem.
When structures are not exhaustive, key factors are silently excluded from analysis. A structure that examines pricing, marketing, and product quality but omits competitive dynamics will produce conclusions that ignore the most important driver in markets where competition is the real force at work. The blind spot is not visible from inside the analysis.
Both requirements must hold simultaneously. A structure that is mutually exclusive but not collectively exhaustive misses real drivers. A structure that is collectively exhaustive but not mutually exclusive double-counts. Only when both requirements are met does the structure provide reliable scaffolding for the analysis that follows.
The almost-MECE traps
One of the most insidious failure patterns in structured thinking is the almost-MECE structure — a decomposition that looks neat (often with parallel categories and consistent labelling) but quietly violates MECE in subtle ways. Almost-MECE is more dangerous than no structure at all, because it creates the illusion of rigor where none exists.
Soft overlaps
Categories that mostly do not overlap, but share a fuzzy boundary the analyst hopes will be ignored. "operational issues" and "process problems" as separate categories — they will overlap in practice.
Mismatched abstraction
Categories at different levels of detail. A structure that lists "pricing," "marketing," "product," and "South region" — the first three are functions; the last is a geography. They do not belong in the same structure.
The catch-all category
A final category labelled "other" or "miscellaneous" added to make the structure feel exhaustive. This is a structural confession — it admits the analyst could not articulate what is missing. Catch-all categories quickly become hiding places for the most important drivers.
Disciplined analysts recognise these patterns and reject them. A structure that is "almost MECE" is not MECE. It must be revised until both requirements are genuinely satisfied — even if that means restarting the decomposition from a different angle.
Practice · MECE Auditor
Each scenario presents an analytical question and a proposed decomposition. Your job is to audit it: for each branch, decide whether it's clean or which trap it falls into. Then identify what's missing from the structure as a whole.
From Bad MECE to Good MECE
Clean structure usually comes from underlying logic, not from listing what comes to mind.
A clean MECE structure for the revenue-decline question starts from the underlying mathematics of revenue. Revenue is the product of how many things were sold and how much each one earned. This forces a decomposition that is both exclusive and exhaustive by construction:
Volume changes — units sold went up or down.
Price changes — average price per unit went up or down.
Mix changes — composition of what was sold shifted toward different products, segments, or channels with different economics.
This structure is genuinely MECE. Every revenue change must come from one or more of these three sources — there is no fourth possibility. And no revenue change can be assigned to more than one source without explicit decomposition. Once the analyst knows whether the decline came from volume, price, or mix, they can drill down within the relevant branch using sub-structures that are themselves MECE.
MECE is iterative, not one-shot
Initial structures are rarely perfect. The first attempt at decomposing a problem usually contains overlaps, gaps, or both. This is not a failure of skill — it is the natural result of working on complex problems. The discipline is not to produce a perfect structure on the first try, but to refine it as understanding deepens.
Three questions drive each refinement:
Where do my categories overlap?
If I find a real example that fits more than one category, the structure must be revised.
What might be missing?
If I imagine an unexpected explanation, where would it go in my structure? If there is no place for it, the structure is not exhaustive.
Are my categories at the same level of abstraction?
If some categories are causes and others are symptoms, the structure mixes layers. Re-anchor at one level.
Structural tools, no numbers yet
At this stage, the only legitimate tools are logical decomposition methods. They don't produce numbers; they produce structure. The discipline of building scaffolding before applying tools is what separates rigorous analysis from data exploration.
| Tool | What it does | When to use |
|---|---|---|
| Issue trees | Hierarchical decomposition of a problem into progressively more specific sub-questions or drivers, branching downward. | A complex business question must be broken into investigable parts. The standard tool for top-down analytical decomposition. |
| Logic trees | Decomposition driven by the logical structure of the problem ("what must be true for X to occur?"), often used for diagnostic root causation. | The analyst needs to map the conditions that could produce an outcome. Useful for diagnostic and hypothesis-generating work. |
| Driver frameworks | Pre-built structures decomposing specific business outcomes (revenue, profit, customer value) into their mathematical or logical components. | A well-understood outcome needs analysis and an established decomposition already exists. |
From Thinking to Measuring
Structure alone never produces findings — the bridge to analysis is the conversion of concepts into variables and hypotheses.
This card addresses a common breakdown in analytical work: analysts invest significant effort in framing problems and building structured issue trees, yet fail to translate those structures into analysis that can actually be executed. The structure exists. The decomposition is clean. The MECE test is passed. And then the work stalls — because the structure itself does not specify what to measure, what to compare, or what would constitute evidence.
The result is a familiar pattern: beautifully decomposed problems sit on slides as conceptual artefacts. They look rigorous. They organise discussion. But they do not produce findings. The branches of the tree describe categories of drivers ("pricing," "product," "competitive dynamics"), but no one knows what data to pull, what test to run, or what answer would settle the question. The structure has been built; the analysis has not begun.
Advanced analytical thinking requires recognising that structure alone does not create analysis. A MECE issue tree is necessary but not sufficient. Without the conversion from concepts to measurable elements, even perfectly structured frameworks remain disconnected from data and from decision-making. The work feels analytical but produces nothing actionable.
Concepts · Variables · Hypotheses
Three categories must be clearly distinguished. Most analysts use these words interchangeably in casual conversation, but in disciplined analytical work they refer to very different things. Confusing them is one of the most common reasons structured frameworks fail to translate into actual analysis.
| Category | What it is | What it produces |
|---|---|---|
| Concept | An idea, theme, or area of concern. Usually a noun phrase. "customer satisfaction," "process efficiency," "market position," "brand strength." | Discussion. Concepts can be talked about and debated, but they cannot be measured directly. They name the thing of interest without specifying how to observe it. |
| Variable | A measurable expression of a concept. Specifies unit, source, time frame, boundary. "NPS score from quarterly customer survey," "average cycle time per order in days," "market share by revenue in the segment." | Numbers. Variables produce data. They can be tracked, compared across periods or segments, and aggregated. They convert ideas into observations. |
| Hypothesis | A specific, testable claim about how variables relate or differ between groups. "Customers who use feature X have lower 90-day churn than those who do not." "Average cycle time has increased materially in the last six months." | Findings. Hypotheses can be confirmed, rejected, or partially supported. They make analysis directional and decisional rather than descriptive. |
These three categories form a sequence. Concepts come from the issue tree. Variables come from translating concepts into measurement. Hypotheses come from connecting variables in claims that can be tested. Skipping any step weakens the analysis. Skipping all of them produces analytical work that never moves beyond description.
The question that drives the conversion
Every branch of an issue tree must answer a single, demanding question:
This question forces the abstract to become concrete. "Customer satisfaction" is a concept; the analyst cannot study it directly. But the question "How would we know if customer satisfaction is changing?" demands a specific answer — a survey score, a complaint rate, a renewal rate, an NPS. The act of answering converts the concept into one or more variables.
If the analyst cannot answer "How would we know if this matters?" for a given branch, the branch cannot be analysed. A branch that cannot be measured is not analytically actionable, regardless of how important it sounds. When this happens, one of three things must occur: refine the branch into something concrete; replace it with a different formulation; or acknowledge it as out of scope for this analysis — deliberately, with reasoning documented.
This prevents a particular failure mode: branches that sound important but float free of any data. When stakeholders ask about "strategic alignment" or "organisational culture" in the context of an analytical question, the analyst's job is not to argue these things matter. It is to ask: how would we know if they are operating on this problem? If no measurement can be defined, the branch belongs to a different conversation — not this one.
Concept → Variable → Hypothesis
Each stage adds specificity. By the end, the analyst knows exactly what data to pull, what comparison to make, and what answer would settle the question.
To illustrate the full conversion, consider an issue tree branch that emerged from a MECE decomposition of customer churn. The branch reads, simply, "customer satisfaction." As a concept, this is reasonable. As an analytical input, it is unusable. The conversion transforms it into something that can drive decisions.
| Stage | Statement | Status |
|---|---|---|
| Concept (raw) | "Customer satisfaction" | Discussion-level only. No measurement defined. Cannot be analysed yet. |
| Concept (refined) | "Customer satisfaction with the product experience over the past 90 days" | Discussion-level, but specific. Tells us what to measure conceptually, but not how. |
| Variables | Quarterly NPS score (survey-based) · Monthly support ticket rate per active account · 90-day product-feature-usage retention | Measurable, but not yet investigative. We can track these numbers, but we have not stated what we expect or what we want to test. |
| Hypothesis | "Accounts with NPS < 7 in the past two quarters have at least 3x the 90-day churn rate of accounts with NPS ≥ 7, after controlling for tenure and contract size." | Testable. This is now an analytical claim that can be confirmed, rejected, or refined with data. |
Notice how each stage adds specificity. The concept becomes a refined concept. The refined concept becomes one or more variables. The variables come together in a hypothesis with a directional claim, a comparison group, and an explicit control structure. By the time the hypothesis is stated, the analyst knows exactly what data to pull, what comparison to make, and what answer would settle the question.
An analyst who stops at the concept level — "customer satisfaction" — typically produces a deck that shows several charts about NPS trends, support tickets, and survey responses. The deck is descriptive. The conclusions are vague: "NPS has declined," "tickets are up," "satisfaction is a concern." Stakeholders nod. Nothing specific is decided.
The analyst who has formulated an explicit hypothesis produces a finding: "Low-NPS accounts churn at three times the rate of high-NPS accounts; addressing satisfaction in this segment is worth approximately $X in retained revenue." The first analyst described the situation. The second answered a question that drives action.
Variables alone are not enough
A common shortcut at this stage is to define variables but skip the hypothesis layer. The analyst identifies what to measure and assumes that examining the data will produce findings. This shortcut is one of the most reliable producers of inconclusive analytical work.
Variables without hypotheses produce description, not investigation. The analyst pulls the data, examines it, and presents what they see. Trends are noted. Differences are reported. But there is no benchmark for what would be surprising, no comparison that would settle a question, no claim that the data could either support or refute. The output is informational at best.
Hypotheses transform variables into investigations. A hypothesis specifies what relationship is expected, what magnitude would be meaningful, and what evidence would support or contradict the claim. With these specifications in place, looking at the data becomes a test rather than a tour.
Three properties of a useful hypothesis
Specificity
The relationship or comparison is named precisely. Vague claims ("X may relate to Y") do not qualify. Directional, magnitude-bearing claims do: "X is associated with at least 2x higher Y, after controlling for Z."
Falsifiability
The hypothesis can be wrong. There must be evidence that, if observed, would force its rejection. A claim that cannot be wrong is not a hypothesis — it is a belief.
Connection to structure
The hypothesis traces back to a specific branch of the MECE issue tree. It is not a free-floating curiosity; it is the analytical expression of a particular driver in the structured decomposition.
Practice · Hypothesis Builder
For each scenario, a branch from a MECE issue tree is given. Select the best variable definition, then select the best hypothesis. Distractors include vague claims, statements that aren't really hypotheses, and variable definitions that don't actually pin down measurement.
Tools Finally Open Up
The tool legitimacy ladder has been climbing all programme — Card F is where inferential methods become legitimate for the first time.
This card marks an important transition in the programme's tool legitimacy progression. Once structure has been translated into variables and hypotheses, formal analytical tools become legitimate for the first time — because there is now something concrete for those tools to operate on. Variables have been defined. Hypotheses have been articulated. The analyst can finally select methods that are designed to test specific claims rather than to explore unstructured data.
Appropriate methods at this stage
| Method | What it does |
|---|---|
| Descriptive statistics | Means, medians, distributions, ranges that characterise each variable. Foundational for understanding the data before testing relationships. |
| Comparisons across segments and periods | Side-by-side examinations that test whether differences exist between groups — the simplest form of hypothesis test. |
| Regression models | Quantify the relationship between variables, control for confounders, test directional claims. Appropriate when hypotheses are specific and data structure is understood. |
| Classification techniques | Predict category membership based on input variables. Note: classification is a tool of prediction, not discovery — see Card D. |
| Hypothesis tests | Formal statistical tests that quantify the strength of evidence for or against a specific claim. |
Tool choice is subordinate to variable definition. No analytical method can compensate for poorly defined variables or ambiguous hypotheses. A regression run on vague variables produces vague results dressed up as precise ones. A hypothesis test applied to a fuzzy claim produces a p-value that does not actually correspond to what the analyst thought they were testing.
Conversely, well-defined variables enable multiple methods to be applied appropriately as the analysis progresses. The same variables that support descriptive statistics can support regression. The same hypothesis that supports a simple comparison can support more sophisticated controls. Disciplined definition expands the range of legitimate methods. Sloppy definition collapses it.
The principle Modules 03 and 04 build on
This card reinforces a principle that will recur throughout the rest of the programme: tools serve hypotheses, not the other way around. The hypothesis defines what is being investigated. The tool is selected because it is the right instrument for that specific investigation. When this order is reversed — when a tool is chosen first and the hypothesis is shaped to fit the tool — analytical errors multiply.