2 · Theory and causality

Wednesday, September 2, 2026

Slides

What you should be able to do after this meeting. Draw the causal graph behind a claim somebody is acting on, name the failure in it, and say which variable you must condition on and which you must not touch. Leave with your team formed, its ground rules in writing, and a first graph of your own project that the room has heard.

NoteWhat to take away

Two questions to ask of every number you are shown. Compared to what? And would those two groups have looked the same anyway?

Three ways a comparison breaks. A confounder causes both the treatment and the outcome, so the two groups were different before anything happened. A collider is caused by both, so looking only inside it invents a link that is not there. A mediator carries the effect, so controlling for it deletes the thing you were measuring.

The rule. Condition on the confounder. Leave the mediator alone. Do not select your sample on the result.

For your proposal: the graph you draw is your conceptual model heading, and the comparison it names, with what must be true for it, is your identification heading.

Due today

Your team’s ground rules, signed, and your project’s first causal graph. Both before you leave.

Session A · 2:05–3:15

The whole hour answers one question: compared to what? Four blocks answer a piece of it.

Theory, and where the comparison comes from

A theory is a set of assumptions that produces a prediction you could be wrong about. That is the entire definition, and the part carrying it is could be wrong. A claim that survives every possible result is a description dressed up.

A theory buys you four things. It tells you which comparison to make. It tells you what sign to expect. It tells you what to hold fixed, and what not to. And it tells you what result would embarrass you, stated before you look.

A prediction is not a theory. A prediction says what will happen, and you score it on accuracy. A theory says why, and therefore what happens if you change something. Only the second one can be acted on. This matters more than it sounds, because a model can fit beautifully and be about nothing. Your churn model says support contacts predict leaving. That is not a reason to believe fewer support contacts would keep anyone.

A prediction scored on accuracy. It can be right for no reason Support calls ? Leaves change something ? A theory scored on whether its assumptions survive Support calls leaves when annoyance > switching cost Leaves change something

It is also why a number travels or does not. A 40% lift measured in one market says nothing about another unless you know the mechanism behind it. Without one you have a local fact with an expiry date.

Google Flu Trends is the cautionary case. It predicted influenza-like illness from search queries, tracked the CDC closely, and was widely admired. By February 2013 it was predicting more than double the CDC figure it had been built to predict. Lazer, Kennedy, King and Vespignani argue the errors were largely avoidable. Part of the problem was that Google’s own search algorithm kept changing underneath the model. Nobody could say why it worked, so nobody could say when it would stop.

% of doctor visits for flu-like illness Jan 2012 Jan 2013 about twice the CDC figure Google Flu Trends CDC count

Schematic, redrawn from Fig. 1 of Lazer, Kennedy, King and Vespignani (2014). Shapes, not data.

Two worked theories, both from markets you will work in.

Akerlof’s market for lemons starts from three assumptions. Sellers know the quality of their car, buyers do not, and one price clears the market. So the price is what an average car is worth, owners of good cars withdraw rather than sell at it, the average falls, and the price follows. The prediction is that good used cars stay off the market. What makes it a good example for this course is that the fixes come free. Akerlof’s own section on counteracting institutions names guarantees, brand names and chains, and licensing. Anything that lets a good seller prove they are one. The intervention list is not a separate brainstorm; it falls out of the assumptions. Look at any used-car marketplace today and you will find all four.

Then a dispute, which is the better example. Spence’s signalling model and human capital both predict that graduates earn more. Human capital says school makes you more productive. Signalling says school reveals who already was. They cannot both be the reason, and the policy stakes are opposite: if human capital is right, subsidising schooling grows the pie, and if signalling is right it mostly reshuffles who gets which job.

The disagreement is what tells you the comparison. Human capital predicts a smooth return, since each year adds learning. Signalling predicts a jump at the certificate, since the final year is the year you get the paper. So compare someone who finished with someone who did three and a half years and left. Same schooling, different paper. That is the sheepskin effect, and substantial premia at the diploma year are robustly found. They are evidence that some of the return is signalling. They do not settle the split, and almost nobody argues it is all one or all the other.

Wage human capital: each year adds the same signalling: flat, then a jump at the diploma 12 years 15½ left 16 finished Years of schooling

Wage against years of schooling under the two theories. Schematic. The comparison is the half-year between the two dashed lines.

The general lesson. A theory earns its keep by telling you which comparison settles the argument. That is the same job the identification section of your proposal does in December.

Try it. Which of these is a theory, which a description, which a prediction, and which a value judgment?

1 Sales are higher in December
2 Customers who contact support twice a month are three times more likely to leave
3 Customers leave when the cost of switching drops below the annoyance of staying
4 Firms should invest in sustainability
5 Firms maximise profit, and whatever they did was the profit-maximising thing
Answers

Description, prediction, theory, value judgment, and a statement that cannot be wrong, which is the tell that it is not a theory. Two and three are about the same thing. Only three tells you what to change, and it names two levers: the cost of switching and the annoyance of staying.

What empirical work can and cannot settle

It can measure a magnitude. It can rule out a sign. And it can choose between two theories when they predict different things about the same comparison. That third one is the one students under-rate, and it is the one Spence illustrates. A comparison both theories predict identically is not evidence, however clean the estimate.

What it cannot do is test a theory on its own. Every test is a joint test of four things: the theory, the measure, the sample, and the model. When the result disagrees with you, it never says which of the four broke. In practice the weakest link is usually the measure, and the first move on a surprising result is to interrogate it rather than to abandon the idea.

Theory Measure Sample Model No the idea the variable the units the equation the result ? ? ? ? Which link broke? The number does not say.

And no amount of data tells you what to want. Data can price a trade-off. It cannot say the trade is worth making. That line is where analysis stops and judgment starts, and crossing it quietly is the most common way an analyst overreaches. Your job is to make the trade-off visible and priced, then say plainly that the choice belongs to the person deciding.

Try it. Can data settle each of these: yes, no, or only with an assumption you would have to defend?

1 Does the loyalty app increase spending?
2 Is the minimum wage too high?
3 Did the ad campaign pay for itself?
4 Will this model still work next year?
5 Should we prioritise growth over margin?
Answers

One, yes, with a design. Two, no: “too high” is a judgment. Three, only with an assumption about what would have happened without the campaign. Four, only with an assumption, because whether the model still works depends on whether the mechanism holds, and that is the Google Flu problem again. Five, no: that is what you want. Two and five are not hard empirical questions. They are not empirical questions.

For your proposal

For your proposal this means something specific. You are not proving your theory. You are showing which comparison your design makes, and stating what would have to be true for that comparison to mean what you claim. That paragraph is the identification section, it is the hardest page in the document, and it is worth more than the results.

Causality

Holland’s framing comes first. Every unit has two outcomes, the one with the treatment and the one without, and you only ever observe one of them. The causal effect for a single customer is the difference between two numbers, one of which does not exist.

That makes it a missing data problem rather than a statistics problem. No estimator recovers a number that was never observed. Every method in this course is a different argument about how to fill the missing column in, and the statistics comes after you have won that argument.

So every comparison you make is a stand-in for a world that did not happen, and the only question worth asking about a result is whether the stand-in was any good. Take the claim the meeting opens with: customers on the loyalty app spend 40% more. The counterfactual is what those same customers would have spent without the app. It is not what the other customers spent. Ask who signs up for a loyalty app and the answer is people who already shop there a lot, which is most of the 40%.

Heavy shopper On the app Spend close it: compare like with like

Two questions do most of the work, for the rest of your career. Compared to what? And would those two groups have looked the same anyway?

Try it. A store gets a refit and sales rise 12% the following quarter. Which stand-in is best?

1 The same store, the quarter before
2 Every other store in the chain, same quarter
3 Stores that were also due a refit but have not had it yet
Answers

One ignores everything else that changed over the quarter. Two compares stores that were picked for a refit with stores that were not, and they were picked for a reason. Three is the closest, if the queue order has nothing to do with sales. It is still an assumption. There is no comparison that needs no assumption, only comparisons whose assumption you can state and defend.

A causal graph is arrows between variables, and it lets you argue about identification before you own a single row of data. That is exactly where your team is this afternoon. A graph is also a theory drawn, so this is the first block in another notation. Every arrow is a claim you have to defend, and every absent arrow is a claim too, usually a stronger one.

Three structures, and on every diagram today the nodes sit in the same places so only the arrows move.

A back door is a second route from the treatment to the outcome that is not the effect you want. Conditioning on a variable means comparing within its values. Controlling for it in a regression is the same act, so the two words are used interchangeably below.

Confounder. Causes both the treatment and the outcome. It opens a back door and the back door has to be closed. In a meeting you say: the two groups were already different before anything happened.

Z X Y condition on Z and the back door closes

Collider. Caused by both. Conditioning on it manufactures a relationship that was never there. The arrows point into Z where the confounder’s pointed out of it. In a meeting you say: you only looked at the winners, so you invented a trade-off.

Z X Y condition on Z and you invent the link

Mediator. Sits on the path from cause to effect. Conditioning on it deletes the effect you came to measure. In a meeting you say: you controlled away the thing you were looking for.

Z X Y condition on Z and you delete the effect

The rule is six words. Close the back doors. Leave the front door alone. Condition on the confounder. Leave the mediator alone. And do not select your sample on the result, because that is conditioning on a collider. All three, in one picture:

Confounder condition on it Treatment X Mediator leave it alone Outcome Y Collider leave it alone box the top one and nothing else

Which is why controlling for everything makes things worse. Every control is a claim about the graph, not a safety measure, and a kitchen-sink regression is a graph nobody drew. If a coefficient moves when a control goes in, that is neither good news nor bad news until you can say which of the three that variable is.

The Berkeley admissions case is the cleanest illustration. In 1973 Berkeley admitted about 44% of male applicants and about 35% of female applicants. Control for department and the gap largely vanishes, because women applied in larger numbers to departments admitting fewer people. Now the question that matters: is department a confounder or a mediator? If applicants chose departments for reasons unconnected to admissions, it is a confounder and you should condition on it. If they were steered there, it sits on the path and conditioning removes part of what you were measuring. Same variable, same data, opposite handling. Only the causal question tells you which.

Translating between plain language and the vocabulary

Most of the causal claims you will meet at work arrive in plain English, and your first job is to hear which structure is hiding in the sentence.

What you actually hear What it is
Firms that do X grow faster confounder
Students who eat breakfast score higher confounder
Among the people we hired, the good testers interviewed badly collider
Nine of the top ten founders dropped out collider, by survivorship
Once we control for X the gap disappears mediator
More police, more crime reverse causality
We sent the team to our worst stores and they improved regression to the mean
Reported cybercrime doubled the measure changed

The reverse direction matters as much. You will have to explain the diagnosis to somebody who has never heard the word collider, and “you only looked at the winners” is the sentence that lands.

Those eight names are not eight separate things. Three are graph structures. One, reverse causality, is an arrow pointing the wrong way. And two cannot be drawn at all: regression to the mean is about picking extremes out of noise, and a changed measure is about measurement. Neither is fixed by conditioning on anything. Selection and survivorship are colliders wearing different clothes.

Three claims, worked in the room

The hiring puzzle. Among the people we hired, the ones with the best test scores had the worst interviews, so the test must be measuring the wrong thing. Either a strong test or a strong interview was enough to be hired. So among the hired, a weak interview implies a strong test, and the trade-off is manufactured by looking only inside the box. In the applicant pool the two are unrelated. Scrapping the test would be acting on a pattern that exists only in the selected sample. The fix is to look at applicants, not hires.

Hired Test score Interview we only ever look inside the box

Good to Great. Jim Collins identified eleven companies that beat the market by at least three times over fifteen years after a transition point, then looked for what they had in common. The book is careful, it uses comparison companies, and the research effort was real. The problem is structural and no care inside the design fixes it: the eleven were chosen because they became great, so whatever else also caused that is now correlated with the practices inside the sample. The missing comparison is the firms that did all the same things and did not become great, and they are not in the book. Circuit City filed for bankruptcy in 2008 and Fannie Mae was placed into conservatorship the same year. Those failures are not the argument, since plenty of good companies fail. They are just why people noticed. The argument is the structure, and it is the same one as the hiring puzzle, arriving from a completely different direction.

Became great The practice Everything else the sample is the eleven inside the box

The promotion gap. Once we control for salary band the gender gap in promotion disappears, so there is no bias. If the bias operates by placing women in lower bands, the band is how the effect travels, and controlling for it controls away the thing you were looking for. The disappearance is not the finding, it is the mechanism. The honest caveat is the Berkeley one: if bands were set before anything the firm did and for reasons outside it, the band could be a confounder instead. Which it is depends on the causal question rather than on the data.

Salary band Gender Promotion the band is how the bias operates

Every research design in this course closes a back door in a different way. You will draw each one as a graph before you hear its name.

Session B · 3:30–4:30

Your team is already yours. Teams are posted before the break and there is a numbered sheet on each table, so the first three minutes are sitting down rather than listening to a list.

Ground rules, one page, signed by everyone. Written while the team still gets along, which is the only time it is easy. Who does what. What happens when somebody misses a deadline. How you settle a disagreement, and who breaks a tie. There is more on what belongs in it under working in a team.

Then your own project, on paper, by hand. Your team has a shared interest and not yet a question. Take the closest thing you have and draw it: the treatment, the outcome, the one back door you would have to close, and the one variable you must not touch.

Then every team reports. Two and a half minutes each, hard stop, on a timer. Your question, and the graph you drew, and nothing else. You hear every other project before you leave, and this is the only afternoon all term when the whole room sees what everybody chose.

Your sheet is disposable and it will be wrong, which is the point of drawing it in September. You redraw it when we reach research designs, and the version that survives the room becomes the figure in your identification section.

Both sheets go in before you leave.

Handouts

Causal claims, eight practice cases — eight claims, each broken in a different way, with a diagnosis sheet. Three of them, the loyalty app, the hiring puzzle and the promotion gap, are worked above. The other five are practice.

Project teams — the nine questions this class is asking, and who is working on each. The four criteria the teams were built on, and what to do if you think you are on the wrong one.

Where to read more

The Effect is the gentlest book on the shelf and the best match for this meeting. Chapter 5 on identification, chapters 6 and 7 on causal diagrams, chapter 8 on closing back doors. Start here.

Causal Inference: The Remix chapter 3 on directed acyclic graphs and chapter 4 on potential outcomes, for the same material with the machinery attached.

Mostly Harmless Econometrics chapter 1, Questions about Questions, is eight pages and needs no mathematics. The rest of that book assumes econometrics you have not been taught, so read chapter 1 or nothing.