AEM 6991

MPS Capstone Project I

Prof. Ariel Ortiz-Bobea

2 · Theory and causality

Wednesday, September 2, 2026

Cornell University

Theory

The board is asked to act on a 40% gap

A grocery chain. Half its customers have installed the loyalty app. The analytics team reports to the board:

$140
a month, customers on the app

$100
a month, customers not on the app

“App customers spend 40% more. Sign the other half up and they will too.”

The board is asked to fund the sign-up campaign. Does it? Thumb up or down.

A theory makes a prediction you could be wrong about

  • A theory is a structured framework of ideas used to explain facts, behaviours or phenomena.
  • It starts from assumptions and ends in predictions.
  • Four checks:
  1. It says why, not only what.
  2. It states its assumptions.
  3. It predicts what you would see if it is right.
  4. It says what you would see if it is wrong.

In a meeting you say: what has to be true for this to work, and what would prove me wrong?

A prediction is not a theory

  • A prediction says what will happen. You judge it by how often it is right. It can be right for no reason.
  • A theory says why it happens. It names its assumptions, so it also says what changes if you change something.
  • Only a theory tells you what to do.

A prediction with no theory behind it fails without warning

  • Google predicted flu doctor visits from search terms.
  • For years it tracked the CDC count well.
  • In February 2013 it predicted over twice the CDC count.
  • Searching changed, not illness. The flu was in the news.
% of doctor visits for flu-like illness Jan 2012 Jan 2013 about twice the CDC figure Google Flu Trends CDC count

Schematic after Lazer et al. (2014), The Parable of Google Flu, Science 343, 1203–1205. Shapes, not data.

Which of these is a theory?

1 Sales are higher in December
2 Customers who contact support twice a month are three times more likely to leave
3 Customers leave when the cost of switching drops below the annoyance of staying
4 Firms should invest in sustainability
5 Firms maximise profit, and whatever they did was the profit-maximising thing

Hold up 1, 2, 3 or 4 fingers. 1 theory · 2 description · 3 prediction · 4 value judgment

Only one of the five is a theory

1 Sales are higher in December description
2 Customers who contact support twice a month are three times more likely to leave prediction
3 Customers leave when the cost of switching drops below the annoyance of staying theory
4 Firms should invest in sustainability value judgment
5 Firms maximise profit, and whatever they did was the profit-maximising thing cannot be wrong

Two and three are the same subject. Only three tells you what to change.

Two theories predict the same thing for different reasons

Human capital

  • School makes you more productive.
  • The degree raises your wage because you learned something.

Signalling

  • School reveals who was already productive.
  • The degree raises your wage because it separates you.

Both predict that graduates earn more. They cannot both be the reason.

Take thirty seconds. What would you look at to tell them apart?

Spence (1973), Job Market Signaling, Quarterly Journal of Economics 87(3), 355–374

Signalling predicts a jump at the diploma, not a slope

  • Human capital: the return is smooth. Each year of school adds learning, so each year adds wage.
  • Signalling: the return jumps at the certificate. The final year is worth more, because that is the year you get the paper.

Compare someone who finished with someone who did three and a half years and left. Same schooling, different paper.

The “sheepskin effect”. Hungerford and Solon (1987), Review of Economics and Statistics 69(1), 175–177, and a literature after it, find substantial premia at the diploma year

What data can settle

Data can measure, rule out a sign, and choose between theories

Measure a magnitude how big, with what uncertainty
Rule out a sign it is not positive, whatever else it is
Choose between two theories when they predict different things about the same comparison

When you test with data, you test four things at once

  • Your theory says the app raises spending. You test it. The number says no.
  1. The theory. The app does not raise spending.
  2. The measure. You counted in-store spend and missed online.
  3. The sample. You looked at new customers only.
  4. The analysis. You fitted a straight line to something that curves.
  • Any one of the four produces that “no”. The number does not say which.

In a meeting you say: the number disagrees with us, and we do not yet know whether the idea is wrong or the data is.

Can data settle this?

1 Does the loyalty programme increase spending
2 Is the minimum wage too high
3 Did the ad campaign pay for itself
4 Will this model still work next year
5 Should we prioritise growth over margin

Hold up 1, 2 or 3 fingers. 1 yes · 2 no · 3 only with an assumption you have to defend

Two of the five are not empirical questions

1 Does the app increase spending yes, with a design
2 Is the minimum wage too high no. “Too high” is a judgment
3 Did the ad campaign pay off only with an assumption
4 Will the model work next year only with an assumption
5 Growth over margin no. That is what you want

Data can price a trade-off. It cannot tell you the trade is worth making.

In December you show the comparison, not prove the theory

  • You are not proving your theory in the proposal.
  • You are showing which comparison your design makes.
  • And what would have to be true for that comparison to mean what you claim.

That paragraph is the identification section. It is the hardest page in the document and it is worth more than the results.

Causality

Every customer has two outcomes, and you see one of them

Every unit has two outcomes, the thing you watch: here, spending. One with the treatment, the thing you change: here, the app. And one without.

Customer On the app Not on the app
Ana $152 ?
Ben ? $98
Chen $131 ?
Dee ? $104

Holland (1986), Statistics and Causal Inference, JASA 81(396), 945–960

Causal inference is a missing data problem

Not a statistics problem.

  • No estimator recovers a number that was never observed.
  • Every method in this course is a different argument about how to fill that column in.
  • So every comparison is a stand-in for a world you cannot see. The only question is whether the stand-in was any good.

The counterfactual is those customers without the app

“App customers spend 40% more. Sign the other half up and they will too.”

  • The counterfactual is what those same customers would have spent without the app.
  • It is not what the other customers spent.
  • The other customers are the stand-in. Were they ever like the app customers?

Who signs up for a loyalty app? Say it out loud.

What is the right comparison?

A store gets a refit. Sales rise 12% the following quarter.

1 The same store, the quarter before
2 Every other store in the chain, same quarter
3 Stores that were also due a refit but have not had it yet

Hold up 1, 2 or 3 fingers. Which stand-in is best?

The best stand-in still needs an assumption

1 Same store, quarter before ignores anything else that changed
2 Every other store refitted stores were picked, not drawn
3 Stores due a refit, not yet done closest, if the queue order is arbitrary

Three is the best available and it is still an assumption. You are claiming the queue order has nothing to do with sales.

Ask two questions of every number

Compared to what?

Would those two groups have looked the same anyway?

Every number in every deck, for the rest of your working life.

A causal graph is a picture of what causes what

  1. Causal graph, or DAG. Circles are things you could measure. An arrow means “causes”. Example: heavy shopper → spend.
  2. Treatment, marked X. The thing you would change. Example: putting a customer on the app.
  3. Outcome, marked Y. The thing you would watch. Example: monthly spending.
  4. Z. A third thing that touches both. Where its arrows point is the whole question.
  5. The red arrow is the effect you want to measure. Grey arrows are everything else.

Five words for how a comparison breaks

  1. Confounder: causes both treatment and outcome. Example: heavy shoppers join the app and spend more.
  2. Collider: caused by both. Example: being hired is caused by the test and by the interview.
  3. Mediator: on the path from treatment to outcome. Example: a price cut brings people in, and they buy.
  4. Back door: a route from treatment to outcome that is not the effect. Example: app ← heavy shopper → spend.
  5. Condition on, or control for: compare within a variable’s values. Example: heavy shoppers with heavy shoppers.

A confounder causes both and opens a back door

  • A confounder causes both the treatment and the outcome. Here: heavy shoppers join the app, and heavy shoppers spend more.
  • That opens a back door. Compare heavy shoppers with heavy shoppers: that is conditioning on it, and it closes the door.
Heavy shopper Z · confounder On the app X · treatment Spend Y · outcome compare heavy shoppers with heavy shoppers

In a meeting you say: the two groups were already different before anything happened.

Selecting on the outcome is conditioning on a collider

  • Among the people we hired
  • Among our customers
  • Of the startups that made it
  • Looking at the funds that survived
  • Selecting on the outcome is conditioning on a collider.
  • It is the most common broken claim in business analytics, and it never looks broken.

In a meeting you say: you only looked at the winners, so you invented a trade-off.

Control for the middle step and the effect disappears

  • A mediator is the middle step from treatment to outcome. Here: a price cut brings people in, and they buy.
  • Control for foot traffic and the price cut has nothing left to do. The effect you came for is gone.
Foot traffic Z · mediator Price cut X · treatment Sales Y · outcome hold foot traffic fixed and the price cut has nothing left to do

In a meeting you say: you controlled away the thing you were looking for.

One rule decides what to condition on

Close the back doors. Leave the front door alone.

The back door is the route through the confounder: close it. The front door is the effect itself, through the mediator: leave it alone. And do not select your sample on the result.

Every control is a claim about the graph

  • Controlling for a variable means comparing within its values. Conditioning on it is the same act.
  • Every variable you control for is a claim about the graph. It is not a safety measure.
  • A regression with every control you could find is a graph nobody drew.

If the coefficient moves when a control goes in, that is neither good news nor bad news until you can say which of the three that variable is.

What you hear, and what it is

What you actually hear What it is
Firms that do X grow faster confounder
Students who eat breakfast score higher confounder
Among the people we hired, the good testers interviewed badly collider
Nine of the top ten founders dropped out collider, by survivorship
Once we control for X the gap disappears mediator
More police, more crime reverse causality
We sent the team to our worst stores and they improved regression to the mean
Reported cybercrime doubled the measure changed

Three are graph structures. Reverse causality is an arrow the wrong way round. The last two cannot be drawn.

Three claims, worked

Firms that hire consultants grow faster, so hire consultants

“Firms that hire consultants grow faster than firms that don’t. So hire consultants.”

A consulting firm’s own marketing deck

Hold up 1, 2, 3 or 4 fingers. 1 confounder · 2 collider · 3 mediator · 4 not a graph problem

Growing firms could afford consultants

Already growing Z · confounder Hires consultants X · treatment Growth Y · outcome compare firms that were growing at the same rate
  • Firms already growing could afford consultants. Growth caused the hiring.
  • So the comparison is between growing firms and the rest. The back door is open.

In a meeting you say: the two groups were already different before anything happened.

Good to Great picks eleven winners and asks what they share

Eleven companies beat the market by at least three times over fifteen years. The book finds what they have in common and recommends it.

The best-selling management book of its generation

Hold up 1, 2, 3 or 4 fingers. 1 confounder · 2 collider · 3 mediator · 4 not a graph problem

The sample was chosen on the outcome

Became great Z · collider The practice Everything else the sample is the eleven inside the box
  • The eleven were chosen because they became great.
  • Whatever else also caused that is now correlated with the practices, inside the sample.

Circuit City filed for bankruptcy in 2008. Fannie Mae was placed in conservatorship the same year.

HR controls for salary band and the gap disappears

“Once we control for salary band, the gender gap in promotion disappears. So there is no bias in promotion.”

An HR review, closing the matter

Hold up 1, 2, 3 or 4 fingers. 1 confounder · 2 collider · 3 mediator · 4 not a graph problem

The bias works through the salary band

Salary band Z · mediator Gender X · treatment Promotion Y · outcome the band is how the bias operates
  • If the bias works by placing women in lower bands, the band is how the effect travels.
  • Control for it and you have controlled away the thing you were looking for.

The disappearance is not the finding. It is the mechanism.

Every design ahead closes a back door

  • Each one closes a back door in a different way.
  • You will draw each of them as a graph before you hear its name.
  • Next week you generate questions. The two questions you ask of every number are the filter that kills the bad ones.

After the break, your own project

Sit at your numbered table your team is on the sheet
Ground rules, one page, signed while you still get along
Your project’s causal graph on paper, by hand
Every team reports two and a half minutes, hard stop

Both sheets go in before you leave.

Three questions, three minutes

On your own device. This one tests the plumbing and does not count.

Six across the term. Lowest one dropped. Each closes session A.