Results under scrutiny — critique cards

Card 1

Complaints per store

Outcome: number of complaints per store per month. Mostly 0, 1 or 2; a few stores hit 15. Estimated by OLS.

"Adding one staff member cuts complaints by 0.4 per month. At our smallest stores that takes us below zero."

Critique. Is the model wrong, or only the sentence?

Card 2

The 0.03% lift

2.4 million users, randomised. Effect on conversion: +0.03 percentage points, p < 0.001.

"Highly significant. Roll it out to everyone."

Critique. Significant of what?

Card 3

The pilot that "failed"

38 stores. Estimated effect +12%, 95% interval from −4% to +28%, p = 0.14.

"No significant effect. Kill the programme."

Critique. What has actually been shown?

Card 4

The elasticity

Three years of the firm's own transaction data. Its price changed twice in that period, both times by about 2%.

"We estimate demand is inelastic, so we can raise prices."

Critique. What did the standard error have to work with?

Card 5

The 112% customer

Outcome: churn, coded 0 or 1. Fitted with a linear model on tenure, spend and support tickets.

"This customer has a 112% probability of churning, so prioritise them."

Critique. Name the specific problem.

Card 6

The log coefficient

Outcome: log of units sold. Coefficient on price: −1.8, standard error 0.3.

"Raising price by one dollar costs us 1.8 units."

Critique. Read it back in the units of the problem.

Card 7

R² of 0.94

Weekly sales regressed on advertising spend, same market, three years. R² = 0.94.

"The model explains 94% of sales. Doubling ad spend will roughly double them."

Critique. Which question does this model answer?

Card 8

The 97%

A single estimate, p = 0.03.

"There is a 97% chance the effect is real, so we can act on it."

Critique. Say precisely what p = 0.03 does mean.

Critique sheet — team ______  ·  card ____

How this runs

Every team works its card at the same time. Then one team is drawn at random to defend at the front while everyone else attacks — and we draw again, and again. You will not know whether you are up, so be ready. Points for an attack that lands, and you have to name what is wrong, not just dislike it.

Name it: wrong model for the shape of the outcome · significant ≠ large · not significant ≠ no effect · not enough variation in X · predictions outside the possible range · coefficient read in the wrong units · fit is not a causal claim · that is not what a p-value says

What they are actually entitled to claim

Where exactly the sentence goes wrong

What would you need to see before you believed it?

Is our own project at risk of this same error? How would we know?

From the floor: the objection that landed