Syllabus

Draft. Settled in structure, not in every detail. Meeting dates follow the Cornell Fall 2026 academic calendar. Anything marked “to be announced” is filled in here as it is confirmed, and each meeting page carries the current version for that day.

Course

AEM 6991 — MPS Capstone Project I

Fall 2026. Wednesdays, 2:00–4:30 pm. First meeting: August 26, 2026. Last meeting: December 2, 2026. Location: Riley-Robb Hall B15. Section: LEC 004 (to be confirmed).

Instructor. Prof. Ariel Ortiz-Bobea. Office: 450B Warren Hall. Email: . Office phone: (607) 255-0220.

Teaching assistant. To be announced; posted here and on Canvas once assigned.

Office hours. To be announced in week 1 and posted on Canvas; also by appointment.

Credit hours. 3. Enrollment limited to AEM MPS students.

Course description

AEM 6991 is the first half of the MPS capstone sequence. Each team moves from a vague interest to a research plan that can be executed in AEM 6992 in the spring. That plan is the deliverable.

The course runs in the order the work actually happens.

Part What it does
Getting started Who is in the room, and what everyone is interested in
Foundations Causal language, where questions come from, research ethics and access, and the tools
Methods & designs What models estimate, how estimates fail, and the designs that survive
Communication Writing, figures, and speaking — each with session B on your own project
Delivery The midterm, and the pitch

Three commitments shape it.

The question comes before the data. Teams who find data first rationalise a question around whatever they found. That is the most common way a capstone goes wrong. The conceptual half is taught on examples from industry and the news, before anyone opens a dataset.

Finding the question is a skill, not a preamble. Most of your education has rewarded execution. AEM 6992 is where the executing happens. A bad question executed beautifully is still a bad project.

Writing and teamwork run through the term, not at the end. Writing, figures and speaking each get a meeting, with an individual piece attached. Each team writes its own ground rules in meeting 2, and every session B opens with a check-in that is written down and graded.

No prior econometrics or programming is assumed.

Audience and prerequisites

Open to, and required of, AEM MPS students.

There are no formal prerequisites. No econometrics, no prior programming, and none is taught. Where the course touches data directly, it does so through an AI assistant. No student is blocked by tooling.

If you arrive with statistics or coding experience, the data work will be easy. The design vocabulary and the verification habit will be new regardless of what you bring.

Learning outcomes

The registrar’s outcomes for AEM 6991 are unchanged. By the end of the semester students will be able to:

  1. Teamwork. Work effectively in small groups, complete tasks on time, and allocate work equitably among team members.
  2. Problem focus. Select and define a project problem that demonstrates relevant knowledge and articulates the significance of the problem.
  3. Project design and workplan. Develop a methodologically rigorous research design and identify key project tasks and outputs.
  4. Presentation of progress. Write about, present, and discuss the project in class with analytical rigor.

This section adds two outcomes that serve the four above:

  1. Evidence sourcing. Establish that the data a project needs exists, can be obtained on the project’s timeline, and measures what the research question requires.
  2. Design vocabulary. State what identification strategy a project relies on, what it assumes, and what it therefore can and cannot claim.

AI policy

Generative AI tools are permitted by default for all coursework. Students are expected to use them, and to use them well. Three rules apply to every deliverable:

  1. Disclose what you used. Every submitted deliverable carries a short disclosure block: which tools, on which parts of the work. The form is brief; what matters is that it exists.
  2. Verify what the AI produced. AI output is your work, exactly as code from a colleague would be. Two failure modes are graded against explicitly, because they are the ones that destroy a research proposal:
    • Fabricated citations. Models invent references, and they do it convincingly. Every reference in every submission must be one you have opened. A citation to a paper that does not exist is a fabrication, not a typo.
    • Plausible-but-wrong design advice. Models will confidently recommend an estimator whose assumptions your data violates. The design you propose is yours.
  3. Do not paste restricted data into a public chatbot. Confidential, IRB-restricted, and proprietary data stay on your machine.

Failure to disclose is an academic integrity violation under the Cornell Code. Failure to verify is graded against the rubric like any other error.

Note. Other sections of AEM 6991 restrict generative AI. This one takes the opposite position, deliberately.

The deliverable here is a research proposal. That is exactly the genre where unverified AI output does the most damage. So the skill worth teaching is verification, not abstinence.

Assessment

Half the grade is individual and half is team-based, so that no student can be carried and no student is sunk by a team.

Component Weight Type
In-class quizzes 15 % Individual
Midterm 15 % Individual
Participation 20 % Individual
Individual subtotal 50 %
Project pitch — 1 page, presented 5 % Team
Proposal outline 15 % Team
Research proposal 30 % Team
Team subtotal 50 %

Grading scale: A+ (≥ 98), A (≥ 92), A- (≥ 88), B+ (≥ 85), B (≥ 80), B- (≥ 75), C+ (≥ 70), C (≥ 65), C- (≥ 60), D (≥ 50), F (< 50).

In-class quizzes (15 %)

Short quizzes, in class, without warning. Six across the semester. Three questions, three minutes, on your own device. Each one closes session A.

They are drawn from the lecture you are sitting in, not from a reading. Most are scenario questions, not definitions. Here is a situation — which design fits, and what would have to be true? Here is a design — what would break it? Here is a step an AI just did for you — how would you check it without reading the code?

They are unannounced on purpose. A quiz you can revise for is a reading check. A quiz you cannot is a reason to be in the room and paying attention.

The lowest score is dropped, and there are no make-ups.

Midterm (15 %)

In class on Wednesday, November 18. Open note, short answers. It is the whole meeting — there is no session B that week. The date is fixed and will not move.

You are given a situation and asked to design the study that would answer it. The question. The design, and the assumption carrying it. What data it requires. And what would undermine it.

It tests whether you can apply the design vocabulary to a situation you have not seen. There is nothing to revise beyond having turned up all term.

Participation (20 %)

You start with full marks and lose them. What costs points:

Research interest survey not submitted, or late −5
Session-B check-in not filed −2 each
Peer evaluation not submitted on time −3 each
Arriving late, or leaving at the break −2 each
Absent without notice −4 each

The research interest survey is due Tuesday September 1 at 11:59 pm. It is on Canvas. It should take under an hour, and most answers are two or three sentences.

This replaces the one-page memo I described in meeting 1. That memo asked you to write about feasibility. You have two days, and no time to find out whether your data exist. I would rather have the question and your reasoning, so an idea with no data plan is completely fine at this stage.

Two of the questions are how I match teams. What you bring is capability: a prior job, an industry, a language, access to somebody. Without it I can only match on topic. Where this fits your plans is motivation, and it is the better predictor of who still cares in November. Write both for me rather than for a grader.

The checkbox questions put complementary skills on the same team. Answer them honestly. Say you have not used a method rather than that you have read about it.

I form the teams from these, and most of these ideas will not survive that. That is expected, and it is not a verdict on yours.

The session-B check-in is ten minutes at the start of every second half: what your team did since last week, what is stuck, and who does what next. Written down and handed in. It is the record that makes an uneven split of the work visible in week 4 rather than week 13.

Peer evaluation runs twice, at meetings 6 and 12.

Team deliverables (50 %)

One paper, reached in three milestones. A one-page pitch, then the whole argument as a detailed outline, then the written proposal. The argument is fixed before the sentences are, so a broken argument is caught at the outline where fixing it costs an afternoon rather than a week.

# Deliverable Stage What it is Weight Due
1 Project pitch One page, and five minutes in front of the room. The question, why it matters, which archetype it is, and what data would answer it. This is the first time the project has to survive an audience, and the page is what you hand in. 5 % Meeting 9 — Oct 21 (week 9), in class*
2 Proposal outline The whole proposal as a detailed outline. One topic sentence per paragraph, in order, with the evidence noted underneath — the citation, the number, the figure you will point at. It has to cover the question, the identification strategy and the assumption carrying it, the unit and sample, the data and exactly how you get it, the threats to validity, and the work plan for the spring. Exchanged with a partner team at meeting 12, who referee it. If the argument does not work here, it will not work at 5,000 words. 15 % Meeting 12 — Nov 11 (week 12)
3 Research proposal ~5 pages, written out from the outline you workshopped. The plan somebody could pick up and execute in January. 30 % Dec 11 (week 16)

The data section of the outline is the one that kills projects. The most common way an MPS capstone fails is not a bad question. It is a good question whose data never arrives. So the outline makes you name the dataset, its holder, the access route, the timeline and your plan B.

How the proposal gets built

You do not write the proposal in November. You accumulate it. Session A introduces a concept; session B makes you use it, first on a shared case and then on your own project. That second half is what you hand in before you leave, and it is a piece of your proposal.

Meeting Session B leaves you with
1  Introductions Your team’s question sheet: candidate questions, each run through the four tests of a good idea, with the ones that fail killed on the record and the reason named
2  Theory and causality Your team’s ground rules, and the first causal graph of your own question — treatment, outcome, the back door you must close, and the variable you must not touch
3  Finding the question Your team’s question sheet, three candidates each labelled with its route, its shape and who acts or pays, run through the four filters with one killed on the record, and the resource map of the survivor with the row that stops them dated and the question it becomes if that row does not close
4  Ethics and access Your project’s ethics-and-access checklist — every gate it trips, who signs, and the date each one has to start
5  Coding tools and AI Your candidate dataset interrogated — or, if you do not hold one yet, the named access route and plan B for getting it
6  Models and uncertainty The overclaim your own project is most at risk of making
7  Model failures The one threat your design must survive — its direction, and whether controls can close it
8  Research designs I Whether an experiment, instrument or discontinuity is open to you
9  Research designs II Your design defended in public, plus the attack that landed — the identification section and the limitation
10  Writing Your problem statement, as an outline somebody outside your team could follow
11  Figures Your project’s key figure, drawn twice — as you expect it to look, and as it looks if you are wrong
12  Speaking Your final-presentation deck, rehearsed and attacked — plus the gap between the claim you made and the claim the room heard
14  The pitch Your plan, pitched — and what your team owes AEM 6992 over the break

What you are graded against

Across AEM 6991 sections, the criteria for written team work are the required section headings. If a heading has nothing under it, that is the finding.

The outline and the proposal must both cover, under these headings: research problem · practical justification · what is already known · conceptual model · methodological approach · study sample · data and access · internal validity · external validity · limitations and self-critique · work plan · research ethics · references.

Excellent Every heading is answered from your own project. The claim is in the first sentence. The evidence supports it. Somebody outside your team could act on it.
Adequate The headings are answered, but some are generic — true of any project, not this one.
Weak Substantial sections restate the prompt or borrow text. The argument does not cohere.
Unacceptable Not submitted, submitted beyond the late policy, or so thin it shows no thought.

Two questions move a grade more than anything else. Is the claim in the first sentence? And could it come out the other way?

Late policy. Team deliverables lose 10 % up to 8 hours late, 50 % from 8 to 72 hours, and are not accepted after 72 hours. Notify me before the deadline if something has gone wrong.

Course schedule

Fourteen meetings. Each runs the same shape:

Time
2:05–3:15 Session A — the concept, on worked examples
3:15–3:30 Break
3:30–4:30 Session B — your own project, submitted before you leave

A quiz closes session A, without warning.

Session B is not a lab exercise on someone else’s data. It applies the day’s work to the team’s actual project: conceptual clinics early in the term — draw your graph, name your design, work out what data it needs — and hands-on work once R arrives. The output is submitted the same day. Fourteen of these, accumulated, are most of the research design paper.

Every meeting has its own page, reachable from the sidebar, carrying the readings, slides, and materials for that day. The table below, the front page, and the sidebar are all generated from one file, so they cannot disagree.

Part 1 — Getting started

Who is in the room, how the course works, and what you are here to build. You leave the first meeting knowing what everyone else is interested in, which is how the teams get formed.

# Date Meeting Due
1 Aug 26 Introductions
Say what you are interested in, in one line somebody else can repeat back to you. Hear the obstacles in it from people who have just met it. Know how this course works and what is due when.
A · 2:05 Three blocks, and the first is the one the rest hangs off. First, what makes a good idea. You already have ideas, so the useful thing is not another way to generate them. It is a way to sort them quickly, before you spend a month on one that was never going to work. Four tests, and a good idea passes all four. Important: somebody is better off knowing the answer, and importance is the size of the decision that changes rather than the size of the topic. Feasible: you can get the data, in the time you have, with the skills on the team. Most projects die here, and they die in March because nobody checked in September. Novel: nobody has already answered it, and you know because you looked. Fit: your team can do the work and will still care in March. The four pull against each other, so a project is a choice about where to sit between them. Value gets its own treatment, because it decides who your reader is. Public value is shared and gets published. Private value can be kept and somebody pays for it. Both are fine here. Private value also decays: McLean and Pontiff followed 97 published rules for beating the market and found they earned 58 percent less once in print. Second, the shapes a project can take. Seven of them, each with the question it asks and an example, because naming yours tells you what you have to go and get. Third, where ideas come from. From the world, a problem you saw. From previous work, a marginal change to something published. From data, a file you can reach worked backwards to a question. Each route leaves something different missing. The first lacks data, the second lacks novelty, the third lacks a decision. And working backwards from data is honest in exactly one form, which meeting 1 gave you.
B · 3:30 Ten-minute team check-in first, in writing, on the sheet on your table. Who is here. What you did since last week. What is stuck. Who does what next, by name and by date. In week three, nothing is stuck yet is an acceptable answer. Then your own project, which is the half you hand in. You write candidate questions and put each one through the four tests, and you kill the ones that fail. The tests you cannot run on your own work are run by another team, because every team thinks its own question matters and a stranger is the only honest judge of that. Your sheet goes in before you leave, and it is what your pitch gets built from. A question you kill in September is a month you do not lose in March.
Research interest survey, on Canvas before meeting 2

Part 2 — Foundations

What a causal claim is, where good questions come from, what you are allowed to do with other people’s data, and the tools you will work through. Teams form here. So does the question.

# Date Meeting Due
2 Sep 2 Theory and causality
Draw the causal graph behind a claim somebody is acting on, name the failure in it, and say which variable you must condition on and which you must not touch. Leave with your team formed, its ground rules in writing, and a first graph of your own project that the room has heard.
A · 2:05 Three blocks. First, theory. A theory is a set of assumptions that produces a prediction you could be wrong about. That is all it is, and it buys you four things. It tells you which comparison to make. It tells you what sign to expect. It tells you what to hold fixed. And it tells you what result would embarrass you. It is also the only thing that lets a number travel. A 40% lift measured in New York says nothing about Chicago unless you know the mechanism that produced it. Two worked examples, both from markets you will work in. Akerlof’s market for lemons predicts that good used cars stay off the market, and it points straight at the fixes — warranties, certification, return policies. Then a dispute. Spence’s signalling model and human capital both predict that a degree raises wages. They disagree about why. That disagreement tells you which comparison settles it: the jump at graduation, not the slope across years of study. Second, what empirical work can and cannot settle. It can measure a magnitude. It can rule out a sign. It can decide between two theories that predict different things about the same comparison. It cannot test a theory on its own. Every test is a joint test of the theory, the measure, the sample and the model, so a rejection never tells you which of the four broke. And no amount of data tells you what to want. Third, causality, which is the rest of the term. Counterfactuals first. Every unit has two outcomes: the one with the treatment and the one without. You only ever see one of them. That is Holland’s fundamental problem, and it is a missing-data problem rather than a statistics problem. So every comparison you make is a stand-in for a world that did not happen, and the only question is whether the stand-in is any good. Customers on the loyalty app spend 40% more. The counterfactual is what those same customers would have spent without the app. It is not what the other customers spent. Then the picture. Causal graphs are arrows between variables, and they let you argue about identification before you own any data. A confounder causes both the treatment and the outcome, opens a back door, and has to be closed. A collider is caused by both, and conditioning on it manufactures an association that was never there. A mediator sits on the path you are trying to measure, and conditioning on it deletes the effect you came for. Close the back doors. Leave the front door alone. That is also why controlling for everything makes things worse. Every extra control is a claim about the graph, not a safety measure, and a kitchen-sink regression is a graph nobody drew. Then the cards, because seventy minutes of me is too many. Eight cards, each carrying a real causal claim somebody is acting on: a board paper, a consulting deck, an HR review closing a matter, a widely shared post. Every card is broken in a different way. In threes, draw the graph behind the claim, name the failure in this morning’s vocabulary, then say what you would condition on and what you must leave alone. Card 3 is the hiring puzzle and it is a collider, because among the people you hired a strong test score and a strong interview trade off, since either one was enough to get in. Card 5 is the promotion gap and salary band is a mediator, so control for it and you have controlled away the thing you were looking for. Those two are where the morning either landed or it did not. Finish with the map. Every design in this course closes a back door in a different way. You will draw each one before you hear its name.
B · 3:30 Your team is already yours. I email the teams on Tuesday night and there is a numbered sheet on each table, so the first three minutes are sitting down rather than listening to a list. Then ground rules, one page, signed by everyone, written while the team still gets along. Who does what. What happens when somebody misses a deadline. How you settle a disagreement, and who breaks a tie. Then your own project, on paper, by hand. Your team has a shared interest and not yet a question, so take the closest thing you have and draw it: the treatment, the outcome, the one back door you would have to close, and the one variable you must not touch. Then every team reports. Two and a half minutes each, hard stop, on a timer. Your question, and the graph you drew. Nine teams means you hear eight other projects before you leave, and this is the only afternoon all term when the whole room sees what everybody chose. Your sheet is disposable and it will be wrong, which is the point of drawing it in September. You redraw it when we reach research designs, and the version that survives the room is the figure in your identification section. Both sheets go in before you leave.
Your team’s written ground rules — agreed and handed in today
3 Sep 9 Finding the question
Say where your question came from, who would act on the answer or pay for it, and which of seven shapes it takes. Map what it needs against what your team has, name the row that binds and the date it has to close by, and say what the question becomes if it does not.
A · 2:05 Five blocks, and the fifth is the one to keep. First, the last two cohorts, on one slide. Thirteen capstone titles from 2023 and 2026, and sixty seconds on paper to say what most of them have in common. Four of the five 2026 projects were the same project. Take a public signal anybody can download, regress a financial outcome on it, report a coefficient. Free data, no design, and nobody who would act on the number. I let you find that rather than saying it, because it lands harder when you have written it down than when I say it. Second, where ideas come from. Finding the question is the job, not the preamble to it. Most of your education has rewarded execution. This course rewards the part before it. In research, having the idea counts as much as answering it. In business the same skill has a different name: a proposition that has value. You cannot see the value of a new answer without knowing what is already known, and that is what reading is for. Ideas arrive by three routes. From the world: a problem you saw, or a decision somebody has to make. From a study: a result you would extend, contest or move to a new setting. From data: a dataset you hold, and a question worked backwards onto it. Each route leaves something different missing. The first lacks data, the second lacks novelty, the third lacks a decision. The data route is honest in one form only, and meeting 1 gave it to you. Third, value. An answer has value when somebody is stuck without it. Public value is shared. Many people are better off, nobody can be excluded from the answer, so it is hard to charge the people who benefit and it gets published. Private value can be kept. One decider’s choice changes, they can keep the answer, so somebody pays for it. Both are fine in this course. Private is easier to monetise. Public is easier to publish. And private value evaporates when it is published. McLean and Pontiff followed 97 published return predictors and found their returns fell by more than half after publication. The test on your slate is naming the role that acts or the firm that pays. Fourth, the shapes and the filter. Five archetypes: evaluation, preference and pricing, prediction and risk, diagnosis and benchmarking, feasibility and landscape. One table: what each asks, what it can claim, and how it dies. The pitch asks for the shape by name. Then six ways a project dies, four of which are the ones a room names before anybody teaches them. Then the filter, which is four questions. Could a team of three to five execute this in one spring semester? Can you name the dataset, who holds it, and how you get it, checked rather than assumed? Could the answer come out either way? Would anybody act differently depending on it? A no on any one is fatal, and it is cheap to discover today. The third question was last week, and last week’s two questions are not a fifth filter: they are how the third one gets teeth, and they are the comparison row of the map. The fourth question is the value block. The first two are the fifth block. And you write three candidates before you filter one, because a filter needs something to choose between. Fifth, the resource map. Start from what you have, not from the goal. That is effectuation, the theory paper behind Sarasvathy’s think-aloud study of 27 expert entrepreneurs. It is a small study, so take it as a habit that works rather than a law, and take the inventory it implies. The map is six rows against what the question needs: data, comparison, skills, time and money, permissions, people and access. Four columns: what the question needs, what you have and how you checked it, the gap, and who closes it by when or how the question moves. A have with no check counts as a need. The permissions row takes a rung from the access ladder and nothing more, because meeting 4 teaches the gates. Two live cases, built row by row on the screen after you have tried the first on paper. Does a brand’s livestream lift its sales that week? The missing row is the stream schedule, and the question moves three times before the plan fails once. Does a plain-language nutrition label change what shoppers put in the basket? Nobody holds that data, so the rows that bind are money and the clock. One case binds on a file that does not exist and the other on a clock, and nothing but the map tells you which one you are in. The map is redrawn every time a row changes. That is not the plan failing. That is the plan working.
B · 3:30 Check in first, in writing, nine minutes, on the sheet already in your envelope. Who is here. What you did since last week. What is stuck. Who does what next, by name and by date. In week three, nothing is stuck yet is an acceptable answer. Then the shared case. Every team maps the same question from the same two dataset cards: is flood risk priced into house prices, and did a local flood change that? Six rows, all of them, with how you checked each thing you claim to have. Then the two lines at the bottom: the row that binds and its date, and what the question becomes if it does not close. The same object for nine teams, so the maps are comparable. Three things about these datasets are easy to miss, and the draw is where you find out whether you missed them. One team is drawn at random to read its map aloud, blanks included, while the room attacks. Naming a specific row that is blank, unchecked or wrong is an attack. Calling the whole map thin is not. A point for every attack that lands, and I close with whichever trap nobody found. Then your own project, which is the half you hand in. Three candidate questions, and label the route of each. Each with its archetype and the role of the person who would act on the answer or pay for it. One of them run through all four filters in writing, not in your heads. A second crossed out with the filter that killed it named, because a team that kills nothing has not used the filter. For the third, the one row that would kill it. Then a resource map for the strongest survivor, with every gap owned by somebody and dated, or the question moved. The last lines on the sheet are the ones I read first: the row that binds and the date it has to close by, and what the question becomes if it does not. A second team defends its own map if the clock allows. Four sheets go in your envelope before you leave: the check-in, the shared map, the slate and your own map. The slate and the map are what your pitch gets built from. A question you kill in September is a month you do not lose in March.
4 Sep 16 Ethics and access
Say whether your project needs IRB review, a data-use agreement, both or neither — and defend the answer. Name the date each gate has to start, counted backwards from January, so the spring is not spent waiting.
A · 2:05 Five blocks, and the last one is the point. First, why the rules exist — briefly, and without the sermon. Every rule in research ethics was written after a specific failure. Nuremberg followed the doctors’ trial. The Belmont Report followed Tuskegee, which ran for forty years and was stopped by a reporter rather than by a scientist. Its three principles — respect for persons, beneficence, justice — are still the frame every IRB applies. But the modern failures are the ones that concern this room, and none of them involved a needle. Facebook ran a mood experiment on 689,000 users and published it. Cambridge Analytica harvested a personality quiz’s friend graph. Second, the IRB. Two questions decide whether it applies. Is this research? Are there human participants? Teams miss the second half, because identifiable private information counts even when you never meet anybody. Scraped posts count when individuals are identifiable. So do purchased consumer records. Then the three levels of review, described in weeks rather than in definitions. Exempt can clear in a fortnight if nothing is wrong with the submission. Expedited runs to a month or more. The full board meets on a schedule you do not control, so its clock is a calendar, not a queue. And the rule that catches somebody every year: you do not decide that you are exempt. The IRB issues that determination, and it issues nothing until every person named on the protocol has finished CITI human-participants training. Exempt protocols included. One teammate who skipped the modules holds up the whole team, so do them this week, before you know whether you need them. Third, the other gate. A data-use agreement is a contract, not an ethics review, and the university signs it rather than you. Expect legal review on both sides. Expect months. Expect restrictions on what you are allowed to publish, which can mean you cannot show your own result. Terms of service are the same species — a contract you accepted by clicking. Scraping is rarely a crime and frequently a breach. Prefer the download, then the API, then nothing. Fourth, whether you may hold the file at all. Direct identifiers are the easy part. Latanya Sweeney showed that ZIP code, birth date and sex together identify most Americans uniquely. So “we removed the names” is not de-identification, and a vendor’s word for it is not a determination. Then conflicts of interest, which are disclosed rather than confessed — a disclosure tells the reader how to weight your claim, nothing more. The clause to settle before anyone signs anything: a sponsor may have the right to review, never the right to veto. Then credit inside the team, agreed now, while there is still nothing to fight over. Fifth, integrity. Fabrication, falsification and plagiarism, one line each, with a case attached to each. Diederik Stapel invented his data and lost 58 papers. Reinhart and Rogoff made a spreadsheet error nobody checked and it moved fiscal policy in several countries. And the one that is live in this room: a fabricated citation is a fabrication. A New York lawyer was sanctioned $5,000 in 2023 for filing six cases an assistant had invented. “The AI generated it” was his defence, and it did not work for him either. Finally the calendar, run backwards from January. That is the argument for holding this meeting in September.
B · 3:30 Ten minutes of team check-in, then the cards. Each team draws a project somebody actually proposed, and every one is blocked in a different place. A team scraping Glassdoor reviews with the usernames attached. A teammate’s employer offering internal HR records, with a clause letting the firm approve the write-up. Interviews with farmers in the Dominican Republic, conducted in Spanish, by a team that has not done CITI. A vendor file sold as de-identified that carries ZIP, birth year and sex. A licensed WRDS extract about to be uploaded to a chatbot for cleaning. A survey of your own classmates about salary expectations and visa status. Three specifications run and only the one that worked reported. For each card, three answers. Which gate does this trip — IRB, DUA, terms of service, conflict of interest, integrity, or none? What is the earliest date it has to start? And the half that carries the marks: what is the version of this project that is legal, is honest, and is still worth doing? Then one team is drawn at random to defend at the front while the room attacks, and we draw again as the clock allows. Nobody knows in advance whether they are up. Two cards trip nothing at all, and the team holding one has to be willing to say so. Then your own project. You fill in the ethics-and-access checklist for it, on the template, and you hand it in signed before you leave. Human participants, yes or no, with the sentence that justifies the answer. Every data source you named, who holds it, and what its licence or terms permit. Whether any field is identifiable, what you do about it, and where the file will live. Any conflict, one sentence each. Who does what, by name. And the date the slowest gate has to start. Most teams will find they need none of it. The teams that do need something find out today, which is the whole reason this meeting is in September.
5 Sep 23* Coding tools and AI
Choose the right AI modality for a data job, and say why. Interrogate a file you did not create — through an assistant, without reading the code — and catch what is wrong with it before it reaches your proposal.
A · 2:05 Four blocks. First, code. Not how to write it — what it is, and why you cannot avoid understanding it. Why there are several languages: R grew out of statistics, Python out of general programming, Stata out of applied economics. SQL is not the same kind of thing at all. It is a way of asking questions of data you never load onto your own machine. Then the argument that matters here. Every serious analysis is code, whether or not you wrote it. A spreadsheet stops being enough at a moment you can name: when you join two files, when the same cleaning has to run again on next month’s extract, or when somebody asks what changed between version 4 and version 7. Two cases make that better than any principle. Reinhart and Rogoff’s growth result averaged fifteen countries where the formula should have covered twenty, and it had already been quoted in austerity debates before anyone opened the sheet. Public Health England lost 15,841 COVID cases in 2020 because an old Excel format stops at 65,536 rows and says nothing at all when it runs out. Neither is a statistics error. Both were invisible in a spreadsheet and would have been visible in code. You will drive code through an assistant. But you cannot check what you cannot read at all. Second, the four things people mean by “use AI”. They fail differently, and that is the whole point. Chat answers from memory. It invents citations that look real — right journal, plausible authors, a DOI that resolves to something else. File upload with code execution is different in kind: the assistant writes real code, runs it on the file you gave it, and shows you both. That is the only modality where the output is checkable, and it is the one this course uses. Retrieval answers from documents you supplied, which bounds the invention without removing it. Agents act over many steps without stopping to ask. Pick the wrong column at step two and by step nine it is buried under eight correct-looking operations. So you choose the modality by whether you will be able to tell it was wrong. Third, the skill this course actually grades: judging an answer without reading the code that produced it. Five checks, and they take four minutes. Row counts across a join — a join that grew has a duplicate key, and a join that shrank quietly dropped the rows that did not match, which is how a result ends up being about a different population than you think. Totals against a published figure. Five raw records read by hand, with your eyes, before any summary. Orders of magnitude: 58 and 58,000 are different claims about household income, and the gap is usually a units column nobody mentioned. And the one nobody runs — ask for the same number a second way and see whether it comes back the same. Then what you must not upload. Restricted, IRB-covered and DUA-covered data do not go into a public chatbot. That is meeting 4, and this is the week it stops being abstract. Fourth, getting hold of data at all. Start with what Cornell has already paid for — Dewey, WRDS, Capital IQ, IBISWorld — and the free public sources that need no login: FRED, IPUMS, EDGAR, data.census.gov. Then the three routes to a file. Downloading is boring and usually fastest. An API is a documented door: a key, a rate limit, stable field names, and a record of what you asked for. Scraping is the last resort dressed up as the clever one, and I will break one in front of you. The page changes and the code returns nothing. The rate limit bites. Records go missing without a single error message. The terms of service say no. And nothing you pulled could be reproduced by anybody next week. Going to a county assessor yourself is worth it only when you need one county in unusual depth, or a field the national vendor does not carry. Last, reproducibility, at the level of awareness rather than tooling. Keep the raw file untouched. Keep one path from raw to result. Keep a note saying where the raw came from and on what date. The test is simple: could a teammate get your number on a different laptop next week? If not, you do not have a result. You have a screenshot.
B · 3:30 Ten-minute check-in first. Then everybody works the same file. store-weeks.csv is a retail extract with four things planted in it. A block of store-weeks is duplicated, so every total is too big. Twelve prices are recorded in cents while the rest are in dollars. Missing values are coded −99, which a mean will swallow without complaint. And one store stops reporting in July, so a year-on-year comparison silently compares nine months against twelve. You find them by driving an assistant with file upload. You do not find them by reading the file. Twenty minutes. Then the half that is actually graded, which is the second question on the sheet: which of the five checks would have caught this one? A team that finds a fault and cannot name the check has not learned the thing. Then I put up a transcript of an assistant getting this same dataset confidently wrong, and the room says where it went off the rails. Second half, your own data. Take the dataset your team is currently betting on and run the same interrogation. What is one row. How many rows are there. What is missing, and how is it coded. What changed over time. And one number checked against a source outside the file. That page is your data interrogation sheet, and it is handed in before you leave, with three lines at the bottom: what you could not answer today, who chases it, and by when. If your team has no file yet, your sheet is the shortest named route to one — which source, which access route, how many days or months, and what plan B is. Not having the data in September is fine. Not knowing how you get it is not.

Part 3 — Research methods and designs

What the standard models estimate and where uncertainty comes from. Then how estimates go wrong and what fixes them. Then the designs applied economists use to get a credible answer — what variation each exploits, what each assumes, and how each fails.

# Date Meeting Due
6 Sep 30 Models and uncertainty
Read a coefficient in the units of the problem. Say what a p-value does and does not tell you. Tell an explanatory question apart from a predictive one.
A · 2:05 Linear regression, non-linear models, count models. Described, not derived. What each estimates, and how to read a coefficient in the units of the problem. Then uncertainty. Residuals and errors, sampling variation, what a standard error actually states, and where p-values come from. More usefully: what they are not. Finally the distinction that decides half the projects in this room. Explaining is not predicting. A model can predict well and still tell you nothing about what to do.
B · 3:30 Ten minutes of simulation first, driven from the front. Draw from a world you control. Watch the p-value distribution under the null. Move the sample size and the variation, and watch it shift. Then fit the same world in and out of sample, and watch an excellent-looking model fail on data it has not seen. Nothing to install. Then the cards. Each team gets an analyst who ran something and said something about it, and every one overclaims differently. A count outcome fitted by OLS and predicted below zero. A 0.03% lift called highly significant across two million users. A wide interval called no effect. An elasticity estimated off two price changes in three years. A churn probability of 112%. Work out what they are actually entitled to claim. Then a team is drawn at random to defend that while the room pushes back. The last line on the sheet is the one that matters: is your own project at risk of the same error?
Rate your teammates
7 Oct 7 Model failures
Take a number apart and name the failure that produced it. Say which way it pushes the estimate, and whether more data would fix it. Then say what your own controls would have to close, and what they cannot.
A · 2:05 Two families of failure, and they are not the same problem. Bias moves the number. False precision leaves the number where it is and shrinks the error bar around it. Bias first, and there are four. Selection and sorting: the retention email went to the customers already most likely to renew, so the lift is who was picked, not what was sent. Omitted variables, with a rule for the direction — the omitted thing’s effect on the outcome, times its correlation with the treatment. Two signs multiplied. You can usually work out which way you are wrong before you have the data. Measurement error next. Classical error in the variable you care about pulls the estimate toward zero, always: self-reported spend, survey income, recalled hours. Non-classical error does not. People round, they under-report, and top-coded income errs one way only. Then it can go anywhere. Fourth, endogeneity and reverse causality. More police, more crime. The firm raised price in the quarter demand was strongest. And then the question I will ask of every one of them: does more data fix it? No. A biased estimate with a million rows is a more precisely wrong number. Then the quieter family, which leaves the estimate untouched. Dependence ignored: five hundred shoppers in ten stores is closer to ten observations than five hundred, and Bertrand, Duflo and Mullainathan found placebo policies in state panels rejecting at 45% instead of 5%. Multiple testing: twenty subgroups, one significant, one slide. Specification search: Silberzahn gave 29 teams one dataset and one question, and got estimates from 0.89 to 2.93. None of the three moves the number. All three make you sure of it. Last, the fixes and their ceiling. Conditioning on observables, drawn on the graph — which back doors a set of controls closes, and the two controls that make things worse. One sits on the path from cause to effect. One opens a path by being conditioned on. Then matching, which is the same claim made by comparison rather than by regression, and it needs somebody to compare to. If no untreated store looks like your treated stores, there is nothing there. Then the ceiling. LaLonde put observational estimates with controls up against a randomised benchmark, and they did not reproduce it. ‘I added controls’ is a claim about the graph, not a technique. So state the claim out loud, and report how far your coefficient moved when the controls went in.
B · 3:30 Ten minutes of team check-in, then the results-critique cards again — the same eight analysts, dealt fresh, and this time each one has come back with a fix. They added controls. Fifteen minutes to rule on it. Name the failure in the morning’s vocabulary. Say which way it pushes the estimate. Say whether more data would fix it. Then rule on the fix itself: does the control close a back door, sit on the path from cause to effect, or open a door by being conditioned on? Then one team is drawn at random to defend at the front while the room attacks, and we draw again as far as the clock allows. Nobody knows in advance whether they are up. Points for attacks that land, and you have to name what is wrong, not just dislike it. Card 7 is the one to watch. The advertising model with an R² of 0.94, now with month and competitor spend added, and still uninterpretable — the budget was set from the sales forecast, so the arrow runs backwards, and no control fixes an arrow that runs backwards. Then your own project. One threat, the one your design actually has to survive. Write down which way it pushes your estimate, whether more data would fix it, the variable that would close it, and whether you can obtain that variable before January. If you cannot, write the sentence you will put in your limitations section instead. One page, per team, handed in before you leave. You will be answering that threat in November.
8 Oct 14 Research designs I
Draw the graph for an experiment, an instrument and a discontinuity. Say whether any of the three is open to your question.
A · 2:05 Three designs, each drawn before it is described. Experiments first. Randomisation cuts every arrow into treatment, so there is no back door left to close. Then the decisions that actually make one work. What is the unit you randomise — the person, the store, the week? How many arms can you afford before each one is too thin? Why you check balance afterwards, and what a failure would mean. And pre-registration, the cheapest honesty device in research: say what you will test before you look, and specification search stops being available to you. Instrumental variables next. An arrow into treatment, and none into the outcome except through it. The exclusion restriction is that missing arrow, which is why it cannot be tested. Then regression discontinuity: nothing else jumps at the cutoff. For each, what variation it exploits, what it assumes, how it fails, and what it looks like done well.
B · 3:30 Each team draws a setting nobody in the room owns. A bank with an email list. A scholarship cutoff. An NGO that gave filters to whoever asked. Fifteen minutes to draw the graph, name a design, and state the assumption carrying it. Then teams are drawn at random to defend for four minutes while the room attacks. What breaks this? Nobody knows in advance whether they are up. One card cannot be answered honestly at all, and the team holding it has to be willing to say so. Last: is an experiment, an instrument or a discontinuity actually open to your own project? Most teams find the answer is no. Finding that out today rather than in November is the point.
9 Oct 21* Research designs II
Draw your project’s design as a graph. Name the single assumption carrying the estimate — the one you would have to defend to a sceptic.
A · 2:05 The same discipline on harder designs, each drawn first. Difference-in-differences: the confounders you cannot measure are the ones that do not move. Parallel trends is what that assumption looks like drawn. Panel data with local fixed effects: which arrows the fixed effects erase, and which survive untouched. Event studies: the same picture with time on it. For each, what identifying variation is left once the fixed effects are in, and what assumption is carrying the estimate.
B · 3:30 Same format, higher stakes. First a card apiece — a staggered rollout, a state border, a fee that changed for accounts opened after a date. Fifteen minutes to draw it and name the design. Then the part that counts. Teams are drawn at random to put up their own project’s design, draw the graph, state the assumption carrying the estimate, and defend it while the room tries to break it. Everyone presents eventually. Nobody knows when. Whatever survives is the figure that goes in your identification section, so draw it to keep. And write down the attack that landed. That is your limitations paragraph.
Project pitch — one page, presented in class

Part 4 — Communicating the work

Half of what you are graded on is whether somebody else can follow the argument. So it is taught, not assumed. One meeting on writing, one on figures, one on standing up and saying it. In all three, session B works on your own project.

# Date Meeting Due
10 Oct 28 Writing
Reconstruct the outline behind somebody else’s paper. Build your own problem statement the same way, topic sentence by topic sentence, and hand it to a stranger who can follow it.
A · 2:05 The method. A paper is built outline first: one topic sentence per paragraph, in order, before any prose exists. We start backwards — take a published paper apart into the outline that must have produced it. That is also this week’s exercise, and the fastest way to see that good writing has a skeleton. Then forwards: outline, sketch, full text. Every written product in this course is submitted that way, so the argument is fixed before the sentences are. Then what a problem statement has to do. Scope, unit, population. The ‘so what’ in the first sentence, not the fourth paragraph. Then the section everybody gets wrong — identification: your design named, its assumption stated, and what would break it. And how to state a limitation without gutting your own project. Finally the failure modes: the topic with no question, hedging, the buried claim, the review that reviews without arguing.
B · 3:30 Your own problem statement, written in the room. One topic sentence per paragraph, in order, and nothing else — no prose. Then swap with another team. They get your skeleton and nothing else, and they have to tell you what your argument is. Where they cannot, your argument is not there yet. Whatever survives that is the spine of your pitch next week, and of the proposal outline in November.
11 Nov 4 Figures
Name the trick in a figure built to mislead you, and redraw it honestly without losing the point. Draw the one figure your own project lives or dies on, before you have the data to draw it with.
A · 2:05 Three blocks. First, which figure answers which question. Distributions: a histogram, a density, a boxplot and a cumulative curve of the same column tell you four different things, and only some of them show you the second bump. Time: levels, indexed to a base year, or on a log scale — that choice is a claim about whether the reader should care about dollars or about growth rates. Space: a map of counts is a map of population unless you divided by something. Comparison: a sorted dot plot beats a bar chart nearly every time, and error bars are part of the estimate rather than decoration. Underneath all of it is one instinct — look at the data before you model it. Anscombe’s quartet makes the case in one slide. Four datasets, the same means, the same variances, the same correlation, the same fitted line. One is a clean relationship, one is a curve, one is a line with a single outlier dragging it, one is a vertical stack where one point does all the work. The regression output is identical for all four. Second block: what figures hide. Aggregation, and the Berkeley graduate admissions case of 1973, where the university looked like it admitted women at a much lower rate overall while most departments taken one at a time did not. Scale, and axes that start where the author wanted them to. Binning, where the width of a histogram bin decides whether there is one group in your data or two. Projection, and Greenland the size of Africa. Area scaled by radius, which doubles the number and quadruples the ink. Cherry-picked windows. Counts where you needed a rate. Then the move that makes those useful: read a published figure for what it is not showing. Where is the denominator. What happened before the window starts. What do the subgroups do. What happened to the units that dropped out of the sample partway through. Third block, and it is the hard one. An honest figure that still makes a point. Honest does not mean neutral. You choose the comparison, sort the categories, show the raw points behind the summary, split one aggregate into small multiples, annotate the single number the argument turns on, and write the claim in the title instead of a label. Zeroing the axis is not a rule — for an index, a growth rate or a difference, zero is the wrong anchor. The test is whether the axis matches the quantity. And a finished figure has a sentence. If you cannot write the sentence, the figure is not finished.
B · 3:30 Six figures, each lying a different way, one per team, twenty minutes. A truncated axis turning three points of satisfaction into a soaring bar chart. Two axes on unrelated scales manufacturing a correlation between ad spend and sales. Four quarters shown out of eight years. An aggregate rising while every single region falls. Circles where two million got twice the radius of one million, so the picture is four times bigger and not two. Complaints by state with no population underneath. Name the trick, say what the figure is entitled to claim, and redraw it honestly on the sheet. The redraw still has to make a point, and that is where most teams stall — naming the lie takes two minutes, replacing it takes the rest. Then one team is drawn at random to put its redraw at the front while the room attacks, and we draw again as far as the clock allows. Points for an attack that lands, and you have to name the trick rather than say you dislike the picture. Card 4 is the aggregation one and card 6 is the rate one. Those two are where the morning either landed or it did not. Then your own project, on paper, by hand. Draw the one figure it lives or dies on. Not a figure — the figure, the one a reader looks at before deciding whether to believe you. For most of you it is one of three: the picture showing that your variation exists, the picture showing your comparison group looked like your treated group before anything happened, or the picture showing the outcome move. You do not have the data yet, so you draw it empty. Both axes labelled with their units. What one point on it is. Which comparison is on the page. The claim written across the top as a sentence. Then draw it a second time as it would look if you are wrong. If the two drawings look the same, the figure cannot support the claim, and finding that out in November is cheap. Both drawings are handed in before you leave.
12 Nov 11 Speaking
Give a five-minute pitch whose argument the room can repeat back to you, on slides where every title is a claim. Answer a question you cannot answer without bluffing.
A · 2:05 Same skeleton, new medium. Your talk is the outline from meeting 10 with pictures on it, and the pictures are the ones you rebuilt in meeting 11. Then the rule that does most of the work: one claim per slide, and the claim is the slide title. “Methodology” is a label. “We compare stores on either side of a state border, so the weather and the season are the same on both sides” is a claim. “Results” is a label. “Basket size rose 4% in relaunch stores and did not move next door” is a claim. Read your titles alone, top to bottom, with nothing else on the screen. You should have the argument. If you do not, the deck does not contain one, and no amount of delivery will put it there. Next, the thing everybody does and nobody defends. No bullet lists. A bullet list is an outline you stopped converting into sentences. The room reads it faster than you can say it, so you lose them for the rest of the slide, and you end up reading your own slide out loud to people who finished it. Then the two thirty-seconds. The first thirty seconds say what the question is and why anybody should care. Not your name. Not an agenda slide. Not thanking anyone. The last thirty seconds say what you want: in this genre, what will exist by May and what you need to get there. Everything in between is negotiable; those two are not. Then structure by length, because they are different objects. Five minutes is five slides — the question, the decision somebody makes differently, the design and the assumption carrying it, the data and how you actually get it, and what exists by May. Fifteen minutes is that same spine with room for the graph, one alternative explanation you considered and rejected, and the limitation stated plainly. It is not five minutes with more slides on the end. Then questions, which is where most pitches are actually won or lost. Three kinds arrive. Clarification: answer it in one sentence and move. Challenge: answer first, then explain, and do not relitigate the whole talk. And the one you cannot answer, which is the one worth practising. The answer is “I don’t know,” followed by what would settle it. That sentence costs you nothing and a bluff costs you the room, because the people asking do this for a living and they can hear it. Know your numbers cold — sample size, the headline magnitude, where the data comes from, what it costs — because those are what get asked, and fumbling them undoes twenty good minutes. Then delivery, briefly, and this is not charm school. You will talk too fast, so rehearse standing up and out loud, once, with a clock; reading it in your head is not rehearsal and it is why teams over-run. Pause after a claim. Silence is what makes it land, and it is the only tool here that works on everybody. Stand where the room can see you, not behind the laptop and not in the projector beam, and never turn your back to read your own slide. And plan for the deck failing, because eventually it does. Know your first three claims by heart. Carry a PDF you can open from a phone. Be able to give the whole thing with no slides at all, which if the titles are claims you nearly can. Finally, the genre you actually face. You are pitching a research plan to people who will never read the paper. They are not checking your work. They are deciding whether to fund it, staff it, or wait. So they ask three things: so what would you do differently, how confident are you, and what does it cost. A talk built for a seminar answers none of them. I will give you the same two-minute pitch twice at the front — once with an agenda slide and bullet lists, once with claim titles and nothing else — and you will tell me which one you can repeat back.
B · 3:30 Ten minutes of team check-in first, with one agenda item today: who says which slide at meeting 14, and who holds the clock. Then the shared case. I give a five-minute pitch from the front and it is deliberately bad — an agenda slide, a title called “Results,” four bullets a slide, the claim arriving on slide 4, forty seconds over, and one question answered with a bluff. Everybody scores the same pitch on the same card. What did the first thirty seconds do? Where did the claim first appear? What was the ask? Then the part that gets compared: every team rewrites my five slide titles as claims, working from the same five labels, and the best set goes on the screen next to mine. Same object, same constraint, so a good rewrite and a lazy one are visible side by side. Then the dress rehearsal, which is the rest of the hour. Order is drawn at random. Every team pitches its own project, three minutes, hard stop, five slides with claim titles. Then the room attacks — and it attacks the argument, not the delivery. Attacks come in four named types and nothing else counts: the data cannot answer this question, the assumption carrying the estimate is false in this setting, the comparison group is not comparable, or the answer would not change anyone’s decision. “You seemed nervous” is not an attack and scores nothing. Points for attacks that land, as always. Each team leaves with four things in writing. Their actual clock time. The attack slips, each naming its type. One line from me. And the one that does the real work: before any discussion, two other teams write down in a single sentence the claim they think you just made, and you write down the claim you meant. You leave holding both. Where those two sentences differ is your meeting-14 rewrite, and you have three weeks to close the gap. Handed in before you leave: the five-slide dress-rehearsal deck, your written thirty-second opening and thirty-second close, and the claim-we-meant against claim-they-heard sheet.
Proposal outline · Rate your teammates

Part 5 — Delivery

The midterm, and then the pitch. A semester of weekly pieces becomes one proposal somebody could pick up and execute in January.

# Date Meeting Due
13 Nov 18* Midterm
Design a study in writing, under exam conditions, for a situation you have not seen before.
A · 2:05 Midterm. You are given a situation and asked to design a study to evaluate it, in writing. The question. The design, and the assumption carrying it. What data it requires. And what would undermine it. Open note. It tests whether you can apply the design vocabulary to a situation you have not seen, not whether you memorised definitions. There is nothing to revise beyond having turned up all term — it is the same move the room has made every week since meeting 2.
B · 3:30 No session B. Class ends when you hand the midterm in.
Nov 25 Thanksgiving recess — no class
14 Dec 2 The pitch
Present a plan somebody could pick up and execute in January. Know what your team owes AEM 6992 over the break.
A · 2:05 Final presentations. Each team reviews another team’s work.
B · 3:30 The handoff to AEM 6992. What each team does over the break.
Final presentation, in class
Dec 11 Research proposal due — the final write-up. Date to be confirmed

* These dates are not final: Sep 23 — Instructor traveling — this meeting is held remotely · Oct 21 — Instructor traveling — moves to Oct 20 or Oct 22; date to be confirmed · Nov 18 — Instructor off campus — the midterm is proctored; arrangements announced by week 10

Thanksgiving recess falls on Wednesday, November 25 — no class. The review of another team’s paper is due in the exam period; the date for the final poster or plan is to be announced.

Readings and handbooks

There is no book to buy. Everything is free, either openly online or through the Cornell library.

Quizzes are drawn from the lecture, not from a reading, so nothing below is a gate. These are the shelf you reach for when a session raises something you want more of, and each meeting page points at the chapters that fit it.

The three handbooks

Cunningham, Causal Inference: The Remix — the second edition of The Mixtape, free to read online. The first edition is still up. This is the main one for us: readable without prior econometrics, and it covers the whole inference half.

1–2 Foundational ideas · which causal inference?
3–4 Directed acyclic graphs · potential outcomes and randomization
5 Unconfoundedness
6–7 Regression discontinuity · instrumental variables
8–11 Causal panel designs · difference-in-differences · complex DiD · synthetic control

Adams, Khan, Raeside & White (2007), Research Methods for Graduate Business and Social Science Students (SAGE) — the shared text across AEM 6991 sections, free through the Cornell library. Sign in with your NetID. It is where the non-inference craft lives: the research cycle, literature reviewing, sampling, surveys, and writing up.

1–2 Introduction to research · methodology and research ethics
3–4 The research cycle · literature review and critical reading
5 Sampling, and sample size determination
6–9 Primary data · secondary data · surveys · interviews
10–13 Qualitative analysis · descriptive statistics · correlation and regression · advanced methods
14 Tests of measurement and quality — reliability, validity, generalisability
15–16 Conducting your research · writing and presenting

Angrist & Pischke, Mostly Harmless Econometrics (Princeton, 2009) — the standard graduate reference for the designs in Part III.

This one assumes econometrics. It is not assigned and you are not expected to read it. It is here because it is the book the designs in this course come from, and because a few of you will want it — or will meet it in a job. Chapter 1, Questions about Questions, is eight pages and needs no maths. Start there or nowhere.

1–2 Questions about questions · the experimental ideal
3 Making regression make sense
4–6 Instrumental variables · fixed effects, DiD and panel data · regression discontinuity
7–8 Quantile regression · standard error issues

Also recurring

Huntington-Klein, The Effect, free online and the gentlest of the four. Sandberg & Alvesson on problematization; Haas & Mortensen on teams; Stanley & Castles, The So What Strategy; Zelazny, Say It With Charts.

Attendance

The format is once a week. So missing one meeting means missing two sessions.

Attendance is recorded by the session-B check-in, which cannot be filed retrospectively. Quizzes are unannounced, so missing a meeting risks missing one — and a missed quiz scores zero. The dropped score absorbs one absence without penalty and without a conversation.

If you are going to miss more than two meetings, talk to me early rather than late.

Data practices

Quiz submissions are tied to your netid and timestamped. They are used for grading and attendance.

The roster, your submissions, and any per-student credentials are held in Cornell-managed or instructor-controlled storage. None of it is published. Course materials in the public repository never contain student-identifying data.

Academic integrity

Every student is expected to abide by the Cornell University Code of Academic Integrity. Work submitted for credit is your own.

AI assistance is permitted under the policy above. Failure to disclose AI use is a violation. So is submitting a fabricated citation as a real one.

Collaboration is expected on team assignments. It is not permitted on individual ones.

Accommodations

If you need academic accommodations, contact Student Disability Services. Then give me your accommodation letter as early in the semester as you can, so arrangements can be made.

Course materials

Slides, readings, assignment specifications and the schedule live on this site. It is generated from a public repository and updated through the term.

Canvas carries grades and submissions.


Last updated: 2026-08-25. Schedule and policies may still change.