Where to get data

Read the last section first if you are short of time. The most expensive mistake in a capstone is designing around data you cannot get in time. Some of what follows takes minutes to obtain. Some takes months. You cannot tell which from the outside.

Dewey

app.deweydata.io · a commercial data marketplace Cornell JCB subscribes to. Roughly 84 datasets, free to you.

Getting an account. Register with your Cornell email address. A personal address will not match the subscription.

Your account then sits pending until an administrator approves it. That takes days. Register in week 1, not the night you need it.

There is no IP-based access and no shared login.

What is in it, by family:

Family Examples Typical unit
Property Real Estate Transaction Data (assessor, sales, distress, 2011–); Commercial Real Estate Listings; Rental Data (41,702 ZIPs, 2014–); Dwellsy rent indices Parcel, listing, ZIP
Foot traffic Weekly Patterns Plus (US + Canada, 2017–); Daily Store Visits; Neighborhood Patterns; Global Places (POI, NAICS) Place × week/day
Spending Card panels (100M+ cards, ~13k merchants); Brand Tracker; Spend Patterns; Convenience Store Transactions (SKU level) Brand, POI, SKU
Labour Job Postings (US, with salary, 2016–2025); Job Records (global, 2007–); Layoff Data (WARN, 1996–); Occupational Wages; People and Resume Posting, notice, person
Firms Company Insights (30M+ firms, headcount, churn); Business Entity (50-state registries); Public Company Financials; ESG Scores Company
Markets Global Equity EOD (170+ exchanges); Bond Prices & Liquidity; Credit Curves; Futures & Options; Social Sentiment Security × day
Place & risk US Climate Risk (5 hazards × ZIP); Weather (daily, 1970–); Nature Score; CDC PLACES; ACS block-group tables ZIP, tract, block group
Health & policy Payer Rates (negotiated rates, 293 codes); SBA Loans (13.4M PPP/504/7a); US Lobbying (1999–); government contract awards Claim code, loan, award

Each dataset’s own page carries the schema, coverage and date range. Read it before you build a question on it — that is feasibility filter 2, and it takes ten minutes.

Through the Cornell library

All of these need only a netID.

Companies and finance

  • WRDS (Wharton Research Data Services) — the standard academic host for historical finance and accounting data. Extensive, with use restrictions; talk to a librarian before pulling anything large.
  • Capital IQ — 88,000 public and 825,000 private companies: financials, transactions, people.
  • Refinitiv Workspace — company financials, market data, transactions, analyst reports. Register with your Cornell email.
  • PitchBook — private equity and venture capital deals, investors, comparables. Log in with SSO using your Cornell address.
  • Mergent Intellect and Data Axle Reference Solutions — private company directories, US and international.
  • EDGAR — every SEC filing, free and public, no login at all.

Bloomberg terminals are physical machines: three at Mann Library on the main level, plus the Management, Hotel and Law libraries. Worth an hour even if your project never uses one.

Markets and industries

Mintel · Passport (Euromonitor) · Statista · IBISWorld (700+ US industries) · Frost & Sullivan · Technavio (20 downloads per Cornell email per month) · Business Source Complete · ABI/Inform.

These are reports, not microdata. They are excellent for sizing a market and for the background half of a feasibility memo. They will not give you anything to estimate on.

Demographics and geography

SimplyAnalytics builds thematic maps and reports from US demographic, business and marketing data — the fastest way to check whether a geographic question is even viable. Then the free federal sources: data.census.gov, BLS, and the Consumer Expenditure Survey.

Restricted data, through CCSS

The Cornell Center for Social Sciences runs Cornell’s secure data services, and holds things you cannot get any other way.

  • Kilts-Nielsen marketing data — Consumer Panel, Retail Scanner (weekly price and volume from 90+ retail chains across all US markets), Ad Intel advertising data from 2010, and the PanelViews surveys. For Business of Food or Behavioral Marketing this is close to the best data in the world for the question.

    Getting it is a project, not a download. The data lives in CCSS’s Regulated Research Environment, and access means applying through Cornell Research Services and requesting project space. Budget months, and start by emailing socialsciences@cornell.edu.

    There is also a public-use extract in the CCSS archive, open to any Cornell netID. Be clear about what it is: retail sales of coffee, laundry detergent and shampoo, by US region, brand, size, packaging, UPC and price, for 2019 only. That is a dataset for practising on, not one to build a capstone on. Do not plan a project around it.

  • FSRDC — confidential Census, BLS and NCHS microdata, through one of 30 Federal Statistical Research Data Centers. Requires federal project approval and takes longer than this degree.

  • IAB@Cornell — German administrative labour-market microdata.

Free and public, worth knowing

  • FRED — 800,000+ economic series, the fastest macro source there is.
  • IPUMS — harmonised census and survey microdata, US and international. Free with registration.
  • ICPSR — the largest social science data archive; Cornell is a member.
  • Roper Center — public opinion polling, hosted at Cornell.
  • World Bank Open Data and UN Comtrade — the backbone for International Development Economics, where Dewey is weakest.
  • Zillow research data and the Redfin Data Center — free housing indices at ZIP and metro level.
  • County assessor and deeds records — public in most US counties. Worth knowing that this is precisely what ATTOM, the source behind Dewey’s Real Estate Transaction Data, already packages nationally and cleans for you. Going to a county directly is worth it only for a specific reason: one county in unusual depth, a longer history than ATTOM carries, or a field it does not hold.

Collecting it yourself

Sometimes nobody has your data. That usually means nobody has asked your question. It is a good sign, not a problem.

  • Surveys. Cornell provides Qualtrics to all students. Fielding to a paid sample through Prolific or MTurk costs real but small money and needs no programming.
  • Online experiments. A message test, a price test, or a conjoint task can be built in Qualtrics and fielded the same way.
  • Asking a firm. The most ambitious project anyone here remembers talked a soda company into donating thousands of cans for a field experiment. People say yes more often than students expect.

All of this needs IRB clearance — see permissions and ethics, and start it early.

The access ladder — read this one

You want How you get it How long
Library databases, Statista, IBISWorld, SimplyAnalytics, FRED, EDGAR netID, or nothing Now
Dewey Register with a Cornell address, wait for approval Days
The Nielsen public-use extract (2019, three product categories) netID Now — but it is a practice dataset
Your own survey or online experiment IRB, then field it Weeks
Full Kilts-Nielsen data Project application through Cornell Research Services, then space in the Regulated Research Environment Months
A data-use agreement with a firm Legal review on both sides Months, and it may fail
Census, BLS or NCHS confidential microdata via FSRDC Federal project approval Longer than this degree

Which row you are in is feasibility filter 2. Getting it wrong is the most common way an MPS capstone dies. Meeting 5 is about this.

Ask a librarian

Tom Ottaviano is Mann Library’s business and economics librarian — tjo65@cornell.edu, and his profile page has a booking link for appointments. Fifteen minutes with him in September is worth more than a week of guessing in November.

For the library generally, mann-ref@cornell.edu reaches the Mann reference desk.