22 · Audio
Tuesday, Nov 10, 2026
Materials for this session are not published yet. They appear here before class.
Objectives
By the end of this session you can:
- Transcribe speech with timestamps and speaker labels.
- Extract the metadata that means something: pauses, disfluencies, tone.
- Build the utterance table.
- Audit a transcript by sampled listening, with a word-error rate.
What we cover
- What audio has measured: the Fed chair’s vocal tone moves markets; the affect in earnings calls predicts.
- Whisper-class transcription with timestamps and speakers.
- Past the words: pauses, disfluencies, pitch and energy, emotion labels — audio is text with meaningful metadata.
- The deliverable shape: one row per utterance — speaker, text, start and end, pause before, tone.
Verification habit. Spot-listen to a sample; report word-error rate on a hand-checked subset.