Open Leaderboards for Benchmarking Automated Speech Recognition in Educational Contexts

Team: Allison Koenecke, Rachel Slama, Rene Kizilcec, and Kirk Vanacore

Overview: This project addresses a key K–12 AI infrastructure gap: ASR systems underperform for children and marginalized speakers. We will expand tutoring audio datasets, release a child-optimized ASR baseline, and create a public leaderboard to benchmark models using education-specific metrics.

Subject Focus: K-12 math tutoring audio involving diverse demographic cohorts

Targeted Universalism Focus: The project targets equity by focusing its evaluation on a highly diverse student baseline, specifically analyzing ASR performance on minoritized speaker groups, students with clinical speech impairments, and students utilizing regional dialects such as African American English

Public Goods & Deliverables:

  • An open leaderboard on HuggingFace tracking aggregate and subgroup-disaggregated metrics across leading proprietary, open-source, and fine-tuned ASR services.
  • Model weights and reproducible training recipes for leading open architectures (such as Whisper and wav2vec) fine-tuned on National Tutoring Observatory partner data.
  • A standard template for measuring and reporting subgroup disparities across age bands, accents, backgrounds, and clinical speech impairments.

Return to see all grantees.