The Storiza Corpus of Early Reader Story Recordings: A Dialect-Inclusive Speech Dataset for Reading Technology Development

Team: Dr. Walter Leite, Dr. Zoey Liu, Dr. Anthony Botelho

Overview: The Storiza Corpus, comprising over 80 hours of early elementary children reading AI-generated stories aligned with specific phonetic patterns will provide word- and sentence-level annotations to advance reading fluency and speech technology development.

Subject Focus: Reading/Literacy, Grades: K-3

Targeted Universalism Focus: The corpus explicitly includes children from all five major U.S. geographic regions to ensure reading AI tools are inclusive of diverse U.S. speech dialects.

Public Goods & Deliverables:

  • Storiza Corpus: 80 total hours of child oral reading audio (.wav) across all major U.S. dialects. Tier 1 includes 35–40 hours of gold-standard, human-verified, word-level time-aligned audio with orthographic and IPA phonemic transcriptions, plus 13 distinct reading error categories. Tier 2 contains ~40 hours of automatically annotated audio.
  • System & xAPI Logs: Anonymized logs covering ~5,400 AI-generated stories, student interaction data, and 1,800 maze comprehension test results.
  • Open Models & Code: Open-source ASR baselines, reading error classification models (published on Hugging Face), and automated annotation scripts/data dictionaries (published on GitHub and OSF.io). All restricted audio assets are governed through the LDC repository under IRB data-banking approvals.

Return to see all grantees.