Team: Dr. Walter Leite, Dr. Zoey Liu, Dr. Anthony Botelho
Overview: The Storiza Corpus, comprising over 80 hours of early elementary children reading AI-generated stories aligned with specific phonetic patterns will provide word- and sentence-level annotations to advance reading fluency and speech technology development.
Subject Focus: Reading/Literacy, Grades: K-3
Targeted Universalism Focus: The corpus explicitly includes children from all five major U.S. geographic regions to ensure reading AI tools are inclusive of diverse U.S. speech dialects.
Public Goods & Deliverables:
- Storiza Corpus: 80 total hours of child oral reading audio (.wav) across all major U.S. dialects. Tier 1 includes 35–40 hours of gold-standard, human-verified, word-level time-aligned audio with orthographic and IPA phonemic transcriptions, plus 13 distinct reading error categories. Tier 2 contains ~40 hours of automatically annotated audio.
- System & xAPI Logs: Anonymized logs covering ~5,400 AI-generated stories, student interaction data, and 1,800 maze comprehension test results.
- Open Models & Code: Open-source ASR baselines, reading error classification models (published on Hugging Face), and automated annotation scripts/data dictionaries (published on GitHub and OSF.io). All restricted audio assets are governed through the LDC repository under IRB data-banking approvals.