Team: Dr. Ying Xu, Dr. Phil Capin, Dr. Zhonghao Shi
Overview: OpenLiteracy builds an open-source AI infrastructure suite to advance speech foundation models for early word reading assessment and instruction.
Subject Focus: Reading/Literacy, Grades K-3
Targeted Universalism Focus: The synthetic samples will represent diverse reading pathologies and ensure the model learns from underrepresented error patterns.
Public Goods & Deliverables:
- Benchmark: A model-agnostic evaluation suite containing 20,000 authentic, phoneme-annotated child speech utterances across diverse tasks and demographics, complete with phoneme error rate (PER/PFER), mispronunciation detection (MDD), and fairness metrics.
- Synthetic Data: A curated inventory of phoneme-level reading errors, a validated text-to-speech generation pipeline (SSML/AWS Polly), 50,000 human-validated synthetic speech utterances, and fine-tuned speech model weights.
- Consortium: Institutional infrastructure providing standardized speech de-identification (PII removal and randomized vocal feature masking), central IRB protocols, data-sharing agreements, and external data access licensing under Harvard’s Child-Centered AI Lab.