OpenLiteracy: An Open-Source AI Infrastructure Suite for Advancing Speech Foundation Models for Early Word Reading Assessment and Instruction

Team: Dr. Ying Xu, Dr. Phil Capin, Dr. Zhonghao Shi

Overview: OpenLiteracy builds an open-source AI infrastructure suite to advance speech foundation models for early word reading assessment and instruction.

Subject Focus: Reading/Literacy, Grades K-3

Targeted Universalism Focus: The synthetic samples will represent diverse reading pathologies and ensure the model learns from underrepresented error patterns.

Public Goods & Deliverables:

  • Benchmark: A model-agnostic evaluation suite containing 20,000 authentic, phoneme-annotated child speech utterances across diverse tasks and demographics, complete with phoneme error rate (PER/PFER), mispronunciation detection (MDD), and fairness metrics.
  • Synthetic Data: A curated inventory of phoneme-level reading errors, a validated text-to-speech generation pipeline (SSML/AWS Polly), 50,000 human-validated synthetic speech utterances, and fine-tuned speech model weights.
  • Consortium: Institutional infrastructure providing standardized speech de-identification (PII removal and randomized vocal feature masking), central IRB protocols, data-sharing agreements, and external data access licensing under Harvard’s Child-Centered AI Lab.

Return to see all grantees.