Building a Proof-of-Concept Dataset for AI-Enabled Formative Writing Feedback in Science: Validating Critical Thinking Through Claims, Evidence, and Reasoning

Team: Yusuf Ahmad, Dr. Nadine Dabby, Hilah Barbot, Michele Leardo, Nidhi Hebbar Overview: Playlab is creating the first open-license dataset of AI-driven formative feedback on middle and high school science writing. Derived from OpenSciEd-aligned tools, this project will build the systems to clean, de-identify, annotate, and prepare a proof-of-concept dataset from thousands of CER (Claims, … Read more

The “Productive Coaching” Benchmark: Validating Coaching Quality in Student-Facing, AI-Powered Learning Tools

Team: Peter Gault, Sarah Akhtar, Dr. Rachel Phillips, Jared Fritz Overview: Quill.org is developing an open benchmark and public dataset to evaluate the quality of AI-powered “productive coaching” in student-facing learning tools. Grounded in a five-dimension rubric, the benchmark assesses whether AI feedback is accurate, actionable, supportive, and aligned with student thinking. Subject Focus: Writing, … Read more

Comprehensive Multimodal Writing and Feedback Dataset

Team: Dr. Xiang Lorraine Li, Dr. Gayle Rogers, Dr. Diane Litman, Dr. Raquel Coelho Overview: This project creates a multimodal writing and feedback dataset capturing the transitional phase between high school and college writing. It incorporates multi-draft student essays, detailed instructor feedback on idea development, and office-hour audio interactions. Subject Focus: Writing, Grade 12-College Freshman … Read more

The Storiza Corpus of Early Reader Story Recordings: A Dialect-Inclusive Speech Dataset for Reading Technology Development

Team: Dr. Walter Leite, Dr. Zoey Liu, Dr. Anthony Botelho Overview: The Storiza Corpus, comprising over 80 hours of early elementary children reading AI-generated stories aligned with specific phonetic patterns will provide word- and sentence-level annotations to advance reading fluency and speech technology development. Subject Focus: Reading/Literacy, Grades: K-3 Targeted Universalism Focus: The corpus explicitly … Read more

Privacy-Preserving, Discipline-Specific AI Writing Feedback: A Public Good for Under-Resourced Rural Districts

Team: Brian Pool, Jen Couch, Dan Studebaker, Sarah Hollingsworth Overview: This project enhances four production open-source Moodle plugins to provide privacy-preserving, discipline-specific AI writing feedback for under-resourced rural schools. It operationalizes formative feedback principles in ELA, Science, and Social Studies across grades 5–12. Subject Focus: Writing, Grade 5-12 Targeted Universalism Focus: Centers under-resourced rural students … Read more

OpenLiteracy: An Open-Source AI Infrastructure Suite for Advancing Speech Foundation Models for Early Word Reading Assessment and Instruction

Team: Dr. Ying Xu, Dr. Phil Capin, Dr. Zhonghao Shi Overview: OpenLiteracy builds an open-source AI infrastructure suite to advance speech foundation models for early word reading assessment and instruction. Subject Focus: Reading/Literacy, Grades K-3 Targeted Universalism Focus: The synthetic samples will represent diverse reading pathologies and ensure the model learns from underrepresented error patterns. … Read more

Enhancing Two Multimodal Classroom Datasets to Advance R&D on Formative Assessment

Team: Dr. Jing Liu, Dr. Meghavarshini Krishnaswamy, Michael Chrzan Overview: This project enhances two existing multimodal classroom datasets (EDSI and NCTE) to support AI-enabled formative assessment in K–8 mathematics. It transcribes small-group student interactions and enriches transcripts with turn-level discourse annotations and observation codes. Subject Focus: Mathematics, Grades K-8 Targeted Universalism Focus: This dataset incorporates … Read more

A Multimodal Dataset for AI-Enhanced Formative Assessment of Additive-to-Multiplicative Reasoning: AMR Pilot

Team: Dr. Heidi Cian, Cheryl Tobey, Dr. Ibrahim Dahlstrom-Hakki, Rebecca Neville Overview: The AMR Pilot project produces a public dataset focused on students’ transition from additive to multiplicative reasoning in grades 2–5. Partnering with TERC and 40 teachers, the project will collect and annotate 2,400 student reasoning episodes from formative fluency interviews. Subject Focus: Mathematics, … Read more