Part 1: AI Harnesses for Education: What is this Emerging Layer of AI Infrastructure?

Tenth-grade boy and girl work on computer in workshop Tenth graders work together on research for a project.

Author: K-12 AI Infrastructure Program Team In an age where budgets are tightening and there is more work on an educator’s plate, there are clear areas of promise where Artificial Intelligence (AI) could automate processes and lighten educator’s loads. However, the inconsistency in the reliability, accuracy and privacy of AI models, as well as the … Read more

The Responsible AI and Tech Justice Guide for K-12 Education

Cover of Responsible AI and Tech Justice for K12 Education Report

From The Kapor Foundation by Shana V. White and Allison Scott, Ph.D. Download the Responsible AI and Tech Justice Guide for K-12 Education Amidst the debates on AI use in schools, its impacts on teaching and learning, and its potential for future job disruption, one fact remains clear: All students will need to be equipped … Read more

Access the Public Goods Scanner

Landing page of the AI scanner showing various public goods in rows

Access the Scanner: Browse and filter K-12 AI resources including datasets, benchmarks, models, and competitions by grade level, academic subject, and source platform to evaluate and improve AI systems in education. Please share any resources you would like added to this database with k12ai-infrastructure [at] digitalpromise [dot] org

K-12 AI Infrastructure: Findings from Educator and Developer Outreach

Cover image for report: Findings from Educator and Developer Outreach image of tall man lecturing in front of classroom

Authors: Rebecca Griffiths, Vanessa Peters Hinton, Shelton Daal, and Megan Pattenhouse https://doi.org/10.51388/20.500.12265/309 Abstract This report presents findings from the K-12 AI Infrastructure Program’s outreach to educators, edtech developers, and other stakeholders, conducted during late 2025 to mid-2026. It examines what educators want from AI tools, where current AI applications fall short, how education technology developers … Read more

The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences

arxiv logo

Authors: Matey Yordanov, Mikhail Bychkov, Andrei Kuchma Abstract: Traditional human marking in upper-secondary STEM education creates a structural bottleneck that restricts the frequency of formative mock examinations. This quasi-experimental, mixed-methods longitudinal study (N = 142) investigates the efficacy of deploying a fully automated, handwritten assessment marking platform to remove this bottleneck. Students preparing for STEM … Read more

KB-TutorBench: A Multimodal Knowledge-Building Dataset for AI-Enabled Formative Assessment

Stanford Graduate School of Education Logo

Team: Dr. Hari Subramonyam, Mayank Sharma, and Dr. Roy Pea Overview: KB-TutorBench is a dataset and benchmark for training and evaluating AI tutors that support knowledge-building and formative assessment, using 6,000 annotated tutoring dialogues from simulated and real classroom interactions. Subject Focus: Earth Science, Chemistry, Biology, ELA, and History across grades K-5, 6-8, and 9-12. … Read more

Open Leaderboards for Benchmarking Automated Speech Recognition in Educational Contexts

Cornell, National Tutoring Observatory logo

Team: Allison Koenecke, Rachel Slama, Rene Kizilcec, and Kirk Vanacore Overview: This project addresses a key K–12 AI infrastructure gap: ASR systems underperform for children and marginalized speakers. We will expand tutoring audio datasets, release a child-optimized ASR baseline, and create a public leaderboard to benchmark models using education-specific metrics. Subject Focus: K-12 math tutoring … Read more

Pipeline for Training Simulated Student Models with Reduced Human Data

Princeton AI Lab logo

Team: Tammy Kwan and Brenden Lake Overview: This project develops a training pipeline for neural-network-based simulated student models. The approach combines synthetic reasoning traces generated by rule-based student models with human student reasoning data to substantially reduce the amount of human data required. Subject Focus: Fraction arithmetic within grades 6-8, with exploratory transfer to decimals, … Read more

A Benchmark for AI Identification of Science Misconceptions in Free-Form Student Responses

Learning Equality logo

Team: Richard Tibbles, Jamie Alexandre Overview: Science misconception research spans decades but isn’t operationalized for classroom use. Three public goods, all openly licensed: a benchmark for identifying science misconceptions in free-form student responses, a structured dataset, and reference classifiers for classroom devices. Subject Focus: Grades 6-12 Earth, Life, and Physical sciences Targeted Universalism Focus: The … Read more