The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences

arxiv logo

Authors: Matey Yordanov, Mikhail Bychkov, Andrei Kuchma Abstract: Traditional human marking in upper-secondary STEM education creates a structural bottleneck that restricts the frequency of formative mock examinations. This quasi-experimental, mixed-methods longitudinal study (N = 142) investigates the efficacy of deploying a fully automated, handwritten assessment marking platform to remove this bottleneck. Students preparing for STEM … Read more

KB-TutorBench: A Multimodal Knowledge-Building Dataset for AI-Enabled Formative Assessment

Stanford Graduate School of Education Logo

Team: Dr. Hari Subramonyam, Mayank Sharma, and Dr. Roy Pea Overview: KB-TutorBench is a dataset and benchmark for training and evaluating AI tutors that support knowledge-building and formative assessment, using 6,000 annotated tutoring dialogues from simulated and real classroom interactions. Subject Focus: Earth Science, Chemistry, Biology, ELA, and History across grades K-5, 6-8, and 9-12. … Read more

Open Leaderboards for Benchmarking Automated Speech Recognition in Educational Contexts

Cornell, National Tutoring Observatory logo

Team: Allison Koenecke, Rachel Slama, Rene Kizilcec, and Kirk Vanacore Overview: This project addresses a key K–12 AI infrastructure gap: ASR systems underperform for children and marginalized speakers. We will expand tutoring audio datasets, release a child-optimized ASR baseline, and create a public leaderboard to benchmark models using education-specific metrics. Subject Focus: K-12 math tutoring … Read more

Pipeline for Training Simulated Student Models with Reduced Human Data

Princeton AI Lab logo

Team: Tammy Kwan and Brenden Lake Overview: This project develops a training pipeline for neural-network-based simulated student models. The approach combines synthetic reasoning traces generated by rule-based student models with human student reasoning data to substantially reduce the amount of human data required. Subject Focus: Fraction arithmetic within grades 6-8, with exploratory transfer to decimals, … Read more

A Benchmark for AI Identification of Science Misconceptions in Free-Form Student Responses

Learning Equality logo

Team: Richard Tibbles, Jamie Alexandre Overview: Science misconception research spans decades but isn’t operationalized for classroom use. Three public goods, all openly licensed: a benchmark for identifying science misconceptions in free-form student responses, a structured dataset, and reference classifiers for classroom devices. Subject Focus: Grades 6-12 Earth, Life, and Physical sciences Targeted Universalism Focus: The … Read more

Formative Assessment and Public Goods for AI in Education

k12AI AI Infrastructure Program Logo

Author: K-12 AI Infrastructure Program Team  Download a PDF copy of this Formative Assessment page Definition Formative assessment is a key aspect of many applications of AI to education. Through production of public goods (as defined in the Request for Proposals), the K-12 AI Infrastructure program seeks to aid developers of AI-enabled products and services. … Read more

Building the Future of EdTech: 5 Essential Reads from The Learning Agency

Building high-quality AI-enabled edtech requires an adaptive understanding of open licensing, structured pipelines that involve humans, and leveraging the open-source benchmarks, datasets, and models already available to the community. If you are looking to learn more about building more responsible and effective AI-enabled edtech for education, here are resources from The Learning Agency to explore: … Read more

Request for Proposals #1 – Formative Assessment

RFP Overview Download the RFP as a PDF Formative Assessment and Public Goods for AI in Education Frequently Asked Questions         Notice: Award decisions for RFP #1 were emailed on May 18th. Please check your personal (and institutional) spam filters for this message. If you did not recieve an email, please reach … Read more

Sandpiper for Education

National Tutoring Observatory Logo

Sandpiper is a new open-source annotation app developed by the National Tutoring Observatory (NTO) to help researchers, educators, and organizations analyze tutoring interactions at scale. Learn more: Sandpiper Launch Recording