KB-TutorBench: A Multimodal Knowledge-Building Dataset for AI-Enabled Formative Assessment

Team: Dr. Hari Subramonyam, Mayank Sharma, and Dr. Roy Pea

Overview: KB-TutorBench is a dataset and benchmark for training and evaluating AI tutors that support knowledge-building and formative assessment, using 6,000 annotated tutoring dialogues from simulated and real classroom interactions.

Subject Focus: Earth Science, Chemistry, Biology, ELA, and History across grades K-5, 6-8, and 9-12.

Targeted Universalism Focus: The dataset integrates diverse learner profiles directly into its data collection protocols, explicitly highlighting neurodiverse learners (such as those with dyscalculia or high verbal/low spatial reasoning), English Language Learners (ELL), and students from varied socioeconomic contexts.

Public Goods & Deliverables:

  • The complete dataset of 6,000+ text dialogues and spoken transcripts under a CC-BY-4.0 license on HuggingFace
  • Standardized scenario seeds and rubrics allowing external developers to submit their AI tutor outputs for multi-dimensional evaluation.
  • Fine-tuned model weights on HuggingFace (such as a Mistral-7B via QLoRA fine-tuning) that perform comparably to frontier proprietary models like Claude Sonnet 4.5 on pedagogical quality.
  • A companion GitHub repository housing data collection tools, annotation code, safety filters, and user notebooks under an Apache 2.0 license.

Return to see all grantees.