The “Productive Coaching” Benchmark: Validating Coaching Quality in Student-Facing, AI-Powered Learning Tools

Team: Peter Gault, Sarah Akhtar, Dr. Rachel Phillips, Jared Fritz

Overview: Quill.org is developing an open benchmark and public dataset to evaluate the quality of AI-powered “productive coaching” in student-facing learning tools. Grounded in a five-dimension rubric, the benchmark assesses whether AI feedback is accurate, actionable, supportive, and aligned with student thinking.

Subject Focus: Writing, Grades 8-12

Targeted Universalism Focus: Title I educators will be involved in dataset annotation and validating benchmark performance across diverse institutional contexts.

Public Goods & Deliverables:

  • Training Dataset: A dataset of 10,000+ de-identified, educator-written feedback examples establishing reference standards for 1:1 coaching.
  • Multi-Turn Evaluation Dataset: A public dataset of 1,000+ real classroom multi-turn AI coaching sessions sampled across at least 4 learning tools, annotated across the 5 coaching dimensions and paired with revision improvement labels.
  • Public Benchmark Leaderboard & Adoption Package: Hosted via Learning Commons, featuring codebooks, open-source scoring pipelines, and an implementation playbook website for district decision-makers. Released under CC BY 4.0.

Return to see all grantees.