Comprehensive Multimodal Writing and Feedback Dataset

Team: Dr. Xiang Lorraine Li, Dr. Gayle Rogers, Dr. Diane Litman, Dr. Raquel Coelho

Overview: This project creates a multimodal writing and feedback dataset capturing the transitional phase between high school and college writing. It incorporates multi-draft student essays, detailed instructor feedback on idea development, and office-hour audio interactions.

Subject Focus: Writing, Grade 12-College Freshman

Targeted Universalism Focus: The dataset centers culturally and linguistically diverse writers, including English as a Second Language (ESL) students, to support research on personalized writing development across diverse backgrounds.

Public Goods & Deliverables:

  • Multimodal Writing Corpus: ~2,000 long-form essays across 500 students with 2–3 aligned drafts per assignment, margin/summative instructor feedback, rubric scores, and anonymized demographic metadata (released under a 90% public / 10% held-out test split).
  • Instructional Dialogue Audio: 300–450 hours of timestamped, Whisper-transcribed student–instructor paper conferences linked directly to specific essay sections.
  • Open Tools & Benchmarks: Open-source PDF-to-JSON conversion scripts, PII obfuscation pipelines, feedback/idea-development annotation schemas, and baseline LLM evaluation models published on GitHub.

Return to see all grantees.