Team: Dr. Xiang Lorraine Li, Dr. Gayle Rogers, Dr. Diane Litman, Dr. Raquel Coelho
Overview: This project creates a multimodal writing and feedback dataset capturing the transitional phase between high school and college writing. It incorporates multi-draft student essays, detailed instructor feedback on idea development, and office-hour audio interactions.
Subject Focus: Writing, Grade 12-College Freshman
Targeted Universalism Focus: The dataset centers culturally and linguistically diverse writers, including English as a Second Language (ESL) students, to support research on personalized writing development across diverse backgrounds.
Public Goods & Deliverables:
- Multimodal Writing Corpus: ~2,000 long-form essays across 500 students with 2–3 aligned drafts per assignment, margin/summative instructor feedback, rubric scores, and anonymized demographic metadata (released under a 90% public / 10% held-out test split).
- Instructional Dialogue Audio: 300–450 hours of timestamped, Whisper-transcribed student–instructor paper conferences linked directly to specific essay sections.
- Open Tools & Benchmarks: Open-source PDF-to-JSON conversion scripts, PII obfuscation pipelines, feedback/idea-development annotation schemas, and baseline LLM evaluation models published on GitHub.