Meg Benner

Author directory

2026

This study reports on an open data science competition based on a benchmark dataset of mathematics misunderstandings, comprising over 52,000 mathematics explanations written to justify answer choices from 15 multiple-choice questions that were labeled by expert human annotators. Competitors were tasked with correctly classifying the explanations and any misunderstandings. By combining stable validation methods with efficient inference and enriched training data, top teams achieved high accuracy scores that correctly classified the explanations as correct, a misunderstanding, or neither and, if it did have a misunderstanding, what type of misunderstanding it was.