[Submitted on 30 Apr 2026 (v1), last revised 24 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason about mathematics across human- and AI-like personifications. MEDS comprises 28,000 runs from 14 LLMs (i.e., Mistral, Qwen, DeepSeek, IBM Granite, Microsoft Phi, and xAI Grok) generated under human-shadow and AI-assistant conditions. Each record (digital shadow) includes a set of prompts; psychological and sociodemographic metadata; and four mathematics tasks: (i) interviews about relationships with mathematics, (ii) three psychometric questionnaires on mathematics self-efficacy and anxiety, (iii) one cognitive network capturing attitudes towards mathematics, and (iv) 18 high-school mathematics quiz items enriched with reasoning explanations and confidence scores. Analyses of the data show that LLMs exhibit differences in attitudes and performance across human-shadow and AI-assistant modes, including human-like negative attitudes towards mathematics, logical fallacies, and overconfidence in mathematics. As a data resource, MEDS can benefit learning scientists and developers of safer AI tutors in mathematics.

Submission history

From: Ali Aghazadeh Ardebili Dr. [view email]
[v1] Thu, 30 Apr 2026 09:08:58 UTC (8,141 KB)
[v2] Fri, 24 Jul 2026 16:21:41 UTC (3,627 KB)