StudentSim:训练基于 LLM 的学生模拟器框架

内容摘要
StudentSim是一种基于LLM的学生模拟器训练框架,旨在通过整合学生数据,生成个性化的模拟器。该框架能够模仿学生的回答,并在导师指导下更新这些回答。StudentSimEval是一个标准化评估协议,涵盖60名学生在国际象棋、英语写作和数学三个领域的表现。评估结果显示,StudentSim在行为忠实度和指导响应度两个指标上均优于GPT-5.4。在国际象棋领域,StudentSim的行为忠实度达到0.51,指导响应度达到0.91,而GPT-5.4的相应指标为0.23和0.72。此外,将StudentSim作为奖励模型用于导师强化学习,产生的国际象棋导师在准确性、指导性和个性化方面均优于无强化学习基线和对抗GPT-5.4模拟器奖励的导师。相关代码可在GitHub上获取。
StudentSim是一种基于LLM的学生模拟器训练框架,旨在通过整合学生数据,生成个性化的模拟器。该框架能够模仿学生的回答,并在导师指导下更新这些回答。StudentSimEval是一个标准化评估协议,涵盖60名学生在国际象棋、英语写作和数学三个领域的表现。评估结果显示,StudentSim在行为忠实度和指导响应度两个指标上均优于GPT-5.4。在国际象棋领域,StudentSim的行为忠实度达到0.51,指导响应度达到0.91,而GPT-5.4的相应指标为0.23和0.72。此外,将StudentSim作为奖励模型用于导师强化学习,产生的国际象棋导师在准确性、指导性和个性化方面均优于无强化学习基线和对抗GPT-5.4模拟器奖励的导师。相关代码可在GitHub上获取。

AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.

原始发布方:HuggingFace Daily Papers(社区热门论文)

原文时间:2026-09-01 08:00:00 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值