Researchers launched CSTutorBench to evaluate small language models as tutors in VEX VR block-based programming, finding models perform well on surface criteria but fail to avoid answer leakage and engage with debugging history, highlighting the need for context-specific pedagogical benchmarks.

1 min read

CSTutorBench: A New Benchmark for Evaluating Small Language Models as Block-Based Programming Tutors

Introduction

FAQ

What is CSTutorBench?

A benchmark for evaluating small language models as tutors in VEX VR block-based programming, consisting of 17 pedagogical scenarios.

Why are small language models important in education?

They offer a cheaper and more private alternative to large models, especially in K-12 settings.

What are the key findings?

Models perform well on surface criteria but fail in deep pedagogical behaviors like avoiding answer leakage.

Can model performance be improved?

Yes, pedagogical prompt revision improved scores for 10 of 11 models.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.