Grading command-line exams at scale is difficult because rule-based autograders can't handle partial credit or syntactic variation, while manual grading doesn't scale with rising enrollments. This study tests whether four frontier LLMs (GPT, Claude Opus, Gemini, GLM) can approximate expert human grading of Linux/bash responses, using a four-level cognitive taxonomy from basic file operations to advanced system management. Gemini 3.0 Pro with rubric-guided prompting achieved the strongest human-AI agreement, though accuracy declined for harder questions. This offers computing educators a practical, evidence-based framework for determining which exam questions are safe to auto-grade with AI.
Authors: Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira
Paper: https://arxiv.org/abs/2607.02432v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
