As code evolves, its test suite must evolve alongside it, but existing benchmarks for evaluating AI coding agents often treat tests and code changes in isolation using unverified static metadata. TestEvo-Bench fixes this by mining real commit histories across 152 open-source Java projects, creating executable tasks for both generating new tests and updating existing ones, evaluated with concrete metrics like pass rate and coverage. Being a "live" benchmark that's continuously updated helps reduce data leakage in model evaluations. This is directly useful for benchmarking and improving AI coding assistants like Claude Code or SWE-Agent on realistic software maintenance work.
Authors: Jiale Amber Wang, Kaiyuan Wang, Pengyu Nie
Paper: https://arxiv.org/abs/2607.02469v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
