Julius Frost

Author directory

2026

Code-editing language-model agents can change training choices beyond the hyper- parameter grids tested in automated essay scoring. Comparing these approaches requires distinguishing gains from a broader search space from evidence of a better search procedure. We compare code-editing agents with grid-restricted search, including random search, in two studies on ASAP-AES with nominally matched 12-hour search budgets. Code-editing produced the configuration with the highest test quadratic-weighted kappa (QWK) point estimate in each primary comparison. However, random search over a grid built afterwards around the Study 2 agent’s backbone, input length and head rule recovered most of its gain over BERT. Requiring a minimum validation-score improvement to accept a trial also left the accepted configuration’s validation score below the highest recorded valid validation score in every run that accepted a trial under this rule. The primary comparisons use one search run per condition, and some test folds reused essays involved in configuration selection, so these results compare selected configurations without establishing search-procedure superiority. These findings motivate reporting the available search choices and both accepted and best-scoring trials, and evaluating search procedures through repeated runs on data independent of configuration selection.
Search
Co-authors
    Venues
    Fix author