View a PDF of the paper titled CriterAlign: Criterion-Centric Rationale Alignment for Code Desire Judging, by Zhenyu Li and three different authors
View PDF
HTML (experimental)
Summary:Pairwise human choice prediction is central to evaluating code-generation methods, the place high quality usually depends upon task-specific trade-offs past practical correctness. Whereas rubric-based LLM judges enhance interpretability by decomposing analysis into express standards, most current pipelines stay pointwise: they rating every response independently and derive preferences by evaluating aggregated scores. We present that this design is poorly matched to pairwise code choice prediction and may underperform a powerful monolithic decide. We suggest CriterAlign, a criterion-centric framework that adapts rubric-based judging to pairwise choice analysis by way of direct criterion-level pairwise judgments, tie-driven criterion refinement, swap-consistency filtering, and ultimate pairwise synthesis. We additional introduce Human-Desire-Aligned Steerage (HPAG), synthesized offline from coaching examples by extracting recurring rationale gaps between human preferences and monolithic decide predictions, and injected into the criterion generator, criterion decide, and ultimate decide. On BigCodeReward, CriterAlign improves a Qwen2.5-VL-32B monolithic decide from 60.4% to 66.3% accuracy, with ablations confirming the contributions of pairwise criterion design and HPAG.
Submission historical past
From: Zhenyu Li [view email] [v1]
Tue, 19 Could 2026 10:59:19 UTC (381 KB)
[v2]
Thu, 9 Jul 2026 06:11:29 UTC (381 KB)
![[2605.19665] CriterAlign: Criterion-Centric Rationale Alignment for Code Desire Judging [2605.19665] CriterAlign: Criterion-Centric Rationale Alignment for Code Desire Judging](http://arxiv.org/static/browse/0.3.4/images/arxiv-logo-fb.png)
