Our paper, co-authored by Ignacio Castro, “CG-TTRL: Context-guided Test-time Reinforcement Learning for On-device Large Language Models,” has been accepted at Transactions of the Association for Computational Linguistics and will be presented at EMNLP 2026. The paper introduces CG-TTRL, which integrates contextual guidance into test-time reinforcement learning to improve pseudo-label accuracy and regulate exploration during adaptation. It also proposes an efficient context-selection mechanism for on-device use. Across mathematical and scientific QA benchmarks, CG-TTRL outperforms standard TTRL, achieving up to a 7% relative accuracy improvement and stronger performance after only a few test-time training steps.

Paper at Transactions of ACL