Combining Large Language Models and Gradient-Free Optimization for Automatic Control Policy Synthesis

Combining Large Language Models and Gradient-Free Optimization for Automatic Control Policy Synthesis

Oct 1, 2025·
Carlo Bosio
Matteo Guarrera
Matteo Guarrera
,
Alberto Sangiovanni-Vincentelli
,
Mark W. Mueller
· 0 min read
Abstract
Control policies can be written as programs, which makes them readable and verifiable, but searching over programs is hard. We let a language model generate the symbolic structure of the policy and hand the continuous parameters to a gradient-free optimizer, so the two halves of the problem are solved by the tools suited to them. The combination reaches target performance in over an order of magnitude fewer search steps than the baselines, and scales to full quadruped locomotion. Qualcomm Innovation Fellowship finalist, 2025.
Type
Publication
arXiv preprint arXiv:2510.00373