Decision-Aware Learning for Context-Dependent LQR: A Smart Predict-then-Optimize Approach
IEEE 65th Conference on Decision and Control (CDC)
TL;DR A smart predict-then-optimize approach that learns predictions tailored to downstream context-dependent LQR decisions.
Abstract & PDF
This paper develops a decision-aware learning framework for constrained linear quadratic regulator (LQR) design. It addresses the challenge of tuning unknown context-dependent cost weighting matrices by directly optimizing downstream control performance. When the prediction accuracy is poor, traditional two-stage methods, separating weight prediction and control optimization, are difficult to meet the demands of downstream high-precision tasks. Thus, we introduce the Smart Predict-then-Optimize (SPO) framework to the constrained LQR problem, which optimizes downstream control quality via direct decision regret minimization. First, we generalize the theory of the SPO framework to fit the LQR control problem. Second, we derive a convex surrogate for SPO loss, which retains interpretability via explicit weight prediction. Third, we propose an end-to-end training pipeline based on gradient-based optimization. Finally, we validate our approach via simulations, showing that our method reduces the cost ratio by approximately 34% compared to the Mean Squared Error (MSE) baseline, achieving significantly better downstream control performance.