Graph-Informed Sequential Decision Making

Shuang Wu
Ph.D., 2025
AMINI, ARASH A
This dissertation studies graph-informed sequential decision making, where graphs enter the bandit problem either as data—actions, contexts, rewards—or as structure that couples decisions, observations and agents. Algorithms that leverage graph priors to accelerate learning under limited feedback are developed with comprehensive theoretical analysis in this work.Part I introduces the backgrounds of the models and concepts in both statistical sequential decision making and machine learning on graphs. Chapter 1 elucidates bandit problems and algorithms, while Chapter 2 introduces graph learning models, from graph spectral theory to graph deep learning.Part II presents the sequential decision making problems where the graph serves as data and our proposed algorithm, GNN-TS. Chapter 3 introduces two online problems in which each round presents a graph and only bandit feedback is revealed. First, in online graph selection, actions are full graphs (e.g., molecules, program graphs); the learner selects a graph and observes a noisy payoff. This framing highlights the need for graph representations and calibrated exploration at decision time. Second, in online graph classification, each input is a graph and the learner must output a multi-class label with only action-dependent bandit feedback, linking the problem to multinomial logistic bandits over graph encodings. Chapter 4 presents the first project, graph neural Thompson Sampling, which pairs graph neural encoders with Thompson sampling as exploration rules. Theoretically, its performance is characterized via an effective-dimension parameter of a graph neural tangent kernel, yielding sublinear regret of order O˜( ˜d T1/2 ).Part III presents the sequential decision making problems where the graph serves as structure and our contribution in novel algorithms and problem unification. Chapter 5 first introduces the problems that decisions are coupled by a known graph. A Laplacian-regularized linear unified view that fuses content features with structural smoothness is presented for this problem. The second bandit problem is under the multi-agent setting, with a set of wide applications in interactive systems (recommendation, advertising, personalization). The second project is detailed in Chapter 6. The Laplacian kernelized bandit algorithms are proposed by inducing a multi-user kernel and Gaussian process style posterior, with confidence bounds derived from a bias–noise decomposition and regret governed by an effective dimension. The proposals are applied into a generalized design of the gang-of-bandits problem and competitive in both preferred regime and the other regimes. Part IV introduces the future works and the conclusion on the study about sequential decision making with graph information. Chapter 7 presents the ongoing works and future investigation on this research topic. A novelty algorithm, GCN-Logistic bandit, is proposed as the ongoing project, for online graph classification with bandit feedback. A foundation work on random graph generation model in sequential decision making as well as the innovation for online recommendation with decision making on the item-user graph, are introduced as future works.
2025