Membership Inference Attacks Against Tabular Synthetic Data Under Conservative Threat Models

Joshua Ward
Ph.D., 2026
CHENG, GUANG
Tabular synthetic data has emerged as a promising mechanism for privacy-preserving data sharing, enabling organizations in sensitive domains such as healthcare and finance to publish data products that preserve statistical utility while limiting risk to individuals in their training sets. The empirical privacy of synthetic data releases is most rigorously assessed through Membership Inference Attacks (MIAs). However, existing MIAs for tabular synthetic data share two critical limitations: they frequently assume unrealistic adversarial knowledge of the generative model's implementation that would be withheld under responsible release practices, and they operate exclusively over the Cartesian feature space of a single observation, failing to capture vulnerabilities introduced by Large Language Models and relational data generators whose inductive biases differ fundamentally from earlier generative models. This dissertation addresses both limitations through three contributions. The first introduces the Generative Likelihood Ratio Attack (Gen-LRA), a No-Box MIA grounded in an influence function framework, accompanied by theoretical results characterizing what the attack measures and when it succeeds. Gen-LRA achieves state-of-the-art performance across a benchmark spanning 35~datasets, 9~generative architectures, and over 1,500 synthetic dataset configurations. The second contribution studies privacy risks specific to Large Language Model (LLM)-based tabular generators, which encode tabular rows as strings and generate records via autoregressive text production—a paradigm invisible to feature-space attacks. We introduce LevAtt, a No-Box MIA that targets memorized sequences of numeric digits in synthetic string representations, achieving near-perfect membership classification on some state-of-the-art generators. The third contribution formalizes membership inference for the multi-table synthetic data setting, where a user's information is distributed across multiple interconnected tables in a relational database. We show that single-table MIAs systematically underestimate user-level privacy leakage in this setting, and introduce MT-MIA, a No-Box attack that represents each user as a heterogeneous subgraph and leverages Graph Neural Networks to perform user-level membership inference. Together, these three works advance a principled framework for realistic, architecture-aware privacy auditing of tabular synthetic data.
2026