[Submitted on 6 Jun 2025 (v1), last revised 22 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:A key tool in developing safe AI models is \emph{data auditing}, i.e., using statistical tools to determine whether harmful content may have been used in the training data of a black-box model. Unfortunately, most \emph{membership inference attacks} (MIAs) used to perform this type of auditing themselves assume \emph{access} to examples of harmful content from the same distribution as the query data. In real-world auditing scenarios, auditors often face legal and ethical restrictions preventing them from accessing a representative set of samples of harmful content to train MIA models effectively. We abstract and formalize this setting into a new data access model, the ``unseen class'' setting, and show that the state of the art MIAs fail due to the lack of access to the full target distribution. We show in this setting, \emph{quantile regression attacks} outperform approaches typically considered to be SoTA. We demonstrate this both empirically and theoretically, showing that quantile regression attacks achieve up to \textbf{11$\times$ the TPR} of shadow model-based approaches in practice, and providing a theoretical model that outlines the generalization properties required for this approach to succeed. Our work identifies an important failure mode in existing MIAs and provides a cautionary tale for practitioners that aim to directly use existing tools for real-world applications of AI safety.

Submission history

From: Pratiksha Thaker [view email]
[v1] Fri, 6 Jun 2025 19:27:52 UTC (2,290 KB)
[v2] Sat, 25 Oct 2025 21:13:10 UTC (2,376 KB)
[v3] Wed, 22 Jul 2026 01:30:58 UTC (2,804 KB)