[Submitted on 6 Jul 2026 (v1), last revised 28 Jul 2026 (this version, v2)]

View PDF HTML (experimental)

Abstract:Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback-Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to approximation error and a statistical error governed by the complexity of the ratio hypothesis class. When full coverage fails, we introduce coverage-stopped FORE, which targets the discounted occupancy accumulated before the first uncovered state-action pair and yields a conservative lower bound on target-policy value for nonnegative rewards. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.

Submission history

From: Lars Van Der Laan [view email]
[v1] Mon, 6 Jul 2026 17:53:32 UTC (122 KB)
[v2] Tue, 28 Jul 2026 16:20:46 UTC (166 KB)