[Submitted on 18 Sep 2025 (v1), last revised 22 Jul 2026 (this version, v2)]
Abstract:Deep reinforcement learning (DRL) methods, though powerful, often lack transparency, which limits their adoption in critical domains. We apply Self-Explaining Neural Networks (SENNs) to RL by parametrizing the policy of a PPO agent with a SENN, producing intrinsic local explanations, and propose a method for aggregating them into global explanations. We evaluate our approach on a mobile network resource allocation problem, our approach performs within a small margin of the state-of-the-art deep learning method and significantly outperforms the best deployed heuristic, while the extracted global explanations correlate strongly with DeepLift and InputXGradient, making SENNs a promising candidate for high-stakes RL.
Submission history
From: Konrad Nowosadko [view email]
[v1]
Thu, 18 Sep 2025 13:04:29 UTC (380 KB)
[v2]
Wed, 22 Jul 2026 11:06:50 UTC (249 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.