[Submitted on 10 Jun 2025 (v1), last revised 29 Jul 2026 (this version, v3)]

View PDF HTML (experimental)

Abstract:Large language models have moved from advising on offensive security to autonomously conducting it. A growing literature presents agents that execute reconnaissance, exploitation, and privilege escalation against real or simulated targets. Such an agent is a deployable, re-pointable capability whose harm potential scales with the underlying model. The papers that introduce it therefore carry an unusual ethical burden, which top security venues have begun to encode as hard policy in 2025-2026 ethics-section mandates. We present a systematic, reproducible audit of ethics-and-risk reporting in this literature. From a pre-registered Scopus query (Channel A, n=35) plus a reproducible forward-snowball of two seed papers via the Semantic Scholar citation graph (Channel B, n=19, all Scopus-absent) we assemble 54 autonomous offensive-LLM penetration-testing prototypes (2023-2026). We score each against a nine-dimension instrument derived both top-down from the Menlo Report, and bottom-up from the 2025-26 venue mandates. Our central result is a recognition-without-mitigation gap: dual-use risk is reported as recognized in 39% of papers but a concrete mitigation is reported in only 7%, roughly a 5:1 gap. Of the papers, 17% are anti-safeguard, reporting the defeat of model safety controls with no countermeasure. The near-universal safeguards reported are research-integrity controls that protect the experiment, not the public; institutional-review (2%) and coordinated-disclosure (6%) practice is almost absent and confined to Channel B. Measured against the new mandates, the corpus defines a pre-regulation baseline: current practice does not meet the substantive requirements. We argue this audit is itself defensive intelligence on the offensive-agent ecosystem, and we distill a minimal containment checklist for future work.

Submission history

From: Andreas Happe [view email]
[v1] Tue, 10 Jun 2025 11:11:55 UTC (19 KB)
[v2] Sun, 31 May 2026 08:37:46 UTC (16 KB)
[v3] Wed, 29 Jul 2026 07:05:45 UTC (46 KB)