[Submitted on 13 Feb 2026 (v1), last revised 22 Jul 2026 (this version, v4)]

Authors:Rong Fu, Xiaowen Ma, Kun Liu, Wangyu Wu, Ziyu Kong, Jia Yee Tan, Tailong Luo, Xianda Li, Yongtai Liu, Youjin Wang, Simon Fong

View PDF HTML (experimental)

Abstract:Deploying expressive learning models directly on programmable dataplanes promises line-rate, low-latency traffic analysis but remains hindered by strict hardware constraints and the need for predictable, auditable behavior. Chimera introduces a principled framework that maps attention-oriented neural computations and symbolic constraints onto dataplane primitives, enabling trustworthy inference within the match-action pipeline. Chimera combines a kernelized, linearized attention approximation with a two-layer key-selection hierarchy and a cascade fusion mechanism that enforces hard symbolic guarantees while preserving neural expressivity. The design includes a hardware-aware mapping protocol and a two-timescale update scheme that together permit stable, line-rate operation under realistic dataplane budgets. The paper presents the Chimera architecture, a hardware mapping strategy, and empirical evidence showing that neuro-symbolic attention primitives can achieve high-fidelity inference within the resource envelope of commodity programmable switches.

Submission history

From: Rong Fu [view email]
[v1] Fri, 13 Feb 2026 11:55:06 UTC (12,973 KB)
[v2] Wed, 4 Mar 2026 03:21:34 UTC (12,973 KB)
[v3] Tue, 21 Apr 2026 14:14:01 UTC (1,323 KB)
[v4] Wed, 22 Jul 2026 02:11:13 UTC (1,322 KB)