[2026年4月13日提出 (v1)、最終改訂 2026年7月29日 (本バージョン, v2)]
Abstract:The integration of frame-based RGB cameras with event streams constitutes a promising paradigm for robust object detection under challenging dynamic conditions. Nevertheless, effectively modeling intricate multi-modal interactions and reconciling the semantic heterogeneity between RGB and event data remain formidable challenges for high-precision detection. In this paper, we present Hyper-FEOD, a novel high-performance detection framework that synergistically strengthens cross-modal representation learning through two core hypergraph-driven components. Specifically, we first design a Sparse Hypergraph-enhanced Cross-Modal Fusion (SHCF) module that exploits event activity cues to identify motion-critical sparse tokens and performs high-order relational reasoning through hypergraph modeling. This design effectively captures intricate high-order dependencies and rich contextual correlations across modalities. Second, we develop a Fine-Grained Mixture-of-Experts (FG-MoE) module tailored to handle the heterogeneous semantic demands arising from distinct visual regions. By deploying specialized hypergraph experts with varying hyperedge connectivity pattern and incorporating a spatial gating mechanism, FG-MoE adaptively routes features to enable precise enhancement at target regions. Coupled with an auxiliary router loss, the proposed framework ensures stable end-to-end training and optimal feature refinement. Comprehensive experiments conducted on widely-adopted RGB-Event benchmarks show that Hyper-FEOD delivers superior detection performance and consistently outperforms existing state-of-the-art approaches by a notable margin.
Submission history
From: Wei Bao [view email]
[v1]
Mon, 13 Apr 2026 07:56:16 UTC (1,808 KB)
[v2]
Wed, 29 Jul 2026 02:42:12 UTC (2,725 KB)
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.