- Publications
- Understanding Emergent Non-Verbal Communication in the Delta Force Competitive Video Game through Multimodal AI Analysis
Non-verbal communication plays a critical role in multiplayer games, players often rely on gestures, movement patterns, item interactions, and UI signals to communicate intent, negotiate cooperation willingness, and avoid conflict. May et al. [May et al. 2013] demonstrated how non-verbal game behaviors are effective at changing player behavior and task success.
Many studies have considered non-verbal communication in competitive esports, though the bulk of them have focused on MOBA games like League of Legends [Leavitt et al. 2016], Dota 2 [Wuertz et al. 2017], and Heroes of the Storm [Zheng and Farzan 2023; Zheng et al. 2023]. MOBA games include a ping feature, that places a mark on the map with intention to communicate warning or helping concepts. While players develop their own intent with these tools, ping largely sticks to the game designer’s intended meaning, and are somewhat more limited in scope than more expressive non-verbal communication. Another line of research has also leveraged MOBAs and online matchmaking that matches team mates with strangers to evaluate how ad hoc teams work [Kou and Gui 2014; Lee et al. 2025; Tan et al. 2022].
However, other genres, like extraction shooter and free-for-all battle royale games, have enabled even more ad hoc teaming, betrayals, and rescue behavior. In most shooter games, teams generally coordinate through verbal communication such as map callouts in Counter-Strike [Rusk et al. 2024; Tang et al. 2012]. In contrast to these prior studies, we endeavor to study non-verbal communication in the shooter game genre.
In our work, we analyze gameplay videos to extract non-verbal communication. Some past work has shown that training a model with data across a range of shooter games it is capable of learning some cross-game features due to the similarities within the genre.[Rašajski et al. 2024]. Perhaps the closest related work to our method is GELID, which performs a somewhat different set of operations, including segmenting based on captions, categorizing and grouping with image features and color analysis [Guglielmi et al. 2023]. Others have shown more general video game understanding using VLMs [Taesiri and Bezemer 2025]. We believe our pipeline to be novel, while being composed of relatively well-understood components.
Authors
Jinyuan Guo (Duke University)
Publication Date
Sunday, July 19, 2026
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.