TMMF: Temporal Multi-modal Fusion for Single-Stage Continuous Gesture Recognition

, , , & (2021) TMMF: Temporal Multi-modal Fusion for Single-Stage Continuous Gesture Recognition. IEEE Transactions on Image Processing, 30, Article number: 9528969 7689-7701.

View at publisher

Description

Gesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods have focused on recognising isolated gestures, and existing continuous gesture recognition methods are limited to two-stage approaches where independent models are required for detection and classification, with the performance of the latter being constrained by detection performance. In contrast, we introduce a single-stage continuous gesture recognition framework, called Temporal Multi-Modal Fusion (TMMF), that can detect and classify multiple gestures in a video via a single model. This approach learns the natural transitions between gestures and non-gestures without the need for a pre-processing segmentation step to detect individual gestures. To achieve this, we introduce a multi-modal fusion mechanism to support the integration of important information that flows from multi-modal inputs, and is scalable to any number of modes. Additionally, we propose Unimodal Feature Mapping (UFM) and Multi-modal Feature Mapping (MFM) models to map uni-modal features and the fused multi-modal features respectively. To further enhance performance, we propose a mid-point based loss function that encourages smooth alignment between the ground truth and the prediction, helping the model to learn natural gesture transitions. We demonstrate the utility of our proposed framework, which can handle variable-length input videos, and outperforms the state-of-the-art on three challenging datasets: EgoGesture, IPN hand and ChaLearn LAP Continuous Gesture Dataset (ConGD). Furthermore, ablation experiments show the importance of different components of the proposed framework.

Impact and interest:

45 citations in Scopus
40 citations in Web of Science®
Search Google Scholar™

Citation counts are sourced monthly from Scopus and Web of Science® citation databases.

These databases contain citations from different subsets of available publications and different time periods and thus the citation count from each is usually different. Some works are not in either database and no count is displayed. Scopus includes citations from articles published in 1996 onwards, and Web of Science® generally from 1980 onwards.

Citations counts from the Google Scholar™ indexing service can be viewed at the linked Google Scholar™ search.

ID Code: 214058
Item Type: Contribution to Journal (Journal Article)
Refereed: Yes
ORCID iD:
Gammulle, Harshalaorcid.org/0000-0003-0670-0406
Denman, Simonorcid.org/0000-0002-0983-5480
Sridharan, Sridhaorcid.org/0000-0003-4316-9001
Fookes, Clintonorcid.org/0000-0002-8515-6324
Measurements or Duration: 13 pages
Keywords: Feature extraction, Gesture Recognition, Gesture recognition, Semantics, Solid modeling, Spatio-temporal Representation Learning, Streaming media, Temporal Convolution Networks, Three-dimensional displays, Visualization
DOI: 10.1109/TIP.2021.3108349
ISSN: 1057-7149
Pure ID: 99924855
Divisions: Current > Research Centres > Centre for Data Science
Current > QUT Faculties and Divisions > Faculty of Science
Current > QUT Faculties and Divisions > Faculty of Engineering
Current > Schools > School of Electrical Engineering & Robotics
Funding:
Copyright Owner: 2021 IEEE
Copyright Statement: This work is covered by copyright. Unless the document is being made available under a Creative Commons Licence, you must assume that re-use is limited to personal use and that permission from the copyright owner must be obtained for all other uses. If the document is available under a Creative Commons License (or other specified license) then refer to the Licence for details of permitted re-use. It is a condition of access that users recognise and abide by the legal requirements associated with these rights. If you believe that this work infringes copyright please provide details by email to qut.copyright@qut.edu.au
Deposited On: 21 Oct 2021 08:23
Last Modified: 27 Sep 2026 05:37