Fine-grained action segmentation using the semi-supervised action GAN
Gammulle, Harshala, Denman, Simon, Sridharan, Sridha, & Fookes, Clinton (2020) Fine-grained action segmentation using the semi-supervised action GAN. Pattern Recognition, 98, Article number: 107039.
|
Accepted Version
(PDF 6MB)
59982746. Available under License Creative Commons Attribution Non-commercial No Derivatives 2.5. |
Description
In this paper we address the problem of continuous fine-grained action segmentation, in which multiple actions are present in an unsegmented video stream. The challenge for this task lies in the need to represent the hierarchical nature of the actions and to detect the transitions between actions, allowing us to localise the actions within the video effectively. We propose a novel recurrent semi-supervised Generative Adversarial Network (GAN) model for continuous fine-grained human action segmentation. Temporal context information is captured via a novel Gated Context Extractor (GCE) module, composed of gated attention units, that directs the queued context information through the generator model, for enhanced action segmentation. The GAN is made to learn features in a semi-supervised manner, enabling the model to perform action classification jointly with the standard, unsupervised, GAN learning procedure. We perform extensive evaluations on different architectural variants to demonstrate the importance of the proposed network architecture, and show that it is capable of outperforming current state-of-the-art on three challenging datasets: 50 Salads, MERL Shopping and Georgia Tech Egocentric Activities dataset.
Impact and interest:
Citation counts are sourced monthly from Scopus and Web of Science® citation databases.
These databases contain citations from different subsets of available publications and different time periods and thus the citation count from each is usually different. Some works are not in either database and no count is displayed. Scopus includes citations from articles published in 1996 onwards, and Web of Science® generally from 1980 onwards.
Citations counts from the Google Scholar™ indexing service can be viewed at the linked Google Scholar™ search.
Full-text downloads:
Full-text downloads displays the total number of times this work’s files (e.g., a PDF) have been downloaded from QUT ePrints as well as the number of downloads in the previous 365 days. The count includes downloads for all files if a work has more than one.
| ID Code: | 200897 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Item Type: | Contribution to Journal (Journal Article) | ||||||||
| Refereed: | Yes | ||||||||
| ORCID iD: |
|
||||||||
| Measurements or Duration: | 12 pages | ||||||||
| Keywords: | Context modelling, Generative adversarial networks, Human action segmentation | ||||||||
| DOI: | 10.1016/j.patcog.2019.107039 | ||||||||
| ISSN: | 0031-3203 | ||||||||
| Pure ID: | 59982746 | ||||||||
| Divisions: | Current > Research Centres > Centre for Data Science Current > Research Centres > Centre for Biomedical Technologies Past > Institutes > Institute for Future Environments Past > QUT Faculties & Divisions > Science & Engineering Faculty Current > QUT Faculties and Divisions > Faculty of Science Current > QUT Faculties and Divisions > Faculty of Engineering Current > Schools > School of Electrical Engineering & Robotics Current > Research Centres > Centre for Tropical Crops and Biocommodities |
||||||||
| Funding Information: | Professor Sridha Sridharan has a B.Sc. (Electrical Engineering) degree and obtained a M.Sc. (Communication Engineering) degree from the University of Manchester, UK and a Ph.D. degree from University of New South Wales, Australia. He is currently with the Queensland University of Technology (QUT) where he is a Professor in the School Electrical Engineering and Computer Science. Professor Sridharan is the Leader of the Research Program in Speech, Audio, Image and Video Technologies (SAIVT) at QUT, with strong focus in the areas of computer vision, pattern recognition and machine learning. He has published over 500 papers consisting of publications in journals and in refereed international conferences in the areas of Image and Speech technologies during the period 1990–2016. During this period he has also graduated 60 Ph.D. students in the areas of Image and Speech technologies. Prof Sridharan has also received a number of research grants from various funding bodies including Commonwealth competitive funding schemes such as the Australian Research Council (ARC) and the National Security Science and Technology (NSST) unit. Several of his research outcomes have been commercialised. | ||||||||
| Copyright Owner: | 2019 Elsevier Ltd. | ||||||||
| Copyright Statement: | This work is covered by copyright. Unless the document is being made available under a Creative Commons Licence, you must assume that re-use is limited to personal use and that permission from the copyright owner must be obtained for all other uses. If the document is available under a Creative Commons License (or other specified license) then refer to the Licence for details of permitted re-use. It is a condition of access that users recognise and abide by the legal requirements associated with these rights. If you believe that this work infringes copyright please provide details by email to qut.copyright@qut.edu.au | ||||||||
| Deposited On: | 09 Jun 2020 15:03 | ||||||||
| Last Modified: | 29 Sep 2026 08:38 |
Export: EndNote | Dublin Core | BibTeX
Repository Staff Only: item control page