Predicting the future: A jointly learnt model for action anticipation
Gammulle, Harshala, Denman, Simon, Sridharan, Sridha, & Fookes, Clinton (2019) Predicting the future: A jointly learnt model for action anticipation. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV 2019). Institute of Electrical and Electronics Engineers Inc., United States of America, pp. 5561-5570.
|
Accepted Version
(PDF 5MB)
59982504. Available under License Creative Commons Attribution Non-commercial 4.0. |
Description
Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current state-of-the-art methods which first learn a model to predict future video features and then perform action anticipation using these features, the proposed framework jointly learns to perform the two tasks, future visual and temporal representation synthesis, and early action anticipation. The joint learning framework ensures that the predicted future embeddings are informative to the action anticipation task. Furthermore, through extensive experimental evaluations we demonstrate the utility of using both visual and temporal semantics of the scene, and illustrate how this representation synthesis could be achieved through a recurrent Generative Adversarial Network (GAN) framework. Our model outperforms the current state-of-the-art methods on multiple datasets: UCF101, UCF101-24, UT-Interaction and TV Human Interaction.
Impact and interest:
Citation counts are sourced monthly from Scopus and Web of Science® citation databases.
These databases contain citations from different subsets of available publications and different time periods and thus the citation count from each is usually different. Some works are not in either database and no count is displayed. Scopus includes citations from articles published in 1996 onwards, and Web of Science® generally from 1980 onwards.
Citations counts from the Google Scholar™ indexing service can be viewed at the linked Google Scholar™ search.
Full-text downloads:
Full-text downloads displays the total number of times this work’s files (e.g., a PDF) have been downloaded from QUT ePrints as well as the number of downloads in the previous 365 days. The count includes downloads for all files if a work has more than one.
| ID Code: | 200892 | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Item Type: | Chapter in Book, Report or Conference volume (Conference contribution) | ||||||||
| Series Name: | Proceedings of the IEEE International Conference on Computer Vision | ||||||||
| ORCID iD: |
|
||||||||
| Measurements or Duration: | 10 pages | ||||||||
| Event Title: | IEEE International Conference on Computer Vision | ||||||||
| Event Dates: | 2019-10-27 - 2019-11-02 | ||||||||
| Event Location: | Seoul, Korea, Republic of | ||||||||
| Additional URLs: | |||||||||
| DOI: | 10.1109/ICCV.2019.00566 | ||||||||
| ISBN: | 978-1-7281-4804-5 | ||||||||
| Pure ID: | 59982504 | ||||||||
| Divisions: | Past > Institutes > Institute for Future Environments Past > QUT Faculties & Divisions > Science & Engineering Faculty |
||||||||
| Funding Information: | This research was supported by an Australian Research Council (ARC) Linkage grant LP140100221 | ||||||||
| Funding: | |||||||||
| Copyright Owner: | Consult author(s) regarding copyright matters | ||||||||
| Copyright Statement: | 2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. | ||||||||
| Deposited On: | 09 Jun 2020 14:43 | ||||||||
| Last Modified: | 01 Oct 2026 00:19 |
Export: EndNote | Dublin Core | BibTeX
Repository Staff Only: item control page