Modeling and visual recognition of human actions and interactions

Ivan Laptev 1
1 WILLOW - Models of visual object recognition and scene understanding
CNRS - Centre National de la Recherche Scientifique : UMR8548, Inria Paris-Rocquencourt, DI-ENS - Département d'informatique de l'École normale supérieure
Abstract : This work addresses the problem of recognizing actions and interactions in realistic video settings such as movies and consumer videos. The first contribution of this thesis (Chapters 2 and 4) is concerned with new video representations for action recognition. We introduce local space-time descriptors and demonstrate their potential to classify and localize actions in complex settings while circumventing the difficult intermediate steps of person detection, tracking and human pose estimation. The material on bag-of-features action recognition in Chapter 2 is based on publications [L14, L22, L23] and is related to other work by the author [L6, L7, L8, L11, L12, L13, L16, L21]. The work on object and action localization in Chapter 4 is based on [L9, L10, L13, L15] and relates to [L1, L17, L19, L20]. The second contribution of this thesis is concerned with weakly-supervised action learning. Chap- ter 3 introduces methods for automatic annotation of action samples in video using readily-available video scripts. It addresses the ambiguity of action expressions in text and the uncertainty of tem- poral action localization provided by scripts. The material presented in Chapter 3 is based on publications [L4, L14, L18]. Finally Chapter 5 addresses interactions of people with objects and concerns modeling and recognition of object function. We exploit relations between objects and co-occurring human poses and demonstrate object recognition improvements using automatic pose estimation in challenging videos from YouTube. This part of the thesis is based on the publica- tion [L2] and relates to other work by the author [L3, L5].
Document type :
Habilitation à diriger des recherches
Liste complète des métadonnées

Cited literature [108 references]  Display  Hide  Download

https://tel.archives-ouvertes.fr/tel-01064540
Contributor : Minsu Cho <>
Submitted on : Tuesday, September 16, 2014 - 11:19:56 PM
Last modification on : Wednesday, January 30, 2019 - 11:07:44 AM
Document(s) archivé(s) le : Wednesday, December 17, 2014 - 11:21:29 AM

Identifiers

  • HAL Id : tel-01064540, version 1

Collections

Citation

Ivan Laptev. Modeling and visual recognition of human actions and interactions. Computer Vision and Pattern Recognition [cs.CV]. Ecole Normale Supérieure de Paris - ENS Paris, 2013. ⟨tel-01064540⟩

Share

Metrics

Record views

584

Files downloads

535