Supervised Statistical Representations for Human Action Recognition in Video

Muhammad Muneeb Ullah

Thèse Année : 2012

Supervised Statistical Representations for Human Action Recognition in Video

Statistiques Supervisées pour la Reconnaissance d'Actions Humaines dans les Vidéos

(1)

Muhammad Muneeb Ullah

Fonction : Auteur

Models of visual object recognition and scene understanding

Résumé

This thesis addresses the problem of human action recognition in realistic video data, such as movies and online videos. Automatic and accurate recognition of human actions in video is a fascinating capability. The potential applications range from surveillance and robotics to medical diagnosis, content-based video retrieval, and intelligent human- computer interfaces. The task is highly challenging due to the large variations in person appearances, dynamic backgrounds, view-point changes, lighting conditions, action styles and other factors. Statistical video representations based on local space-time features have been recently shown successful for action recognition in realistic scenarios. Their success can be at- tributed to the mild assumptions about the data and robustness to several variations in the video. Such representations, however, often encode videos by disordered collection of low-level primitives. This thesis extends current methods by developing more discrimi- native features and integrating additional supervision into Bag-of-Features based video representations, aiming to improve action recognition in unconstrained and challenging video data. We start by evaluating a range of available local space-time feature detectors and descriptors under the standard Bag-of-Features framework. We then propose to improve the basic Bag-of-Features model by integrating additional supervision in the form of non-local region-level information. We further investigate an attribute-based representation, wherein the attributes range from objects (e.g., car, chair, table, etc.) to human poses and actions. We demonstrate that such representation captures high-level information in video, and provides complementary information to the low-level features. We finally propose a novel local representation for human action recognition in video, denoted as Actlets. Actlets are body part detectors undergoing characteristic motion patterns. We train Actlets using a large synthetic video dataset of rendered avatars and demonstrate the advantages of Actlets for action recognition in realistic data. All methods proposed and developed in this thesis represent alternative ways of construct- ing supervised video representations and demonstrate improvements of human action recognition in realistic settings.

Mots clés

computer vision action recognition

Domaines

Vision par ordinateur et reconnaissance de formes [cs.CV]

Fichier principal

2012thesisUllah.pdf (21.48 Mo)

Minsu Cho : Connectez-vous pour contacter le contributeur

https://theses.hal.science/tel-01063349

Soumis le : jeudi 11 septembre 2014-18:54:30

Dernière modification le : vendredi 19 avril 2024-16:18:57

Archivage à long terme le : vendredi 12 décembre 2014-11:01:20

Dates et versions

tel-01063349 , version 1 (11-09-2014)

Identifiants

HAL Id : tel-01063349 , version 1

Citer

Muhammad Muneeb Ullah. Supervised Statistical Representations for Human Action Recognition in Video. Computer Vision and Pattern Recognition [cs.CV]. Université Européenne de Bretagne, 2012. English. ⟨NNT : ⟩. ⟨tel-01063349⟩

Exporter

BibTeX XML-TEI Dublin Core DC Terms EndNote DataCite

Collections

ENS-PARIS CNRS INRIA THESES-ENS INRIA2 PSL

281 Consultations

233 Téléchargements

Supervised Statistical Representations for Human Action Recognition in Video

Statistiques Supervisées pour la Reconnaissance d'Actions Humaines dans les Vidéos

Résumé

Mots clés

Domaines

Dates et versions

Identifiants

Citer

Exporter

Collections

Partager