Supervised Statistical Representations for Human Action Recognition in Video - TEL - Thèses en ligne Accéder directement au contenu
Thèse Année : 2012

Supervised Statistical Representations for Human Action Recognition in Video

Statistiques Supervisées pour la Reconnaissance d'Actions Humaines dans les Vidéos

Résumé

This thesis addresses the problem of human action recognition in realistic video data, such as movies and online videos. Automatic and accurate recognition of human actions in video is a fascinating capability. The potential applications range from surveillance and robotics to medical diagnosis, content-based video retrieval, and intelligent human- computer interfaces. The task is highly challenging due to the large variations in person appearances, dynamic backgrounds, view-point changes, lighting conditions, action styles and other factors. Statistical video representations based on local space-time features have been recently shown successful for action recognition in realistic scenarios. Their success can be at- tributed to the mild assumptions about the data and robustness to several variations in the video. Such representations, however, often encode videos by disordered collection of low-level primitives. This thesis extends current methods by developing more discrimi- native features and integrating additional supervision into Bag-of-Features based video representations, aiming to improve action recognition in unconstrained and challenging video data. We start by evaluating a range of available local space-time feature detectors and descriptors under the standard Bag-of-Features framework. We then propose to improve the basic Bag-of-Features model by integrating additional supervision in the form of non-local region-level information. We further investigate an attribute-based representation, wherein the attributes range from objects (e.g., car, chair, table, etc.) to human poses and actions. We demonstrate that such representation captures high-level information in video, and provides complementary information to the low-level features. We finally propose a novel local representation for human action recognition in video, denoted as Actlets. Actlets are body part detectors undergoing characteristic motion patterns. We train Actlets using a large synthetic video dataset of rendered avatars and demonstrate the advantages of Actlets for action recognition in realistic data. All methods proposed and developed in this thesis represent alternative ways of construct- ing supervised video representations and demonstrate improvements of human action recognition in realistic settings.
Fichier principal
Vignette du fichier
2012thesisUllah.pdf (21.48 Mo) Télécharger le fichier
Loading...

Dates et versions

tel-01063349 , version 1 (11-09-2014)

Identifiants

  • HAL Id : tel-01063349 , version 1

Citer

Muhammad Muneeb Ullah. Supervised Statistical Representations for Human Action Recognition in Video. Computer Vision and Pattern Recognition [cs.CV]. Université Européenne de Bretagne, 2012. English. ⟨NNT : ⟩. ⟨tel-01063349⟩
281 Consultations
233 Téléchargements

Partager

Gmail Facebook X LinkedIn More