Making Use of Existing Lexical Resources to Build a Verbnet like Classification of French Verbs

Ingrid Falk 1
1 SYNALP - Natural Language Processing : representations, inference and semantics
LORIA - NLPKD - Department of Natural Language Processing & Knowledge Discovery
Abstract : Classifications which group together verbs and a set of shared syntactic and semantic properties have proven useful both in linguistics and in Natural Language Processing tasks. However, for French this type of classifications is not available in a format suitable for automated processing. In addition, most existing approaches for automatically acquiring verb classes fail to associate the verb classes produced with an explicit characterisation of the syntactic and semantic properties shared by the class members. Here we propose a novel approach to verb clustering which addresses these shortcomings. We classify French verbs using two clustering methods, a symbolic method called Formal Concept Analysis (FCA) and a probabilistic neural clustering method called Incremental Growing Neural Gas with Feature Maximisation (IGNGF). The obtained classes group together verbs, subcategorisation frames and thematic grids. We apply this approach to French data consisting of roughly 4000 verbs and 350 subcategorisation frames, and evaluate both the clusters obtained (i.e., verb classes) and the features labeling each cluster (i.e., syntactic frames and thematic grids). The results suggest that both classification methods can be used to bootstrap a Verbnet style classification for French such that the verb classes it contains (i) are reasonably clean and (ii) associate verbs with partial information about subcategorisation frames and thematic grids. The obtained classifications are complementary. While the FCA classification better represents verb polysemy (better F-measure and recall compared to reference data) the IGNGF classification performed better with respect to the produced verb classes and when used in a task based evaluation.
Document type :
Theses
Complete list of metadatas

Cited literature [3 references]  Display  Hide  Download

https://tel.archives-ouvertes.fr/tel-00714737
Contributor : Ingrid Falk <>
Submitted on : Thursday, July 5, 2012 - 2:47:32 PM
Last modification on : Tuesday, December 18, 2018 - 4:38:01 PM
Long-term archiving on : Saturday, October 6, 2012 - 2:41:34 AM

Identifiers

  • HAL Id : tel-00714737, version 1

Citation

Ingrid Falk. Making Use of Existing Lexical Resources to Build a Verbnet like Classification of French Verbs. Computation and Language [cs.CL]. Université Nancy II, 2012. English. ⟨tel-00714737⟩

Share

Metrics

Record views

767

Files downloads

640