CAp 2017 challenge: Twitter Named Entity Recognition

07/24/2017
by   Cédric Lopez, et al.
0

The paper describes the CAp 2017 challenge. The challenge concerns the problem of Named Entity Recognition (NER) for tweets written in French. We first present the data preparation steps we followed for constructing the dataset released in the framework of the challenge. We begin by demonstrating why NER for tweets is a challenging problem especially when the number of entities increases. We detail the annotation process and the necessary decisions we made. We provide statistics on the inter-annotator agreement, and we conclude the data description part with examples and statistics for the data. We, then, describe the participation in the challenge, where 8 teams participated, with a focus on the methods employed by the challenge participants and the scores achieved in terms of F_1 measure. Importantly, the constructed dataset comprising ∼6,000 tweets annotated for 13 types of entities, which to the best of our knowledge is the first such dataset in French, is publicly available at <http://cap2017.imag.fr/competition.html> .

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset