View article

[PDF] from aclanthology.org

Creating a dataset for named entity recognition in the archaeology domain

Authors

Alex Brandsen, Suzan Verberne, Milco Wansleeben, Karsten Lambers

Publication date

2020/5

Conference

Proceedings of the Twelfth Language Resources and Evaluation Conference

Pages

4573-4577

Description

In this paper, we present the development of a training dataset for Dutch Named Entity Recognition (NER) in the archaeology domain. This dataset was created as there is a dire need for semantic search within archaeology, in order to allow archaeologists to find structured information in collections of Dutch excavation reports, currently totalling around 60,000 (658 million words) and growing rapidly. To guide this search task, NER is needed. We created rigorous annotation guidelines in an iterative process, then instructed five archaeology students to annotate a number of documents. The resulting dataset contains~ 31k annotations between six entity types (artefact, time period, place, context, species & material). The inter-annotator agreement is 0.95, and when we used this data for machine learning, we observed an increase in F1 score from 0.51 to 0.70 in comparison to a machine learning model trained on a dataset created in prior work. This indicates that the data is of high quality, and can confidently be used to train NER classifiers.

Total citations

Cited by 88

2021202220232024202515 20 27 21 5

Scholar articles

Creating a dataset for named entity recognition in the archaeology domain

A Brandsen, S Verberne, M Wansleeben, K Lambers - Proceedings of the Twelfth Language Resources and …, 2020

zproxy.org