ASET: Ad-hoc Structured Exploration of Text Collections [Extended Abstract]

03/09/2022
by   Benjamin Hättasch, et al.
3

In this paper, we propose a new system called ASET that allows users to perform structured explorations of text collections in an ad-hoc manner. The main idea of ASET is to use a new two-phase approach that first extracts a superset of information nuggets from the texts using existing extractors such as named entity recognizers and then matches the extractions to a structured table definition as requested by the user based on embeddings. In our evaluation, we show that ASET is thus able to extract structured data from real-world text collections in high quality without the need to design extraction pipelines upfront.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset