An Ensemble Approach to Acronym Extraction using Transformers

Prashant Sharma; Hadeel Saadany; Leonardo Zilio; Diptesh Kanojia; Constantin Orăsan

doi:10.48550/arxiv.2201.03026

Back

Preprint

An Ensemble Approach to Acronym Extraction using Transformers

Prashant Sharma, Hadeel Saadany, Leonardo Zilio, Diptesh Kanojia and Constantin Orăsan

arXiv (Cornell University)

09/01/2022

DOI: https://doi.org/10.48550/arxiv.2201.03026

Abstract

Computer Science - Computation and Language

Acronyms are abbreviated units of a phrase constructed by using initial components of the phrase in a text. Automatic extraction of acronyms from a text can help various Natural Language Processing tasks like machine translation, information retrieval, and text summarisation. This paper discusses an ensemble approach for the task of Acronym Extraction, which utilises two different methods to extract acronyms and their corresponding long forms. The first method utilises a multilingual contextual language model and fine-tunes the model to perform the task. The second method relies on a convolutional neural network architecture to extract acronyms and append them to the output of the previous method. We also augment the official training dataset with additional training samples extracted from several open-access journals to help improve the task performance. Our dataset analysis also highlights the noise within the current task dataset. Our approach achieves the following macro-F1 scores on test data released with the task: Danish (0.74), English-Legal (0.72), English-Scientific (0.73), French (0.63), Persian (0.57), Spanish (0.65), Vietnamese (0.65). We release our code and models publicly.

Metrics

13 Record Views

Details

Title: An Ensemble Approach to Acronym Extraction using Transformers
Creators: Prashant Sharma - Hitachi CRL, Japan
Hadeel Saadany - University of Surrey, School of Literature and Languages
Leonardo Zilio - University of Surrey, School of Literature and Languages
Diptesh Kanojia - University of Surrey, School of Literature and Languages
Constantin Orăsan - University of Surrey, School of Literature and Languages
Publication Details: arXiv (Cornell University)
Identifiers: 99783722502346
Academic Unit: School of Computer Science and Electronic Engineering; School of Literature and Languages
Language: English
Resource Type: Preprint

An Ensemble Approach to Acronym Extraction using Transformers

Abstract

Metrics

Details

Usage Policy