LAME: Layout-Aware Metadata Extraction Approach for Research Articles

  • Choi, Jongyun; 
  • Kong, Hyesoo; 
  • Yoon, Hwamook; 
  • Oh, Heungseon; 
  • Jung, Yuchul
Citations

WEB OF SCIENCE

1

초록

The volume of academic literature, such as academic conference papers and journals, has increased rapidly worldwide, and research on metadata extraction is ongoing. However, high-performing metadata extraction is still challenging due to diverse layout formats according to journal publishers. To accommodate the diversity of the layouts of academic journals, we propose a novel LAyout-aware Metadata Extraction (LAME) framework equipped with the three characteristics (e.g., design of automatic layout analysis, construction of a large meta-data training set, and implementation of metadata extractor). In the framework, we designed an automatic layout analysis using PDFMiner. Based on the layout analysis, a large volume of metadata-separated training data, including the title, abstract, author name, author affiliated organization, and keywords, were automatically extracted. Moreover, we constructed a pre-trained model, Layout-MetaBERT, to extract the metadata from academic journals with varying layout formats. The experimental results with our metadata extractor exhibited robust performance (Macro-F1, 93.27%) in metadata extraction for unseen journals with different layout formats.

키워드

Automatic layout analysis; layout-MetaBERT; metadata extrac-tion; research article
제목
LAME: Layout-Aware Metadata Extraction Approach for Research Articles
저자
Choi, Jongyun; Kong, Hyesoo; Yoon, Hwamook; Oh, Heungseon; Jung, Yuchul
DOI
10.32604/cmc.2022.025711
발행일
2022-03
유형
Article
저널명
Computers, Materials and Continua
권
72
호
2
페이지
4019 ~ 4037