Fine-tuning BERT Models for Keyphrase Extraction in Scientific Articles

Fine-tuning BERT Models for Keyphrase Extraction in Scientific Articles

초록

Despite extensive research, performance enhancement of keyphrase (KP) extraction remains a challenging problem in modern informatics. Recently, deep learning-based supervised approaches have exhibited state-of-the-art accuracies with respect to this problem, and several of the previously proposed methods utilize Bidirectional Encoder Representations from Transformers (BERT)-based language models. However, few studies have investigated the effective application of BERT-based fine-tuning techniques to the problem of KP extraction. In this paper, we consider the aforementioned problem in the context of scientific articles by investigating the fine-tuning characteristics of two distinct BERT models — BERT (i.e., base BERT model by Google) and SciBERT (i.e., a BERT model trained on scientific text). Three different datasets (WWW, KDD, and Inspec) comprising data obtained from the computer science domain are used to compare the results obtained by fine-tuning BERT and SciBERT in terms of KP extraction.

키워드

keyphrase extraction; BERT; fine-tuning; embedding; scientific articles
제목
Fine-tuning BERT Models for Keyphrase Extraction in Scientific Articles
제목 (타언어)
Fine-tuning BERT Models for Keyphrase Extraction in Scientific Articles
저자
임연수; 서덕진; 정유철
발행일
2020-01
저널명
한국정보기술학회 영문논문지
권
10
호
1
페이지
45 ~ 56