KoBERT–LLM 합의 기반 기타 라벨 재라벨링 기법

Consensus-based Relabeling of the “Other” Class using KoBERT and Large Language Models
  • 권수민; 
  • 곽지호; 
  • 박지영; 
  • 정유철

초록

The “other” label in text classification aggregates heterogeneous and ambiguous samples, often limiting overall performance. We propose an automatic relabeling method that combines a KoBERT classifier with an LLM, using their agreement and confidence to reassign “other” documents to more suitable labels. Retraining on the relabeled data redistributes “other” samples and improves overall performance as well as the stability of “other” predictions. This suggests model-consensus relabeling as a practical way to mitigate the “other” label issue with minimal manual validation. Future work will extend this into a universal data refinement framework via hierarchical classification and prompt optimization.

키워드

large language model; text classification; label noise reduction; KoBERT; automatic relabeling; label correction; .
제목
KoBERT–LLM 합의 기반 기타 라벨 재라벨링 기법
제목 (타언어)
Consensus-based Relabeling of the “Other” Class using KoBERT and Large Language Models
저자
권수민; 곽지호; 박지영; 정유철
DOI
10.14801/jkiit.2026.24.5.1
발행일
2026-05
유형
Y
저널명
한국정보기술학회논문지
권
24
호
5
페이지
1 ~ 13