상세 보기
Utility-Based Preference Training for Effective Synthetic Text Classification
- Gwak, Jiho;
- Jung, Yuchul
WEB OF SCIENCE
0SCOPUS
0초록
High-quality synthetic text can mitigate annotation scarcity in text classification. However, standard preference optimization often produces samples that are fluent but weakly label-specific. We present Utility-weighted Direct Preference Optimization (U-DPO), a preference-optimization framework for class-conditional synthetic data generation. In U-DPO, a task-specific classifier provides a margin-based external score for each candidate generation, which is combined with an embedding-based internal similarity score to form an overall utility. These utilities are used (i) to mine preference pairs from multiple candidates per class and (ii) to weigh each DPO update by the utility gap between preferred and dispreferred samples. This design encourages the generator to concentrate on learning informative, label-discriminative preference comparisons rather than treating all pairs equally. Across two multiclass scientific-abstract benchmarks (arXiv and WOS-11967), U-DPO consistently improves downstream SciBERT classification accuracy compared with both vanilla synthetic generation and standard DPO fine-tuning, with gains up to 0.88 percentage points on arXiv and 0.83 percentage points on WOS-11967 depending on the generator. An additional GPT-4.5-based evaluation also indicates a higher mean quality score for U-DPO samples with reduced variance.
키워드
- 제목
- Utility-Based Preference Training for Effective Synthetic Text Classification
- 저자
- Gwak, Jiho; Jung, Yuchul
- 발행일
- 2026-01
- 유형
- Article
- 저널명
- MATHEMATICS
- 권
- 14
- 호
- 3
- 언어
- ENG
- 출판사
- MDPI
- 발행국가
- 스위스
- ISSN
- E 2227-7390