Utility-Based Preference Training for Effective Synthetic Text Classification

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

High-quality synthetic text can mitigate annotation scarcity in text classification. However, standard preference optimization often produces samples that are fluent but weakly label-specific. We present Utility-weighted Direct Preference Optimization (U-DPO), a preference-optimization framework for class-conditional synthetic data generation. In U-DPO, a task-specific classifier provides a margin-based external score for each candidate generation, which is combined with an embedding-based internal similarity score to form an overall utility. These utilities are used (i) to mine preference pairs from multiple candidates per class and (ii) to weigh each DPO update by the utility gap between preferred and dispreferred samples. This design encourages the generator to concentrate on learning informative, label-discriminative preference comparisons rather than treating all pairs equally. Across two multiclass scientific-abstract benchmarks (arXiv and WOS-11967), U-DPO consistently improves downstream SciBERT classification accuracy compared with both vanilla synthetic generation and standard DPO fine-tuning, with gains up to 0.88 percentage points on arXiv and 0.83 percentage points on WOS-11967 depending on the generator. An additional GPT-4.5-based evaluation also indicates a higher mean quality score for U-DPO samples with reduced variance.

키워드

Direct Preference Optimization (DPO); synthetic data generation; text classification; Large Language Models (LLM); utility-based learning; margin-based optimization
제목
Utility-Based Preference Training for Effective Synthetic Text Classification
저자
Gwak, Jiho; Jung, Yuchul
DOI
10.3390/math14030507
발행일
2026-01
유형
Article
저널명
MATHEMATICS
권
14
호
3

파일 다운로드

Thumbnail