Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text Classification

About

Cross-domain text classification aims to transfer models from label-rich source domains to label-poor target domains, giving it a wide range of practical applications. Many approaches promote cross-domain generalization by capturing domain-invariant features. However, these methods rely on unlabeled samples provided by the target domains, which renders the model ineffective when the target domain is agnostic. Furthermore, the models are easily disturbed by shortcut learning in the source domain, which also hinders the improvement of domain generalization ability. To solve the aforementioned issues, this paper proposes TACIT, a target domain agnostic feature disentanglement framework which adaptively decouples robust and unrobust features by Variational Auto-Encoders. Additionally, to encourage the separation of unrobust features from robust features, we design a feature distillation task that compels unrobust features to approximate the output of the teacher. The teacher model is trained with a few easy samples that are easy to carry potential unknown shortcuts. Experimental results verify that our framework achieves comparable results to state-of-the-art baselines while utilizing only source domain data.

Rui Song, Fausto Giunchiglia, Yingji Li, Mingjie Tian, Hao Xu• 2023

Related benchmarks

TaskDatasetResultRank
Machine-generated text detectionSQuAD--
30
MGT detectionWP
Accuracy87.22
16
MGT detectionHSWAG
Accuracy94.13
16
MGT detectionYelp
Accuracy87.89
16
MGT detectionCMV
Accuracy84.88
16
MGT detectionXsum
Accuracy85.23
16
MGT detectionELI5
Accuracy81.09
16
MGT detectionSCI
Accuracy88.36
16
MGT detectionTLDR
Accuracy64.73
16
MGT detectionROCT
Accuracy57.99
16
Showing 10 of 20 rows

Other info

Follow for update