Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution

About

Probing learned concepts in large language models (LLMs) is crucial for understanding how semantic knowledge is encoded internally. Training linear classifiers on probing tasks is a principle approach to denote the vector of a certain concept in the representation space. However, the single vector identified for a concept varies with both data and training, making it less robust and weakening its effectiveness in real-world applications. To address this challenge, we propose an approach to approximate the subspace representing a specific concept. Built on linear probing classifiers, we extend the concept vectors into Gaussian Concept Subspace (GCS). We demonstrate GCS's effectiveness through measuring its faithfulness and plausibility across multiple LLMs with different sizes and architectures. Additionally, we use representation intervention tasks to showcase its efficacy in real-world applications such as emotion steering. Experimental results indicate that GCS concept vectors have the potential to balance steering performance and maintaining the fluency in natural language generation tasks.

Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani, Fan Yang, Mengnan Du• 2024

Related benchmarks

TaskDatasetResultRank
Classification ProbingCounterFact (test)
Probe Acc (Best Layer)87.7
21
Classification ProbingCities (test)
Probe Accuracy (Best Layer)99.7
21
Classification ProbingSarcasm (test)
Probe Acc (Best Layer)94.2
21
Classification ProbingCommon (test)
Probe Accuracy (Best Layer)74.4
21
Classification ProbingHateXplain (test)
Probe Accuracy (Best Layer)76.6
21
Classification ProbingSTSA (test)
Probe Accuracy (Best Layer)0.945
21
Concept vector stabilitySTSA
Mean Absolute-Cosine Similarity0.98
9
Concept vector stabilitySarcasm
Mean Absolute-Cosine Similarity0.99
6
Concept vector stabilityHateXplain
Mean Abs-Cosine Similarity0.98
3
Showing 9 of 9 rows

Other info

Follow for update