Rater State Bias in RLHF Preference Data: An Audit Framework
Rater State Bias in RLHF Preference Data
FAQ
What is rater state bias in RLHF?
It is a structural bias in preference data caused by raters' psychological state (e.g., stress) during evaluation, affecting model accuracy.
How does this bias differ from random noise?
It is state-dependent and can be shared across raters under similar conditions, unlike random noise.
Can the audit framework be applied to Arabic models?
Yes, with adjustments for linguistic and cultural characteristics of the region.
What are the practical implications of this bias?
It can lead to unfair or inaccurate models, especially in sensitive applications like healthcare or law.
Source: arXiv cs.AI
AI-assisted content, human-reviewed.