Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning
Researchers have discovered a way to create a 'backdoor' in multimodal large language models (MLLMs) that can induce fairness violations even after the model is updated through continual learning. This attack, called Persistent Fairness Backdoor Attack (PFBA), works by reshaping the model's deep feature geometry and optimizing the trigger against simulated parameter drift to ensure persistence across future updates. The PFBA induces severe fairness disparities that evade stan
Researchers have discovered a way to create a 'backdoor' in multimodal large language models (MLLMs) that can induce fairness violations even after the model is updated through continual learning. This attack, called Persistent Fairness Backdoor Attack (PFBA), works by reshaping the model's deep feature geometry and optimizing the trigger against simulated parameter drift to ensure persistence across future updates. The PFBA induces severe fairness disparities that evade standard backdoor defenses.
---
Why it matters: This matters because MLLMs are increasingly deployed in high-stakes domains where fairness is critical, such as hiring or loan applications. Engineers need to be aware of this vulnerability to prevent potential biases and discrimination.
Source: https://arxiv.org/abs/2608.21577
This article was originally published at: https://arxiv.org/abs/2608.21577