ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models
A new framework called ExplainGuard aims to ensure the integrity of explanations generated by blackbox AI models. The current approach to auditing these models relies on an implicit 'chain of trust' that assumes third-party auditors are trustworthy. However, recent research has shown that this assumption is flawed and can be exploited by adversarial auditors. ExplainGuard uses a Zero-Trust architecture design to verify the integrity of explanations before they are released to
A new framework called ExplainGuard aims to ensure the integrity of explanations generated by blackbox AI models. The current approach to auditing these models relies on an implicit 'chain of trust' that assumes third-party auditors are trustworthy. However, recent research has shown that this assumption is flawed and can be exploited by adversarial auditors. ExplainGuard uses a Zero-Trust architecture design to verify the integrity of explanations before they are released to users. It does this through three distinct pillars of verification: detecting model substitution, rejecting mathematically impossible explanations, and verifying feature faithfulness. The authors claim that ExplainGuard can effectively neutralize state-of-the-art explanation manipulation attacks.
---
Why it matters: This matters because it addresses a critical weakness in the current auditing paradigm for blackbox AI models. By ensuring the integrity of explanations, ExplainGuard helps to build trust in these models, which is essential for their deployment in high-stakes environments such as healthcare and finance.
Source: https://arxiv.org/abs/2608.21803
This article was originally published at: https://arxiv.org/abs/2608.21803