Online Conformal Abstention for Factuality Control Under Adversarial Bandit Feedback
Minjae Lee*, Yoonjae Jung*, and Sangdon Park
arXiv preprint arXiv:2506.14067, Jun 2025
*Equal contribution
We address reliability concerns in interactive generative systems by proposing ExAUL, an online learning framework that ensures systems answer only when confident, handling partial user feedback in adversarial settings. A conversion lemma links bandit algorithm regret to false discovery rate bounds, and a feedback unlocking strategy achieves O(sqrt(T) ln|H|) regret, translating to O(sqrt(T)) false discovery rate risk control despite incomplete feedback. We validate the framework on question-answering tasks with large language models under varied challenging conditions.