Reliable models should not only predict correctly, but also base their decisions on acceptable evidence. Yet class-level supervision can reward shortcuts that never touch the intended object or interface element.
We periodically expose compact decision-supporting regions with subset-selection attribution, compare them with human-provided masks or boxes, and penalize influential evidence that falls off-prior.
Across image classification and MLLM-based GUI clicking, the resulting models improve observed task performance and attribution–prior alignment.