Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity
A principled interpretability framework based on token-level perplexity to test whether LLMs rely on linguistically relevant cues in benchmark evaluations.


