Invisible Reasoning
Vatsal Baherwani, Tom Goldstein, Ashwinee Panda
This paper demonstrates that frontier language models can perform 'invisible reasoning' using semantically irrelevant filler tokens, improving accuracy by up to 13 percentage points and hiding objectives from CoT monitoring.