Efficiently Detecting Hidden Reasoning with a Small Predictor Model

See this post on LessWrong here.