Exploring Reinforcement Learning Effects on Chain-of-Thought Legibility

See this post on LessWrong here.