Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
Uses a custom tool to elicit intermediate reasoning from frontier models, comparing performance with native reasoning and analysing differences in token efficiency and reasoning structure.