-Take very hard math problem with a verifiable answer
-Have frontier model explain to a tiny model like (0.5-1B params and provably bad score on the problem) how to solve but not the solution, and reward the frontier model for prompts/explanations that helped the tiny model solve the problem
Obviously some amount of human supervision is needed to weed out it giving too much information