Research & Papers

LLM creativity hinges on evaluator, not iterations, new study finds

Iterative generation alone won't make AI more creative; the in-loop evaluator matters most.

Deep Dive

A new paper from Leiden University researchers (Anderson, Verhoef, Zohrehvand) asks a fundamental question: does iterative generation actually make large language models more creative? To test this, they adapted FunSearch, DeepMind's evolutionary-search method, to recipe generation for the 2024 Pillsbury Bake-Off competition. They scored outputs using a TTCT-based (Torrance Tests of Creative Thinking) LLM evaluator, which measures dimensions like fluency, flexibility, originality, and elaboration.

Across two experiments, the researchers varied the number of generation-iteration loops, the temperature of the generator, and the size of the in-loop selection-scorer model. The results are counterintuitive: running more iterations did not improve creativity scores at all. Instead, the choice of evaluator model was the dominant factor. A smaller selection-scorer model produced significantly higher creativity scores on most TTCT dimensions than a larger one. Generator temperature had only a minor effect, and that was limited to originality. The team concludes that evaluator design is a first-order design variable in subjective creative search—more important than raw compute cycles or sampling diversity.

Key Points
  • Adapted FunSearch to Pillsbury Bake-Off recipe generation with TTCT-based LLM evaluation
  • Increasing iterations alone did not improve creativity scores; evaluator size mattered more
  • A smaller in-loop selection scorer yielded significantly higher creativity across most TTCT dimensions
  • Generator temperature had limited effects, only impacting originality scores

Why It Matters

Shows AI creativity pipelines should focus on building better evaluators, not just adding more iteration loops.

📬 Get the top 10 AI stories daily