The mathematical capabilities of AI systems are complex and multifaceted. Most existing research has focused on the correctness of AI-generated solutions to mathematical problems. In this work, we argue that beyond producing correct answers, AI systems should also be capable of, or assist humans in, developing novel solutions to mathematical challenges.
We introduce CreativeMath (AAAI 2025), a novel framework and benchmark encompassing problems ranging from middle school curricula to Olympic-level competitions, designed to assess LLMs’ ability to propose innovative solutions after some known solutions have been provided. Our experiments demonstrate that, while LLMs perform well on standard mathematical tasks, their capacity for creative problem-solving varies considerably — with Gemini-1.5-Pro outperforming other LLMs in generating novel solutions.
