Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation

Published in International Conference on Learning Representations (ICLR), 2025

First-order logic reasoning is a valuable task for evaluating the reasoning capabilities of language models, but existing benchmarks rely on extensive human annotation or handcrafted templates. ProverGen synergizes the generative strengths of LLMs with the rigor of symbolic provers to create ProverQA, a scalable, diverse, and high-quality FOL reasoning dataset that includes accessible and logically coherent intermediate reasoning steps for each problem.

State-of-the-art LLMs struggle on ProverQA even with chain-of-thought prompting, and models finetuned on a separate ProverGen-generated training set improve consistently on both in-distribution and out-of-distribution test sets.