✨✨Latest Papers and Benchmarks in Reasoning with Foundation Models
Awesome-Reasoning-Foundation-Models

survey.pdf | A curated list of awesome large AI models, or foundation models, for reasoning.
We organize the current foundation models into three categories: language foundation models, vision foundation models, and multimodal foundation models. Further, we elaborate the foundation models in reasoning tasks, including commonsense, mathematical, logical, causal, visual, audio, multimodal, agent reasoning, etc. Reasoning techniques, including pre-training, fine-tuning, alignment training, mixture of experts, in-context learning, and autonomous agent, are also summarized.
We welcome contributions to this repository to add more resources. Please submit a pull request if you want to contribute! See CONTRIBUTING.
Table of Contents
table of contents
0 Survey

This repository is primarily based on the following paper:
>A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
>
[Paper][ArXiv]>
[Jiankai Sun](),Chuanyang Zheng, Enze Xie, Zhengying Liu, [Ruihang Chu](), [Jianing Qiu](), [Jiaqi Xu](), [Mingyu Ding](), Hongyang Li, [Mengzhe Geng](), [Yue Wu](), Wenhai Wang, [Junsong Chen](), [Zhangyue Yin](), [Xiaozhe Ren](), Jie Fu, Junxian He, Wu Yuan, Qi Liu, Xihui Liu, Yu Li, Hao Dong, Yu Cheng, Ming Zhang, Pheng Ann Heng, Jifeng Dai, Ping Luo, Jingdong Wang, Ji-Rong Wen, Xipeng Qiu, Yike Guo, Hui Xiong, Qun Liu, and Zhenguo Li
If you find this repository helpful, please consider citing:
@article{sun2025survey,
author = {Sun, Jiankai and Zheng, Chuanyang and Xie, Enze and Liu, Zhengying and Chu, Ruihang and Qiu, Jianing and Xu, Jiaqi and Ding, Mingyu and Li, Hongyang and Geng, Mengzhe and Wu, Yue and Wang, Wenhai and Chen, Junsong and Yin, Zhangyue and Ren, Xiaozhe and Fu, Jie and He, Junxian and Wu, Yuan and Liu, Qi and Liu, Xihui and Li, Yu and Dong, Hao and Cheng, Yu and Zhang, Ming and Heng, Pheng Ann and Dai, Jifeng and Luo, Ping and Wang, Jingdong and Wen, Ji-Rong and Qiu, Xipeng and Guo, Yike and Xiong, Hui and Liu, Qun and Li, Zhenguo},
title = {A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook},
year = {2025},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
issn = {0360-0300},
url = {https://doi.org/10.1145/3729218},
doi = {10.1145/3729218},
abstract = {Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It serves as a fundamental methodology in the field of Artificial General Intelligence (AGI). With the ongoing development of foundation models, there is a growing interest in exploring their abilities in reasoning tasks. In this paper, we introduce seminal foundation models proposed or adaptable for reasoning, highlighting the latest advancements in various reasoning tasks, methods, and benchmarks. We then delve into the potential future directions behind the emergence of reasoning abilities within foundation models. We also discuss the relevance of multimodal learning, autonomous agents, and super alignment in the context of reasoning. By discussing these future research directions, we hope to inspire researchers in their exploration of this field, stimulate further advancements in reasoning with foundation models, e.g. Large Language Models (LLMs), and contribute to the development of AGI.},
journal = {ACM Comput. Surv.},
month = apr,
keywords = {Reasoning, Foundation Models, Multimodal, AI Agent, Artificial General Intelligence, LLM}
}
1 Relevant Surveys and Links
relevant surveys
- Combating Misinformation in the Age of LLMs: Opportunities and Challenges
- The Rise and Potential of Large Language Model Based Agents: A Survey
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants
- A Survey on Multimodal Large Language Models
- Interactive Natural Language Processing
- A Survey of Large Language Models
- Self-Supervised Multimodal Learning: A Survey
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Towards Reasoning in Large Language Models: A Survey
- Reasoning with Language Model Prompting: A Survey
- Awesome Multimodal Reasoning
2 Foundation Models
foundation models

Table of Contents - 2
foundation models (table of contents)
2.1 Language Foundation Models
LFMs
Foundation Models (Back-to-Top)
2023/10|Mistral| Mistral 7B
2023/09|Qwen| Qwen Technical Report
2023/07|Llama 2| Llama 2: Open Foundation and Fine-Tuned Chat Models
2023/07|InternLM| InternLM: A Multilingual Language Model with Progressively Enhanced Capabilities
2023/05|PaLM 2| PaLM 2 Technical Report
2023/03|PanGu-Σ| PanGu-Σ: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing
2023/03|Vicuna| Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
2023/03|GPT-4| GPT-4 Technical Report
2023/02|LLaMA| LLaMA: Open and Efficient Foundation Language Models
2022/11|ChatGPT| Chatgpt: Optimizing language models for dialogue
2022/04|PaLM| PaLM: Scaling Language Modeling with Pathways
2021/09|FLAN| Finetuned Language Models Are Zero-Shot Learners
2021/07|Codex| Evaluating Large Language Models Trained on Code
2021/05|GPT-3| Language Models are Few-Shot Learners
2021/04|PanGu-α| PanGu-α: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
2019/08|Sentence-BERT| Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
2019/07|RoBERTa| RoBERTa: A Robustly Optimized BERT Pretraining Approach
2.2 Vision Foundation Models
VFMs
Foundation Models (Back-to-Top)
2024/01|Depth Anything
Yang et al.
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
[arXiv] [paper] [code] [project]
2023/05|SAA+
Cao et al.
Segment Any Anomaly without Training via Hybrid Prompt Regularization
[arXiv] [paper] [code]
2023/05|Explain Any Concept| Explain Any Concept: Segment Anything Meets Concept-Based Explanation
2023/05|SAM-Track| Segment and Track Anything
2023/04|Edit Everything| Edit Everything: A Text-Guided Generative System for Images Editing
2023/04|Inpaint Anything| Inpaint Anything: Segment Anything Meets Image Inpainting
2023/04|SAM
Kirillov et al., ICCV 2023
Segment Anything
[arXiv] [paper] [code] [blog]
2023/03|VideoMAE V2| VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking
2023/03|Grounding DINO
Liu et al.
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
[arXiv] [paper] [code]
2022/03|VideoMAE| VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
2021/12|Stable Diffusion
Rombach et al., CVPR 2022
High-Resolution Image Synthesis with Latent Diffusion Models
[arXiv] [paper] [code] [stable diffusion
2021/09|LaMa| Resolution-robust Large Mask Inpainting with Fourier Convolutions
2021/03|Swin
Liu et al., ICCV 2021
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
[arXiv] [paper] [code]
2020/10|ViT
Dosovitskiy et al., ICLR 2021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
[arXiv] [paper] [Implementation]
2.3 Multimodal Foundation Models
MFMs
Foundation Models (Back-to-Top)
2024/01|LLaVA-1.6
Liu et al. LLaVA-1.6: Improved reasoning, OCR, and world knowledge
[code] [blog]
2024/01|MouSi
Fan et al.
MouSi: Poly-Visual-Expert Vision-Language Models
[arXiv] [paper] [code]
2023/12|InternVL
Chen et al.
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
[arXiv] [paper] [code]
2023/12|Gemini| Gemini: A Family of Highly Capable Multimodal Models
2023/10|LLaVA-1.5
Liu et al.
Improved Baselines with Visual Instruction Tuning
[arXiv] [paper] [code] [project]
2023/09|GPT-4V| GPT-4V(ision) System Card
2023/08|Qwen-VL| Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
2023/05|InstructBLIP| InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
2023/05|Caption Anything| Caption Anything: Interactive Image Description with Diverse Multimodal Controls
2023/05|SAMText| Scalable Mask Annotation for Video Text Spotting
2023/04|Text2Seg| Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
2023/04|MiniGPT-4| MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
2023/04|LLaVA| Visual Instruction Tuning
2023/04|CLIP Surgery| CLIP Surgery for Better Explainability with Enhancement in Open-Vocabulary Tasks
2023/03|UniDiffuser| One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale
2023/01|GALIP| GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis
2023/01|BLIP-2| BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
2022/12|Img2Prompt| From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models
2022/05|CoCa| CoCa: Contrastive Captioners are Image-Text Foundation Models
2022/01|BLIP| BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
2021/09|CoOp| Learning to Prompt for Vision-Language Models
2.4 Reasoning Applications
reasoning applications
Foundation Models (Back-to-Top)
2022/06|Minerva| Solving Quantitative Reasoning Problems with Language Models
2022/06|BIG-bench| Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
2022/05|Zero-shot-CoT| Large Language Models are Zero-Shot Reasoners
2022/03|STaR| STaR: Bootstrapping Reasoning With Reasoning
2021/07|MWP-BERT| MWP-BERT: Numeracy-Augmented Pre-training for Math Word Problem Solving
2017/05|AQUA-RAT| Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
3 Reasoning Tasks
reasoning tasks
Table of Contents - 3
reasoning tasks (table of contents)
3.1 Commonsense Reasoning
commonsense reasoning
2023/05|LLM-MCTS| Large Language Models as Commonsense Knowledge for Large-Scale Task Planning
2023/05| Bridging the Gap between Pre-Training and Fine-Tuning for Commonsense Generation
2022/10|CoCoGen| Language Models of Code are Few-Shot Commonsense Learners
3.1.1 Commonsense Question and Answering (QA)
2019/06|CoS-E| Explain Yourself! Leveraging Language Models for Commonsense Reasoning
2016/12|ConceptNet| ConceptNet 5.5: An Open Multilingual Graph of General Knowledge
3.1.2 Physical Commonsense Reasoning
2025/05|PhyX| PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
2023/10|NEWTON| NEWTON: Are Large Language Models Capable of Physical Reasoning?
2022/03|PACS| PACS: A Dataset for Physical Audiovisual CommonSense Reasoning
2021/10|VRDP| Dynamic Visual Reasoning by Learning Differentiable Physics Models from Video and Language
2020/05|ESPRIT| ESPRIT: Explaining Solutions to Physical Reasoning Tasks
2019/11|PIQA| PIQA: Reasoning about Physical Commonsense in Natural Language
3.1.3 Spatial Commonsense Reasoning
2024/01|SpatialVLM
Chen et al.
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
[arXiv] [paper] [project] \- [Paper] [Code]
2021/06|PROST| PROST: Physical Reasoning of Objects through Space and Time
2019/02|GQA| GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
3.1.x Benchmarks, Datasets, and Metrics
2023/06|CConS| Probing Physical Reasoning with Counter-Commonsense Context
2023/05|SummEdits| LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond
2021/03|RAINBOW| UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark
2020/11|ProtoQA| ProtoQA: A Question Answering Dataset for Prototypical Common-Sense Reasoning
2020/10|DrFact| Differentiable Open-Ended Commonsense Reasoning
2019/11|CommonGen| CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning
2019/08|Cosmos QA| Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning
2019/08|αNLI| Abductive Commonsense Reasoning
2019/08|PHYRE| PHYRE: A New Benchmark for Physical Reasoning
2019/07|WinoGrande| WinoGrande: An Adversarial Winograd Schema Challenge at Scale
2019/05|MathQA| MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
2019/05|HellaSwag| HellaSwag: Can a Machine Really Finish Your Sentence?
2019/04|Social IQa| SocialIQA: Commonsense Reasoning about Social Interactions
2002/07|BLEU| BLEU: a Method for Automatic Evaluation of Machine Translation
3.2 Mathematical Reasoning
mathematical reasoning
2023/10|MathVista| MathVista: Evaluating Math Reasoning in Visual Contexts with GPT-4V, Bard, and Other Large Multimodal Models
Lu et al., ICLR 2024
2022/11| Tokenization in the Theory of Knowledge
2022/06|MultiHiertt| MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual Data
2021/04|MultiModalQA| MultiModalQA: Complex Question Answering over Text, Tables and Images
2017/05| Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
2004| Wittgenstein on philosophy of logic and mathematics
1989|CLP| Connectionist Learning Procedures
3.2.1 Arithmetic Reasoning
Mathematical Reasoning (Back-to-Top)
2022/09|PromptPG| Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning
2021/03|SVAMP| Are NLP Models really able to Solve Simple Math Word Problems?
2021/03|MATH| Measuring Mathematical Problem Solving With the MATH Dataset
2016/08| How well do Computers Solve Math Word Problems? Large-Scale Dataset Construction and Evaluation
2014/06|Alg514| Learning to Automatically Solve Algebra Word Problems
3.2.2 Geometry Reasoning
Mathematical Reasoning (Back-to-Top)
2024/01|AlphaGeometry| Solving olympiad geometry without human demonstrations
Trinh et al., Nature
2022/12|UniGeo/Geoformer| UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression
2021/05|GeoQA/NGS| GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning
2021/05|Geometry3K/Inter-GPS| Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
2015/09|GeoS| Solving Geometry Problems: Combining Text and Diagram Interpretation
3.2.3 Theorem Proving
Mathematical Reasoning (Back-to-Top)
2020/10|Prover| LEGO-Prover: Neural Theorem Proving with Growing Libraries
2023/09|Lyra| Lyra: Orchestrating Dual Correction in Automated Theorem Proving
2023/06|DT-Solver| DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-level Value Function
2023/03|Magnushammer| Magnushammer: A Transformer-based Approach to Premise Selection
2022/05|HTPS| HyperTree Proof Search for Neural Theorem Proving
2021/07|Lean 4| The Lean 4 Theorem Prover and Programming Language
2021/02|TacticZero| TacticZero: Learning to Prove Theorems from Scratch with Deep Reinforcement Learning
2021/02|PACT| Proof Artifact Co-training for Theorem Proving with Language Models
2020/09|GPT-f|Generative Language Modeling for Automated Theorem Proving
2019/06|Metamath| A Computer Language for Mathematical Proofs
2019/05|CoqGym| Learning to Prove Theorems via Interacting with Proof Assistants
2018/12|AlphaZero| A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
2018/04|TacticToe| TacticToe: Learning to Prove with Tactics
2015/08|Lean| The Lean Theorem Prover (system description)
2010/07| Three Years of Experience with Sledgehammer, a Practical Link between Automatic and Interactive Theorem Provers
2010/04| Formal Methods at Intel - An Overview
2005/07| Combining Simulation and Formal Verification for Integrated Circuit Design Validation
2003| Extracting a Formally Verified, Fully Executable
1996|Coq| The Coq Proof Assistant-Reference Manual
1994|Isabelle| Isabelle: A Generic Theorem Prover
3.2.4 Scientific Reasoning
Mathematical Reasoning (Back-to-Top)
2023/07|SciBench| SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
2022/09|ScienceQA| Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
2022/03|ScienceWorld| ScienceWorld: Is your Agent Smarter than a 5th Grader?
2012| Current Topics in Children's Learning and Cognition
3.2.x Benchmarks, Datasets, and Metrics
Mathematical Reasoning (Back-to-Top)
2024/01|MathBench
MathBench: A Comprehensive Multi-Level Difficulty Mathematics Evaluation Dataset
[code]
2023/08|Math23K-F/MAWPS-F/FOMAS| Guiding Mathematical Reasoning via Mastering Commonsense Formula Knowledge
2023/07|ARB| ARB: Advanced Reasoning Benchmark for Large Language Models
2023/05|SwiftSage| SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks
2023/05|TheoremQA| TheoremQA: A Theorem-driven Question Answering dataset
2022/10|MGSM| Language Models are Multilingual Chain-of-Thought Reasoners
2021/10|GSM8K| Training Verifiers to Solve Math Word Problems
2021/10|IconQA| IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
2021/09|FinQA| FinQA: A Dataset of Numerical Reasoning over Financial Data
2021/08|MBPP/MathQA-Python| Program Synthesis with Large Language Models
2021/08|HiTab/EA| HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation
2021/07|HumanEval/Codex| Evaluating Large Language Models Trained on Code
2021/06|ASDiv/CLD| A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers
2021/05|APPS| Measuring Coding Challenge Competence With APPS
2021/05|TAT-QA| TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
2021/03|SVAMP| Are NLP Models really able to Solve Simple Math Word Problems?
2021/01|TSQA/MAP/MRR| TSQA: Tabular Scenario Based Question Answering
2020/04|HybridQA| HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data
2019/03|DROP| DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
2019|NaturalQuestions| Natural Questions: A Benchmark for Question Answering Research
2018/09|HotpotQA| HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
2018/09|Spider| Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
2018/03|ComplexWebQuestions| The Web as a Knowledge-base for Answering Complex Questions
2017/12|MetaQA| Variational Reasoning for Question Answering with Knowledge Graph
2017/09|GEOS++| From Textbooks to Knowledge: A Case Study in Harvesting Axiomatic Knowledge from Textbooks to Solve Geometry Problems
2017/09|Math23k| Deep Neural Solver for Math Word Problems
2017/08|WikiSQL/Seq2SQL| Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
2017/05|TriviaQA| TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
2017/05|GeoShader| Synthesis of Solutions for Shaded Area Geometry Problems
2016/09|DRAW-1K| Annotating Derivations: A New Evaluation Strategy and Dataset for Algebra Word Problems
2016/08|WebQSP| The Value of Semantic Parse Labeling for Knowledge Base Question Answering
2016/06|SQuAD| SQuAD: 100,000+ Questions for Machine Comprehension of Text
2016/06|WikiMovies| Key-Value Memory Networks for Directly Reading Documents
2016/06|MAWPS| MAWPS: A Math Word Problem Repository
2015/09|Dolphin1878| Automatically Solving Number Word Problems by Semantic Parsing and Reasoning
2015/08|WikiTableQA| Compositional Semantic Parsing on Semi-Structured Tables
2015|SingleEQ| Parsing Algebraic Word Problems into Equations
2015|DRAW| DRAW: A Challenging and Diverse Algebra Word Problem Set
2014/10|Verb395| Learning to Solve Arithmetic Word Problems with Verb Categorization
2013/10|WebQuestions| Semantic Parsing on Freebase from Question-Answer Pairs
2013/08|Free917| Large-scale Semantic Parsing via Schema Matching and Lexicon Extension
1990|ATIS| The ATIS Spoken Language Systems Pilot Corpus
3.3 Logical Reasoning
logical reasoning
2023/10|LogiGLUE| Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
2023/05|LogicLLM| LogicLLM: Exploring Self-supervised Logic-enhanced Training for Large Language Models
2023/05|Logic-LM| Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
2023/03|LEAP| Explicit Planning Helps Language Models in Logical Reasoning
2022/10|Entailer| Entailer: Answering Questions with Faithful and Truthful Chains of Reasoning
2022/06|NeSyL| Weakly Supervised Neural Symbolic Learning for Cognitive Tasks
2022/05|NeuPSL| NeuPSL: Neural Probabilistic Soft Logic
2022/05|NLProofS| Generating Natural Language Proofs with Verifier-Guided Search
2022/05|Least-to-Most Prompting| Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
2022/05|SI| Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
2022/05|MERIt| MERIt: Meta-Path Guided Contrastive Learning for Logical Reasoning
2021/09|DeepProbLog| Neural probabilistic logic programming in DeepProbLog
2021/08|GABL| Abductive Learning with Ground Knowledge Base
2021/05|LReasoner| Logic-Driven Context Extension and Data Augmentation for Logical Reasoning of Text
2020/02|RuleTakers| Transformers as Soft Reasoners over Language
2019/12|NMN-Drop| Neural Module Networks for Reasoning over Text
2019/04|NS-CL| The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
2012| Logical Reasoning and Learning
3.3.1 Propositional Logic
2022/09| Propositional Reasoning via Neural Transformer