Explore concepts like Self-Correct, Self-Refine, Self-Improve, Self-Contradict, Self-Play, and Self-Knowledge, alongside o1-like reasoning elevation🍓 and hallucination alleviation🍄.
Internal Consistency and Self-Feedback in Large Language Models: A Survey
Explore concepts like Self-Correct, Self-Refine, Self-Improve, Self-Contradict, Self-Play, and Self-Knowledge, alongside o1-like reasoning elevation🍓 and hallucination alleviation🍄.
Xun Liang1*, Shichao Song1*, Zifan Zheng2*, Hanyu Wang1, Qingchen Yu2, Xunkai Li3, Rong-Hua Li3, Yi Wang4, Zhonghao Wang4, Feiyu Xiong2, Zhiyu Li2†
1RUC, 2IAAR, 3BIT, 4Xinhua
*Equal contribution, †Corresponding author (lizy@iaar.ac.cn)
[!IMPORTANT]
- Consider giving our repository a 🌟, so you will receive the latest news (paper list updates, new comments, etc.);
- If you want to cite our work, here is our bibtex entry: CITATION.bib.
📰 News
- 2024/10/26 We have created a relevant WeChat Group (微信群) for discussing reasoning and hallucination in LLMs.
- 2024/09/18 Paper v3.0 and a relevant Twitter thread.
- 2024/08/24 Updated paper list for better user experience. Link. Ongoing updates.
- 2024/07/22 Our paper ranks first on Hugging Face Daily Papers! Link.
- 2024/07/21 Our paper is now available on arXiv. Link.
🎉 Introduction
Welcome to the GitHub repository for our survey paper titled "Internal Consistency and Self-Feedback in Large Language Models: A Survey." The survey's goal is to provide a unified perspective on the self-evaluation and self-updating mechanisms in LLMs, encapsulated within the frameworks of Internal Consistency and Self-Feedback.

This repository includes three key resources:
- expt-consistency-types: Code and results for measuring consistency at different levels.
- expt-gpt4o-responses: Results from five different GPT-4o responses to the same query.
- Paper List: A comprehensive list of references related to our survey.
Click Me to Show the Table of Contents
- Related Survey Papers - Section IV: Consistency Signal Acquisition - Confidence Estimation - Hallucination Detection - Uncertainty Estimation - Verbal Critiquing - Faithfulness Measurement - Consistency Estimation - Section V: Reasoning Elevation - Reasoning Topologically - Refining with Responses - Multi-Agent Collaboration - Section VI: Hallucination Alleviation - Mitigating Hallucination while Generating - Refining the Response Iteratively - Activating Truthfulness - Decoding Truthfully - Section VII: Other Tasks - Preference Learning - Knowledge Distillation - Continuous Learning - Data Synthesis - Consistency Optimization - Decision Making - Event Argument Extraction - Inference Acceleration - Machine Translation - Negotiation Optimization - Retrieval Augmented Generation - Text Classification - Self-Repair - Section VIII.A: Meta Evaluation - Consistency Evaluation - Self-Knowledge Evaluation - Uncertainty Evaluation - Feedback Ability Evaluation - Reflection Ability Evaluation - Theoretical Perspectives📚 Paper List
Here we list the most important references cited in our survey, as well as the papers we consider worth noting. This list will be updated regularly.
Related Survey Papers
These are some of the most relevant surveys related to our paper.
- A Survey on the Honesty of Large Language Models
- Awesome LLM Reasoning
- Awesome LLM Strawberry
- Extrinsic Hallucinations in LLMs
- When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
- A Survey on Self-Evolution of Large Language Models
- Demystifying Chains, Trees, and Graphs of Thoughts
- Automatically Correcting Large Language Models: Surveying the Landscape of Diverse Automated Correction Strategies
- Uncertainty in Natural Language Processing: Sources, Quantification, and Applications
Section IV: Consistency Signal Acquisition
For various forms of expressions from an LLM, we can obtain various forms of consistency signals, which can help in better updating the expressions.
Confidence Estimation
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Linguistic Calibration of Long-Form Generations
- InternalInspector I2: Robust Confidence Estimation in LLMs through Internal States
- Cycles of Thought: Measuring LLM Confidence through Stable Explanations
- TrustScore: Reference-Free Evaluation of LLM Response Trustworthiness
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- Quantifying Uncertainty in Answers from any Language Model and Enhancing their Trustworthiness
- Teaching models to express their uncertainty in words
- Language Models (Mostly) Know What They Know
Hallucination Detection
- Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
- Prompt-Guided Internal States for Hallucination Detection of Large Language Models
- Detecting hallucinations in large language models using semantic entropy
- INSIDE: LLMs' Internal States Retain the Power of Hallucination Detection
- LLM Internal States Reveal Hallucination Risk Faced With a Query
- Teaching Large Language Models to Express Knowledge Boundary from Their Own Signals
- Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method
- LM vs LM: Detecting Factual Errors via Cross Examination
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
Uncertainty Estimation
- Enhancing Trust in Large Language Models with Uncertainty-Aware Fine-Tuning
- Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space
- Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
- Uncertainty Estimation of Large Language Models in Medical Question Answering
- To Believe or Not to Believe Your LLM
- Shifting Attention to Relevance: Towards the Uncertainty Estimation of Large Language Models
- Active Prompting with Chain-of-Thought for Large Language Models
- Uncertainty Estimation in Autoregressive Structured Prediction
- On Hallucination and Predictive Uncertainty in Conditional Language Generation
Verbal Critiquing
- LLM Critics Help Catch LLM Bugs
- Reasons to Reject? Aligning Language Models with Judgments
- Self-critiquing models for assisting human evaluators
Faithfulness Measurement
- Are self-explanations from Large Language Models faithful?
- On Measuring Faithfulness or Self-consistency of Natural Language Explanations
Consistency Estimation
- Semantic Consistency for Assuring Reliability of Large Language Models
Section V: Reasoning Elevation
Enhancing reasoning ability by improving LLM performance on QA tasks through Self-Feedback strategies.
Reasoning Topologically
- SRA-MCTS: Self-driven Reasoning Augmentation with Monte Carlo Tree Search for Code Generation
- Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions
- Dynamic Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
- Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains
- DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation
- Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models
- RATT: A Thought Structure for Coherent and Correct LLM Reasoning
- Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
- Chain-of-Thought Reasoning Without Prompting
- Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
- Training Language Models to Self-Correct via Reinforcement Learning
- LLMs cannot find reasoning errors, but can correct them given the error location
- Forward-Backward Reasoning in Large Language Models for Mathematical Verification
- LeanReasoner: Boosting Complex Logical Reasoning with Lean
- Just Ask One More Time! Self-Agreement Improves Reasoning of Language Models in (Almost) All Scenarios
- Soft Self-Consistency Improves Language Model Agents
- Self-Evaluation Guided Beam Search for Reasoning
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
- Universal Self-Consistency for Large Language Model Generation
- Enhancing Large Language Models in Coding Through Multi-Perspective Self-Consistency
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
- Making Language Models Better Reasoners with Step-Aware Verifier
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations
Refining with Responses
- Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
- SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
- Small Language Models Need Strong Verifiers to Self-Correct Reasoning
- Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models
- Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
- Teaching Language Models to Self-Improve by Learning from Language Feedback
- Large Language Models Can Self-Improve At Web Agent Tasks
- Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
- Can LLMs Learn from Previous Mistakes? Investigating LLMs’ Errors to Boost for Reasoning
- Fine-Grained Self-Endorsement Improves Factuality and Reasoning
- Mirror: A Multiple-perspective Self-Reflection Method for Knowledge-rich Reasoning
- Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-Evaluation
- Self-Rewarding Language Models
- Learning From Mistakes Makes LLM Better Reasoner
- Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
- Large Language Models Can Self-Improve
- Improving Logical Consistency in Pre-Trained Language Models using Natural Language Inference
- Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference
Multi-Agent Collaboration
- The Consensus Game: Language Model Generation via Equilibrium Search
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
- Scaling Large-Language-Model-based Multi-Agent Collaboration
- AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
- REFINER: Reasoning Feedback on Intermediate Representations
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
- Towards CausalGPT: A Multi-Agent Approach for Faithful Knowledge Reasoning via Promoting Causal Consistency in LLMs
Section VI: Hallucination Alleviation
Improving factual accuracy in open-ended generation and reducing hallucinations through Self-Feedback strategies.
Mitigating Hallucination while Generating
- Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
- Mitigating Entity-Level Hallucination in Large Language Models
- Know the Unknown: An Uncertainty-Sensitive Method for LLM Instruction Tuning
- Fine-grained Hallucination Detection and Editing for Language Models
- EVER: Mitigating Hallucination in Large Language Models through Real-Time Verification and Rectification
- Chain-of-Verification Reduces Hallucination in Large Language Models
- PURR: Efficiently Editing Language Model Hallucinations by Denoising Language Model Corruptions
- RARR: Researching and Revising What Language Models Say, Using Language Models
Refining the Response Iteratively
- An Evolutionary Large Language Model for Hallucination Mitigation
- From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
- Teaching Large Language Models to Self-Debug
- LLMs can learn self-restraint through iterative self-reflection
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Generating Sequences by Learning to Self-Correct
- MAF: Multi-Aspect Feedback for Improving Reasoning in Large Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- PEER: A Collaborative Language Model
- Re3: Generating Longer Stories With Recursive Reprompting and Revision
Activating Truthfulness
- Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
- Look Within, Why LLMs Hallucinate: A Causal Perspective
- Retrieval Head Mechanistically Explains Long-Context Factuality
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Fine-tuning Language Models for Factuality
Decoding Truthfully
- Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
- Diver: Large Language Model Decoding with Span-Level Mutual Information Verification
- SED: Self-Evaluation Decoding Enhances Large Language Models for Better Generation
- Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
- Trusting Your Evidence: Hallucinate Less with Context-aware Decoding
- Contrastive Decoding: Open-ended Text Generation as Optimization
Section VII: Other Tasks
In addition to tasks aimed at improving consistency (enhancing reasoning and alleviating hallucinations), there are other tasks that also utilize Self-Feedback strategies.
Preference Learning
- Language Imbalance Driven Rewarding for Multilingual Self-improving
- Aligning Large Language Models via Self-Steering Optimization
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- Aligning Large Language Models from Self-Reference AI Feedback with one General Principle
- Aligning Large Language Models with Self-generated Preference Data
- Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation
- Self-Improving Robust Preference Optimization
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- Self-Play Preference Optimization for Language Model Alignment
- ChatGLM-Math: Improving Math Problem-Solving in Large Language Models with a Self-Critique Pipeline
- SALMON: Self-Alignment with Instructable Reward Models
- Self-Specialization: Uncovering Latent Expertise within Large Language Models
- BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Aligning Large Language Models through Synthetic Feedback
- OpenAssistant Conversations -- Democratizing Large Language Model Alignment
- The Capacity for Moral Self-Correction in Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Knowledge Distillation
- Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models
- Personalized Distillation: Empowering Open-Sourced LLMs with Adaptive Learning for Code Generation
- SelFee: Iterative Self-Revising LLM Empowered by Self-Feedback Generation
- Reinforced Self-Training (ReST) for Language Modeling
- Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
- Self-Knowledge Distillation with Progressive Refinement of Targets
- Revisiting Knowledge Distillation via Label Smoothing Regularization
- Self-Knowledge Distillation in Natural Language Processing
Continuous Learning
- Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
- Self-Evolving GPT: A Lifelong Autonomous Experiential Learner
Data Synthesis
- Self-Taught Evaluators
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- Self-training Improves Pre-training for Natural Language Understanding
Consistency Optimization
- Improving the Robustness of Large Language Models via Consistency Alignment
Decision Making
- Can Large Language Models Play Games? A Case Study of A Self-Play Approach
Event Argument Extraction
- ULTRA: Unleash LLMs' Potential for Event Argument Extraction through Hierarchical Modeling and Pair-wise Refinement
Inference Acceleration
- Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
Machine Translation
- TasTe: Teaching Large Language Models to Translate through Self-Reflection
Negotiation Optimization
- Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
Retrieval Augmented Generation
- Improving Retrieval Augmented Language Model with Self-Reasoning
Text Classification
- Text Classification Using Label Names Only: A Language Model Self-Training Approach
Self-Repair
- Explorations of Self-Repair in Language Models
Section VIII.A: Meta Evaluation
Some common evaluation benchmarks.
Consistency Evaluation
- Evaluating Consistencies in LLM responses through a Semantic Clustering of Question Answering
- Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones?
- Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models
- Predicting Question-Answering Performance of Large Language Models through Semantic Consistency
- BECEL: Benchmark for Consistency Evaluation of Language Models
- Measuring and Improving Consistency in Pretrained Language Models
Self-Knowledge Evaluation
- Can I understand what I create? Self-Knowledge Evaluation of Large Language Models
- Can AI Assistants Know What They Don't Know?
- Do Large Language Models Know What They Don’t Know?
Uncertainty Evaluation
- UBENCH: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions
- Benchmarking LLMs via Uncertainty Quantification
Feedback Ability Evaluation
- CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
Reflection Ability Evaluation
- Reflection-Bench: probing AI intelligence with reflection
Theoretical Perspectives
Some theoretical research on Internal Consistency and Self-Feedback strategies.
- Think-to-Talk or Talk-to-Think? When LLMs Come Up with an Answer in Multi-Step Reasoning
- AI models collapse when trained on recursively generated data
- A Theoretical Understanding of Self-Correction through In-context Alignment
- Large Language Models Cannot Self-Correct Reasoning Yet
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- When Can Transformers Count to n?
- Large Language Models as Reliable Knowledge Bases?
- States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
- Large Language Models have Intrinsic Self-Correction Ability
- What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
- Large Language Models Must Be Taught to Know What They Don't Know
- Are LLMs classical or nonmonotonic reasoners? Lessons from generics
- On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
- Calibrating Reasoning in Language Models with Internal Consistency
- Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
- Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization
- SELF-[IN]CORRECT: LLMs Struggle with Refining Self-Generated Responses
- Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
- Do Large Language Models Latently Perform Multi-Hop Reasoning?
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
- The Impact of Reasoning Step Length on Large Language Models
- Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
- GPT-4 Doesn’t Know It’s Wrong: An Analysis of Iterative Prompting for Reasoning Problems
- Lost in the Middle: How Language Models Use Long Contexts
- How Language Model Hallucinations Can Snowball
- On the Principles of Parsimony and Self-Consistency for the Emergence of Intelligence
- On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
- How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
- Language Models as Knowledge Bases?
📝 Citation
@article{liang2024internal,
title={Internal consistency and self-feedback in large language models: A survey},
author={Liang, Xun and Song, Shichao and Zheng, Zifan and Wang, Hanyu and Yu, Qingchen and Li, Xunkai and Li, Rong-Hua and Wang, Yi and Wang, Zhonghao and Xiong, Feiyu and Li, Zhiyu},
journal={arXiv preprint arXiv:2407.14507},
year={2024}
}