LJungang
Awesome-Video-Reasoning-Landscape

πŸ”₯An open-source survey of the latest video reasoning tasks, paradigms, and benchmarks.

Last updated Aug 8, 2026
190
Stars
8
Forks
0
Issues
0
Stars/day
Attention Score
74
Language breakdown
No language data available.
β–Έ Files click to expand
README

The Landscape of Video Reasoning: Tasks, Paradigms and Benchmarksβ€” An Open-Source Survey

Awesome

πŸ—ΊοΈ Overview

This Awesome list systematically curates and tracks the latest progress in Video Reasoning, covering diverse modalities, tasks, and modeling paradigms. Rather than focusing on a single line of research, we organize the landscape from multiple complementary perspectives. Following the emerging taxonomy of the field, current works are grouped into four major paradigms:

  • πŸ—’οΈ CoT-based Video Reasoning β€” language-centric, chain-of-thought reasoning with Video-LMMs
  • πŸ•ΉοΈ CoF-based Video Reasoning β€” vision-centric reasoning grounded in world models or video generation
  • 🌈 Interleaved Video Reasoning β€” unified models that integrate multimodal interaction and iterative inference
  • πŸ” Streaming Video Reasoning β€” continuous, low-latency reasoning over long or unbounded video streams with online perception and incremental state updates.
We additionally maintain a dedicated Benchmark section that summarizes datasets, evaluation settings, and standardized tasks to support fair comparison across paradigms.
[!Note]
This repository aims to provide a structured, up-to-date, and open-source overview of the evolving landscape of video reasoning.
Contributions and PRs are warmly welcome β€” preferably in reverse chronological order (newest first) to keep the list fresh and easy to browse.

πŸ“– Contents

- πŸ“‘ Task Definition - 😎 Paradigms - πŸ—’οΈ CoT-based Video Reasoning - πŸ•ΉοΈ CoF-based Video Reasoning - 🌈 Interleaved Video Reasoning - πŸ” Streaming Video Reasoning - ✨️ Benchmarks - ✈ Related Surveys - 🌟 Star History - β™₯️ Contributors

πŸ“‘ Task Definition

TBD

😎 Paradigms

πŸ•ΉοΈ CoT-based Video Reasoning

| Title | Model & Code | Checkpoint | Input Modalities | Time | Venue | | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :---------------------------------------------------------------------------------------------: | :-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :------------: | :--------------------------: | | Rethinking Chain-of-Thought Reasoning for Videos | GitHub | N/A | Text Video | 2025-12 | Arxiv | | 1+1 > 2 : Detector-Empowered Video Large Language Model for Spatio-Temporal Grounding and Reasoning | GitHub | N/A | Text Video | 2025-12 | Arxiv | | TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning | N/A | N/A | Text Video | 2025-12 | Arxiv | | OneThinker: All-in-one Reasoning Model for Image and Video | GitHub | Hugging Face | Text Video | 2025-12 | Arxiv | | WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning | GitHub | N/A | Text Video | 2025-12 | Arxiv | | Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding | N/A | N/A | Text Video | 2025-12 | Arxiv | | Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models | GitHub | N/A | Text Video | 2025-11 | Arxiv | | Video-CoM: Interactive Video Reasoning via Chain of Manipulations | GitHub | N/A | Text Video | 2025-11 | Arxiv | | VideoSeg-R1: Reasoning Video Object Segmentation via Reinforcement Learning | GitHub | N/A | Text Video | 2025-11 | Arxiv | | AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning | N/A | N/A | Audio Video | 2025-11 | Arxiv | | Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding | N/A | N/A | Text Video | 2025-11 | Arxiv | | Video Spatial Reasoning with Object-Centric 3D Rollout | N/A | N/A | Text Video | 2025-11 | Arxiv | | ViSS-R1: Self-Supervised Reinforcement Video Reasoning | N/A | N/A | Text Video | 2025-11 | Arxiv | | Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning | GitHub | Hugging Face | Text Video | 2025-10 | Arxiv | | Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence | GitHub | Hugging Face | Text Video | 2025-10 | Arxiv | | VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception | GitHub | Hugging Face | Text Video | 2025-09 | Arxiv | | MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning | N/A | N/A | Text Video | 2025-09 | Arxiv | | Kwai Keye-VL 1.5 Technical Report | GitHub | Hugging Face | Text Video | 2025-09 | Arxiv | | Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data | GitHub | Google_Drive | Text Video | 2025-09 | Arxiv | | Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding | N/A | N/A | Text Video | 2025-08 | Arxiv | | Ovis2.5 Technical Report | GitHub | Hugging Face | Text Video | 2025-08 | Arxiv | | Veason-R1: Reinforcing Video Reasoning Segmentation to Think Before It Segments | N/A | N/A | Text Video | 2025-08 | Arxiv | | ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking | GitHub | N/A | Text Video | 2025-08 | Arxiv | | TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding | N/A | N/A | Text Video | 2025-08 | Arxiv | | Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning | GitHub | Hugging Face | Text Video | 2025-08 | Arxiv | | AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video | GitHub | Hugging Face | Audio Video Text | 2025-08 | Arxiv | | ReasonAct: Progressive Training for Fine-Grained Video Reasoning in Small Models | N/A | N/A | Text Video | 2025-08 | Arxiv | | VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering | N/A | N/A | Text Video | 2025-08 | ACM-MM 2025 | | ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts | GitHub | Hugging Face | Text Audio Video | 2025-07 | Arxiv | | METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark | N/A | N/A | Text Audio Video | 2025-07 | Arxiv | | CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks | N/A | N/A | Text Video | 2025-07 | Arxiv | | EmbRACE-3K: Embodied Reasoning and Action in Complex Environments | N/A | N/A | Text Video | 2025-07 | Arxiv | | Scaling RL to Long Videos | GitHub | Hugging Face | Text Video | 2025-07 | NeurIPS 2025 | | Kwai Keye-VL Technical Report | GitHub | N/A | Text Video | 2025-07 | Arxiv | | ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models | GitHub | N/A | Text Video | 2025-07 | ACM-MM 2025 | | Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning | GitHub | Hugging Face | Text Video | 2025-07 | EMNLP 2025 | | Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames | N/A | N/A | Text Video | 2025-07 | Arxiv | | VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning | N/A | N/A | Text Video | 2025-06 | Arxiv | | Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning | GitHub | N/A | Text Video | 2025-06 | Arxiv | | DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning | N/A | N/A | Text Video | 2025-06 | Arxiv | | VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks | GitHub | Hugging Face | Text Video | 2025-06 | Arxiv | | HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context | GitHub | N/A | Audio Video Text | 2025-06 | Arxiv | | MiMo-VL Technical Report | GitHub | Hugging Face | Text Video | 2025-06 | Arxiv | | Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning | GitHub | N/A | Text Video | 2025-06 | EMNLP 2025 (Findinds) | | EgoVLM: Policy Optimization for Egocentric Video Understanding | GitHub | Hugging Face | Text Video | 2025-06 | Arxiv | | Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency | GitHub | Hugging Face | Text Video | 2025-06 | Arxiv | | VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking | N/A | N/A | Text Video | 2025-06 | Arxiv | | ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding | GitHub | N/A | Text Video | 2025-06 | NeurIPS 2025 | | ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding | N/A | N/A | Text Video | 2025-06 | Arxiv | | DIVE: Deep-search Iterative Video Exploration | Github | N/A | Text Video | 2025-06 | CVPR 2025 | | VideoDeepResearch: Long Video Understanding With Agentic Tool Using | Github | N/A | Text Video | 2025-06 | Arxiv | | Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency | N/A | N/A | Text Video | 2025-06 | Arxiv | | DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO | Github | N/A | Text Video | 2025-06 | NeurIPS 2025 | | Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought | N/A | Project_Page | Text Video | 2025-06 | Arxiv | | VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning | N/A | N/A | Text Video | 2025-06 | Arxiv | | Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning | Github | Hugging Face | Text Video | 2025-06 | Arxiv | | Reinforcing Video Reasoning with Focused Thinking | Github | Hugging Face | Text Video | 2025-05 | Arxiv | | A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding | Github | N/A | Text Video | 2025-05 | Arxiv | | Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration | Github | Hugging Face | Text Audio Video | 2025-05 | NeurIPS 2025 | | Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought | Github | N/A | Text Video | 2025-05 | NeurIPS 2025 | | VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Guided Iterative Policy Optimization | Github | Hugging Face | Text Video | 2025-05 | Arxiv | | Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning | Github | N/A | Text Speech Video | 2025-05 | NeurIPS 2025 | | Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning | Github | Hugging Face | Text Video | 2025-05 | NeurIPS 2025 | | UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning | Github | Hugging Face | Text Video | 2025-05 | Arxiv | | VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning | Github | Hugging Face | Text Video | 2025-05 | NeurIPS 2025 | | Seed1.5-VL Technical Report | N/A | N/A | Text Video | 2025-05 | Arxiv | | TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action


README truncated. View on GitHub

Β© 2026 GitRepoTrend Β· LJungang/Awesome-Video-Reasoning-Landscape Β· Updated daily from GitHub