#Vision-language-models
Showing 11 of 11 repositories tagged #vision-language-models, ranked by stars
Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.
[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that can magically restore any image to perfect-4K!
ParseBench - A Document Parsing Benchmark for AI Agents
π Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
EVE Series: Encoder-Free Vision-Language Models from BAAI
Official Repository for PosterGen - CVPR Findings 2026
[ICCV 2025] Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
This repo is a live list of papers on game playing and large multimodality model - "A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges".
[AAAI-2026] Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
[ICLR 2024 Spotlight π₯ ] - [ Best Paper Award SoCal NLP 2023 π] - Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models
A comprehensive survey of VisionβLanguage Models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets