#Vision-language-models

Showing 11 of 11 repositories tagged #vision-language-models, ranked by stars

gokayfem
gokayfem
awesome-vlm-architectures

Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

Score
0
β˜… 1.3k β‘‚ 55 +4/day
Markdown
taco-group
taco-group
4KAgent

[NeurIPS 2025] 4KAgent: Agentic Any Image to 4K Super-Resolution. An intelligent computer vision agent that can magically restore any image to perfect-4K!

Score
62
β˜… 819 β‘‚ 45 +1/day
Python
run-llama
run-llama
ParseBench

ParseBench - A Document Parsing Benchmark for AI Agents

Score
100
β˜… 543 β‘‚ 78 +1/day
Python
worldbench
worldbench
awesome-vla-for-ad

🌐 Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future

Score
88
β˜… 461 β‘‚ 39 +4/day
HTML
baaivision
baaivision
EVE

EVE Series: Encoder-Free Vision-Language Models from BAAI

Score
38
β˜… 376 β‘‚ 12 β€”
Python
Y-Research-SBU
Y-Research-SBU
PosterGen

Official Repository for PosterGen - CVPR Findings 2026

Score
75
β˜… 252 β‘‚ 23 +2/day
Python
worldbench
worldbench
DriveBench

[ICCV 2025] Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives

Score
50
β˜… 246 β‘‚ 15 β€”
Python
BAAI-Agents
BAAI-Agents
GPA-LM

This repo is a live list of papers on game playing and large multimodality model - "A Survey on Game Playing Agents and Large Models: Methods, Applications, and Challenges".

Score
25
β˜… 163 β‘‚ 7 β€”
Wenchuan-Zhang
Wenchuan-Zhang
Patho-R1

[AAAI-2026] Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner

Score
0
β˜… 101 β‘‚ 7 β€”
Python
erfanshayegani
erfanshayegani
Jailbreak-In-Pieces

[ICLR 2024 Spotlight πŸ”₯ ] - [ Best Paper Award SoCal NLP 2023 πŸ†] - Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Score
12
β˜… 95 β‘‚ 6 β€”
Python
SufyanDanish
SufyanDanish
VLM-Survey-

A comprehensive survey of Vision–Language Models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets

Score
0
β˜… 16 β‘‚ 0 β€”
Related Topics
#llm#vlm#large-language-models#awesome-list#mllm#multimodal#computer-vision#foundation-models#transformers#agent#benchmark#autonomous-driving

Β© 2026 GitRepoTrend Β· GitHub repositories by topic Β· Updated weekly