#Image-to-text
Showing 11 of 11 repositories tagged #image-to-text, ranked by stars
Implementation of CoCa, Contrastive Captioners are Image-Text Foundation Models, in Pytorch
Flame is an open-source multimodal AI system designed to translate UI design mockups into high-quality React code. It leverages vision-language modeling, automated data synthesis, and structured training workflows to bridge the gap between design and front-end development.
macOS CLI for OCR and searchable PDFs using Apple's Vision framework
An out-of-the-box local Web UI for DeepSeek-OCR. Built with FastAPI + Vue.js, it supports PDF/Image uploads, progress tracking, and result visualization with bounding boxes. Easily experience the power of a top-tier OCR model.
Solution to im2latex request for research of openai
OCR with Google's AI technology (Cloud Vision API)
The largest multilingual image-text classification dataset. It contains fashion products.
Demonstrates Voice Recognition, Text to Speech, Language Translation, OAuth2, Image Generation, Face Detection and Voice Chatbot.
Tesseract.js OCR
Easiest way to use AI models without coding (Web UI & API support)
Stable Diffusion with Text-to-Image and Image-to-Text