Albumentations on GitHub in Documents and OCR
Used in 549 public GitHub repositories
Document understanding, OCR, handwriting, scene text, and layout analysis.

Public examples
Top public GitHub repositories
Showing the top 100
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
81,114 stars11,056 forkspix2tex: Using a ViT to convert images of equations into LaTeX code.
16,449 stars1,308 forksConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
15,181 stars1,277 forksTranslate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)
9,999 stars1,035 forksEffortless data labeling with AI support from Segment Anything and other awesome models.
9,343 stars1,067 forksThe official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
9,003 stars775 forksOfficial code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
8,132 stars706 forksAll-in-One Development Tool based on PaddlePaddle
6,160 stars1,208 forksPDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
6,015 stars417 forksOpenMMLab Text Detection, Recognition and Understanding Toolbox
4,736 stars779 forksCnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】
3,764 stars539 forks🛠 All-in-one web-based IDE specialized for machine learning and data science.
3,540 stars458 forksInternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
3,206 stars233 forksAn Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
3,148 stars275 forksOptical character recognition for Japanese text, with the main focus being Japanese manga
2,677 stars135 forksDocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
2,175 stars170 forks基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.
1,804 stars200 forks基于 manga-image-translator 的开源漫画翻译工具。支持日/韩/美漫自动翻译,内置 OpenAI、Gemini 等 5 种翻译引擎,并提供可视化编辑器自由调整文本样式。一键安装,开箱即用。如果喜欢,欢迎点亮 ⭐ Star 支持!
1,748 stars107 forksOpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
1,360 stars134 forksA central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.
1,237 stars128 forksParsing-free RAG supported by VLMs
960 stars76 forksGenerate text line images for training deep learning OCR models
912 stars175 forks#23
alephpi/Texo
A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。
826 stars51 forksCnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包
792 stars115 forksTransformer OCR
779 stars252 forksOCR toolbox from Davar-Lab
760 stars152 forks天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取
657 stars113 forks[Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting
586 stars42 forksLightweight & fast OCR models for license plate text recognition.
578 stars77 forksOfficial Code for ECCV 2024 paper — One-Shot Diffusion Mimicker for Handwritten Text Generation
555 stars56 forksA scene text recognition toolbox based on PyTorch
534 stars99 forks2nd solution of ICDAR 2021 Competition on Scientific Literature Parsing, Task B.
470 stars106 forksPLPR utilizes YOLOv5 and custom models for high-accuracy Persian license plate recognition, featuring real-time processing and an intuitive interface in an open-source framework.
449 stars126 forkscrnn chinese_plate_recognition
404 stars80 forksCode base of 'PAPER2WEB: LET’S MAKE YOUR PAPER ALIVE!' integrating the other work package throughout the entire "paper2present" process.
377 stars30 forksMakeACopy is an open-source document scanner app for Android that allows you to digitize paper documents with OCR functionality. The app is designed to be privacy-friendly, working completely offline without any cloud connection or tracking.
371 stars28 forksThe Infosys Responsible AI toolkit incorporates various features including safety, security, explainability, fairness, bias and hallucination detection to ensure AI solutions are trustworthy and transparent.
297 stars74 forksMLEvolve is an open-source autonomous system for end-to-end machine learning algorithm design and optimization powered by progressive search and experience-driven memory.
293 stars54 forksThe official repo for [CVPR'23] "DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting" & [ArXiv'23] "DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting"
290 stars44 forksPytorch re-implementation of Paper: SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition (CVPR 2022)
289 stars41 forksAutomatic Number Plate Detection YOLOv8
277 stars73 forksProceed with text detection only in the selected area of the image
254 stars48 forks[NeurIPS2023] This is the official code of the paper "GlyphControl: Glyph Conditional Control for Visual Text Generation"
238 stars16 forks[ICCV 2023] Code base for Revisiting Scene Text Recognition: A Data Perspective
205 stars9 forks[AAAI'23 Oral] DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer
202 stars28 forks#46
wenwenyu/TCM
Turning a CLIP Model into a Scene Text Detector (CVPR2023) | Turning a CLIP Model into a Scene Text Spotter (TPAMI)
201 stars20 forks[CVPR2023] Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution
199 stars29 forksAutomatic Number Plate Recognition pipeline — fast-alpr + fast-plate-ocr + FastAPI. Privacy-aware (HMAC + KVKK + GDPR). pip install anpr-pipeline
197 stars44 forksICDAR 2019: MaskRCNN on PubLayNet datasets. Paragraph detection, table detection, figure detection,...
183 stars38 forksHandwritten Text Recognition and Character Detection
166 stars19 forksA high-performance, open-source PDF data extraction tool. 一站式开源高性能数据提取工具,将复杂 PDF 文档转换为 Markdown 和 JSON 格式,使用onnx模型。
162 stars34 forksMinimal sharded dataset loaders, decoders, and utils for multi-modal document, image, and text datasets.
161 stars10 forksCodebase for fine-tuning / evaluating nougat-based image2latex generation models
160 stars19 forksYOLO models trained by DocLayNet - power your Document Intelligent by Layout Analysis
157 stars20 forksTableNet Implementation on Pytorch
150 stars43 forksHyperMotion is a pose guided human image animation framework based on a large-scale video diffusion Transformer.
147 stars12 forksThe official code of CornerTransformer (ECCV 2022, Oral) on top of MMOCR.
146 stars15 forks一个相对完整的文档分析和识别项目
144 stars38 forksThis repository is the code of our paper "DiffUTE: Universal Text Editing Diffusion Model" (NeurIPS'2023).
144 stars11 forksan OCR tool to translate Old Persian cuneiform (Achaemenid language) by AI
144 stars10 forksA convenient way to link, deduplicate, aggregate and cluster data(frames) in Python using deep learning
139 stars14 forksA toolbox for Vietnamese Optical Character Recognition.
138 stars33 forksPython3 package for Chinese/English OCR,use paddleocr-v5 onnx model(~20MB), with ultra-fast inference speed. 基于ppocr-v5-onnx模型推理,中英文OCR开源SOTA,推理速度超快。
131 stars21 forksVehicle-Rear: A New Dataset to Explore Feature Fusion For Vehicle Identification Using Convolutional Neural Networks
118 stars19 forksReinforcement Learning Framework for Visual Generation
118 stars7 forksswin-transformer custom for OCR
117 stars20 forksCurated collections of sample applications designed to help you develop optimized AI solutions. Tailored to specific use cases, covering retail, manufacturing, metro, and media & entertainment.
110 stars147 forksA comic localizer built with python
107 stars19 forksThe open-source universal adapter for LLMs. Turn messy real-world data into clean, agent-ready context.
105 stars14 forksA computer vision project for image segmentation
104 stars46 forksZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.
103 stars26 forksMTL-TabNet: Multi-task Learning based Model for Image-based Table Recognition
103 stars12 forksGCNet, official pytorch implementation of our paper "GCNet: Graph Completion Network for Incomplete Multimodal Learning in Conversation"
102 stars8 forks🔍 Table Extraction Tool: A powerful open-source solution combining OCR and computer vision for extracting structured tabular data from images. Ideal for LLM preprocessing, data analysis, and automation. 🚀
98 stars26 forksThis is the official code of the paper "A Multi-Agent System Enables Versatile Information Extraction from the Chemical Literature"
95 stars20 forksCheckbox Detection Model for Scanned Documents
94 stars6 forksMAT: Multi-modal Agent Tuning 🔥 ICLR 2025 (Spotlight)
94 stars6 forks- 93 stars23 forks
Complex data extraction and orchestration framework designed for processing unstructured documents. It integrates AI-powered document pipelines (GenAI, LLM, VLLM) into your applications, supporting various tasks such as document cleanup, optical character recognition (OCR), classification, splitting, named entity recognition, and form processing
92 stars11 forksRe-implementation of MASTER by mmocr
90 stars18 forks[ICML 2026]A Large-scale Dataset for training and evaluating model's ability on Dense Text Image Generation
90 stars0 forksSimple Pytorch framework to train OCRs. Supports CRNNs, Attention, CTC and Cross Entropy Loss.
88 stars17 forksAutomatic Manga Translator
84 stars25 forksOCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers
83 stars2 forksOfficial code for paper Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
81 stars0 forks(ICCV 2023) ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer
78 stars7 forksLicense plate recognition
73 stars8 forks[AAAI 2024] SRFormer: Text Detection Transformer with Incorporated Segmentation and Regression
70 stars9 forks[IEEE TPAMI 2025] Privacy-Preserving Biometric Verification With Handwritten Random Digit String
69 stars0 forks#90
qcf-568/OSTF
[AAAI2025] Revisiting Tampered Scene Text Detection in the Era of Generative AI
68 stars3 forks基于 BallonsTranslator 的漫画/网漫/韩漫/国漫计算机辅助翻译工具,扩展了 OCR、文本检测、|修复、工作流程、字体和导出选项。Computer-aided manga/comic/manhwa/manhua translation tool based on BallonsTranslator, with expanded OCR, detection, inpainting, workflow, font, and export options.
66 stars8 forksPytorch Implementation of TableNet
65 stars23 forksWeKnora‑pro是基于原始 WeKnora 的二次开发版本,核心在于提升文档解析能力。 主要改进:1. 支持扫描件通过 (CPU/GPU 自动优化)进行 OCR 与表格提取;且兼容WeKnora多模态增加 2. 文档大小上限提升至 300 MB
65 stars22 forksUnofficial implementation of ''BEDSR-Net: A Deep Shadow Removal from a Single Document Image'' with PyTorch
65 stars12 forksA model(ing framework) for sample efficient OCR
65 stars9 forksClassification of KYC documents and OCR extraction
63 stars20 forksFrom a chemical reaction image, detect and classify molecules, text and arrows by using the Vision Transformer DETR. Comparisons with well-established CNNs (RetinaNet and FasterRCNN also provided). The detections are then translated into text "OCR" or into SMILES. The direction of the reaction is learned and preserved into the output files.
63 stars6 forks[CVPR 2026] Official data synthesis code of the paper "What’s Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution".
63 stars0 forks- 61 stars14 forks
[CVPR 26] MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
60 stars11 forks