Albumentations on GitHub in Documents and OCR
Used in 495 public GitHub repositories
Document understanding, OCR, handwriting, scene text, and layout analysis.

Public examples
Top public GitHub repositories
Showing the top 100
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
90,495 stars11,434 forksOCR, layout analysis, reading order, table recognition in 90+ languages
21,431 stars1,550 forkspix2tex: Using a ViT to convert images of equations into LaTeX code.
16,581 stars1,311 forksConvert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
15,181 stars1,277 forksX-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
10,579 stars1,171 forksTranslate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)
10,464 stars1,035 forksThe official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
9,059 stars777 forksOfficial code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
8,227 stars716 forksAll-in-One Development Tool based on PaddlePaddle
6,270 stars1,208 forksPDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
6,015 stars417 forksOpenMMLab Text Detection, Recognition and Understanding Toolbox
4,752 stars779 forksCnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】
3,764 stars539 forksAn Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
3,265 stars283 forks基于manga-image-translator 实现的开源漫画AI翻译桌面工具。支持日、韩、英文漫画自动处理,集成OpenAl、Gemini等多翻译引擎;实现OCR文字检测、原文擦除、AI翻译、图像修复、译文排版完整链路,自带可视化编辑器,支持自定义文本样式,一键部署开箱即用。
2,893 stars138 forksOptical character recognition for Japanese text, with the main focus being Japanese manga
2,793 stars142 forksDocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
2,286 stars177 forks基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.
1,874 stars200 forksOpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
1,464 stars149 forksA central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.
1,235 stars127 forksEverything about ComfyUI, including workflow sharing, resource sharing, knowledge sharing, tutorial sharing, and more.关于ComfyUI的一切,工作流分享、资源分享、知识分享、教程分享等
1,181 stars168 forksTurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
1,098 stars113 forksParsing-free RAG supported by VLMs
981 stars76 forksGenerate text line images for training deep learning OCR models
915 stars174 forks#24
alephpi/Texo
A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。
903 stars53 forks天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取
830 stars126 forksTransformer OCR
817 stars257 forksCnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包
794 stars114 forksLightweight & fast OCR models for license plate text recognition.
766 stars99 forksOCR toolbox from Davar-Lab
762 stars152 forks[Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting
590 stars42 forksOfficial Code for ECCV 2024 paper — One-Shot Diffusion Mimicker for Handwritten Text Generation
566 stars57 forksMakeACopy is an open-source document scanner app for Android that allows you to digitize paper documents with OCR functionality. The app is designed to be privacy-friendly, working completely offline without any cloud connection or tracking.
545 stars37 forksA scene text recognition toolbox based on PyTorch
532 stars99 forks2nd solution of ICDAR 2021 Competition on Scientific Literature Parsing, Task B.
470 stars105 forksPLPR utilizes YOLOv5 and custom models for high-accuracy Persian license plate recognition, featuring real-time processing and an intuitive interface in an open-source framework.
449 stars126 forkscrnn chinese_plate_recognition
423 stars80 forksPure-Rust, CPU-only OCR engine for Baidu Unlimited-OCR (a DeepSeek-OCR-derived 3B MoE VLM). Five-model zoo, custom int8 kernels, no ML framework, no Python, no GPU.
332 stars39 forksThe official repo for [CVPR'23] "DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting" & [ArXiv'23] "DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting"
296 stars44 forksPytorch re-implementation of Paper: SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition (CVPR 2022)
287 stars41 forksAutomatic Number Plate Detection YOLOv8
279 stars72 forksProceed with text detection only in the selected area of the image
254 stars48 forks[NeurIPS2023] This is the official code of the paper "GlyphControl: Glyph Conditional Control for Visual Text Generation"
237 stars16 forksA high-performance, open-source PDF data extraction tool. 一站式开源高性能数据提取工具,将复杂 PDF 文档转换为 Markdown 和 JSON 格式,使用onnx模型。
232 stars34 forks[CVPR2023] Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution
220 stars29 forks[AAAI'23 Oral] DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer
205 stars28 forks[ICCV 2023] Code base for Revisiting Scene Text Recognition: A Data Perspective
204 stars9 forksAutomatic Number Plate Recognition pipeline — fast-alpr + fast-plate-ocr + FastAPI. Privacy-aware (HMAC + KVKK + GDPR). pip install anpr-pipeline
202 stars46 forks#48
wenwenyu/TCM
Turning a CLIP Model into a Scene Text Detector (CVPR2023) | Turning a CLIP Model into a Scene Text Spotter (TPAMI)
200 stars20 forksICDAR 2019: MaskRCNN on PubLayNet datasets. Paragraph detection, table detection, figure detection,...
180 stars37 forksHandwritten Text Recognition and Character Detection
174 stars20 forksYOLO models trained by DocLayNet - power your Document Intelligent by Layout Analysis
165 stars20 forksCodebase for fine-tuning / evaluating nougat-based image2latex generation models
162 stars19 forksMinimal sharded dataset loaders, decoders, and utils for multi-modal document, image, and text datasets.
162 stars10 forksTableNet Implementation on Pytorch
149 stars43 forksThe official code of CornerTransformer (ECCV 2022, Oral) on top of MMOCR.
148 stars15 forksan OCR tool to translate Old Persian cuneiform (Achaemenid language) by AI
147 stars10 forks一个相对完整的文档分析和识别项目
144 stars38 forksThis repository is the code of our paper "DiffUTE: Universal Text Editing Diffusion Model" (NeurIPS'2023).
144 stars12 forksA toolbox for Vietnamese Optical Character Recognition.
143 stars33 forksPython3 package for Chinese/English OCR,use paddleocr-v5 onnx model(~20MB), with ultra-fast inference speed. 基于ppocr-v5-onnx模型推理,中英文OCR开源SOTA,推理速度超快。
134 stars21 forksWeKnora‑pro是基于原始 WeKnora 的二次开发版本,核心在于提升文档解析能力。 主要改进:1. 支持扫描件通过 (CPU/GPU 自动优化)进行 OCR 与表格提取;且兼容WeKnora多模态增加 2. 文档大小上限提升至 300 MB
132 stars22 forksVehicle-Rear: A New Dataset to Explore Feature Fusion For Vehicle Identification Using Convolutional Neural Networks
119 stars19 forksZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.
117 stars26 forksswin-transformer custom for OCR
116 stars20 forks- 111 stars2 forks
A comic localizer built with python
108 stars19 forksThe open-source universal adapter for LLMs. Turn messy real-world data into clean, agent-ready context.
107 stars14 forks🔍 Table Extraction Tool: A powerful open-source solution combining OCR and computer vision for extracting structured tabular data from images. Ideal for LLM preprocessing, data analysis, and automation. 🚀
104 stars26 forksMTL-TabNet: Multi-task Learning based Model for Image-based Table Recognition
103 stars12 forksCheckbox Detection Model for Scanned Documents
97 stars6 forksComplex data extraction and orchestration framework designed for processing unstructured documents. It integrates AI-powered document pipelines (GenAI, LLM, VLLM) into your applications, supporting various tasks such as document cleanup, optical character recognition (OCR), classification, splitting, named entity recognition, and form processing
96 stars13 forks基于 BallonsTranslator 的漫画/网漫/韩漫/国漫计算机辅助翻译工具,扩展了 OCR、文本检测、|修复、工作流程、字体和导出选项。Computer-aided manga/comic/manhwa/manhua translation tool based on BallonsTranslator, with expanded OCR, detection, inpainting, workflow, font, and export options.
96 stars8 forksRe-implementation of MASTER by mmocr
89 stars18 forksSimple Pytorch framework to train OCRs. Supports CRNNs, Attention, CTC and Cross Entropy Loss.
89 stars17 forksAutomatic Manga Translator
87 stars25 forksDeepParseX 是一个强大的多模态文档解析与知识管理平台,支持 PDF、Word、Excel、PPT、图片、视频、音频 等多种文件格式的智能解析,自动提取关键信息,并构建 检索增强生成(RAG) 和 知识图谱(Knowledge Graph) 系统,实现结构化数据的智能检索与推理。
87 stars17 forksOCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers
86 stars2 forksZotero AI plugin Research assistant for Zotero 9. Chat with your library, run federated scholarly search, RAG, OCR, systematic reviews, and manage cloud storage. Includes standalone MCP, Agentic capabilities, and skills library.
83 stars5 forks(ICCV 2023) ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer
78 stars7 forks#80
qcf-568/OSTF
[AAAI2025] Revisiting Tampered Scene Text Detection in the Era of Generative AI
77 stars4 forksLocal, privacy-first CAPTCHA recognition for Chrome and Edge. 本地、隐私优先的验证码识别扩展。
76 stars11 forksLicense plate recognition
73 stars8 forks[IEEE TPAMI 2025] Privacy-Preserving Biometric Verification With Handwritten Random Digit String
71 stars0 forks[CVPR 2026] Official data synthesis code of the paper "What’s Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution".
71 stars0 forks[AAAI 2024] SRFormer: Text Detection Transformer with Incorporated Segmentation and Regression
70 stars9 forksUnofficial implementation of ''BEDSR-Net: A Deep Shadow Removal from a Single Document Image'' with PyTorch
67 stars12 forksA model(ing framework) for sample efficient OCR
66 stars9 forksPytorch Implementation of TableNet
65 stars23 forksClassification of KYC documents and OCR extraction
64 stars20 forksFrom a chemical reaction image, detect and classify molecules, text and arrows by using the Vision Transformer DETR. Comparisons with well-established CNNs (RetinaNet and FasterRCNN also provided). The detections are then translated into text "OCR" or into SMILES. The direction of the reaction is learned and preserved into the output files.
64 stars6 forksAn online document authentication portal used to detect morphed images, handwriting forgeries, fake certificates, ID proofs and all the documents issued by the Government on the go.
62 stars25 forks- 61 stars14 forks
💡💡💡awesome compute vision app in gradio
60 stars9 forks[IEEE TIFS 2024] Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach
60 stars1 forks#95
uakarsh/latr
Implementation of LaTr: Layout-aware transformer for scene-text VQA,a novel multimodal architecture for Scene Text Visual Question Answering (STVQA)
56 stars6 forksLarge scale training of Latex formula recognition model, currently being organized and open source
56 stars4 forksThe official code for the CVPR 2024 paper: Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
55 stars5 forksUnofficial implementation of the paper "Full Page Handwriting Recognition via Image to Sequence Extraction" by Singh et al. (2021).
54 stars6 forksAn easy-to-run OCR model pipeline based on CRNN and CTC loss
49 stars20 forkspytorch implementation of crnn. A sample training of license plate is provided.
48 stars9 forks