Documents and OCR
GitHub

Albumentations on GitHub in Documents and OCR

Used in 495 public GitHub repositories

Document understanding, OCR, handwriting, scene text, and layout analysis.

Hands writing on a document beside books and a desk lamp
Photo: Kindel Media on Pexels

Public examples

Top public GitHub repositories

Showing the top 100

  1. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

    90,495 stars11,434 forks
  2. OCR, layout analysis, reading order, table recognition in 90+ languages

    21,431 stars1,550 forks
  3. pix2tex: Using a ViT to convert images of equations into LaTeX code.

    16,581 stars1,311 forks
  4. Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

    15,181 stars1,277 forks
  5. X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

    10,579 stars1,171 forks
  6. Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)

    10,464 stars1,035 forks
  7. The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

    9,059 stars777 forks
  8. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

    8,227 stars716 forks
  9. All-in-One Development Tool based on PaddlePaddle

    6,270 stars1,208 forks
  10. PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

    6,015 stars417 forks
  11. OpenMMLab Text Detection, Recognition and Understanding Toolbox

    4,752 stars779 forks
  12. CnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】

    3,764 stars539 forks
  13. An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

    3,265 stars283 forks
  14. 基于manga-image-translator 实现的开源漫画AI翻译桌面工具。支持日、韩、英文漫画自动处理,集成OpenAl、Gemini等多翻译引擎;实现OCR文字检测、原文擦除、AI翻译、图像修复、译文排版完整链路,自带可视化编辑器,支持自定义文本样式,一键部署开箱即用。

    2,893 stars138 forks
  15. Optical character recognition for Japanese text, with the main focus being Japanese manga

    2,793 stars142 forks
  16. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

    2,286 stars177 forks
  17. 基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.

    1,874 stars200 forks
  18. OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

    1,464 stars149 forks
  19. A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.

    1,235 stars127 forks
  20. Everything about ComfyUI, including workflow sharing, resource sharing, knowledge sharing, tutorial sharing, and more.关于ComfyUI的一切,工作流分享、资源分享、知识分享、教程分享等

    1,181 stars168 forks
  21. TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC

    1,098 stars113 forks
  22. Parsing-free RAG supported by VLMs

    981 stars76 forks
  23. Generate text line images for training deep learning OCR models

    915 stars174 forks
  24. A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。

    903 stars53 forks
  25. 天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取

    830 stars126 forks
  26. Transformer OCR

    817 stars257 forks
  27. CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包

    794 stars114 forks
  28. Lightweight & fast OCR models for license plate text recognition.

    766 stars99 forks
  29. OCR toolbox from Davar-Lab

    762 stars152 forks
  30. [Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting

    590 stars42 forks
  31. Official Code for ECCV 2024 paper — One-Shot Diffusion Mimicker for Handwritten Text Generation

    566 stars57 forks
  32. MakeACopy is an open-source document scanner app for Android that allows you to digitize paper documents with OCR functionality. The app is designed to be privacy-friendly, working completely offline without any cloud connection or tracking.

    545 stars37 forks
  33. A scene text recognition toolbox based on PyTorch

    532 stars99 forks
  34. 2nd solution of ICDAR 2021 Competition on Scientific Literature Parsing, Task B.

    470 stars105 forks
  35. PLPR utilizes YOLOv5 and custom models for high-accuracy Persian license plate recognition, featuring real-time processing and an intuitive interface in an open-source framework.

    449 stars126 forks
  36. crnn chinese_plate_recognition

    423 stars80 forks
  37. Pure-Rust, CPU-only OCR engine for Baidu Unlimited-OCR (a DeepSeek-OCR-derived 3B MoE VLM). Five-model zoo, custom int8 kernels, no ML framework, no Python, no GPU.

    332 stars39 forks
  38. The official repo for [CVPR'23] "DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting" & [ArXiv'23] "DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting"

    296 stars44 forks
  39. Pytorch re-implementation of Paper: SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition (CVPR 2022)

    287 stars41 forks
  40. Automatic Number Plate Detection YOLOv8

    279 stars72 forks
  41. Proceed with text detection only in the selected area of ​​the image

    254 stars48 forks
  42. [NeurIPS2023] This is the official code of the paper "GlyphControl: Glyph Conditional Control for Visual Text Generation"

    237 stars16 forks
  43. A high-performance, open-source PDF data extraction tool. 一站式开源高性能数据提取工具,将复杂 PDF 文档转换为 Markdown 和 JSON 格式,使用onnx模型。

    232 stars34 forks
  44. [CVPR2023] Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution

    220 stars29 forks
  45. [AAAI'23 Oral] DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer

    205 stars28 forks
  46. [ICCV 2023] Code base for Revisiting Scene Text Recognition: A Data Perspective

    204 stars9 forks
  47. Automatic Number Plate Recognition pipeline — fast-alpr + fast-plate-ocr + FastAPI. Privacy-aware (HMAC + KVKK + GDPR). pip install anpr-pipeline

    202 stars46 forks
  48. Turning a CLIP Model into a Scene Text Detector (CVPR2023) | Turning a CLIP Model into a Scene Text Spotter (TPAMI)

    200 stars20 forks
  49. ICDAR 2019: MaskRCNN on PubLayNet datasets. Paragraph detection, table detection, figure detection,...

    180 stars37 forks
  50. Handwritten Text Recognition and Character Detection

    174 stars20 forks
  51. YOLO models trained by DocLayNet - power your Document Intelligent by Layout Analysis

    165 stars20 forks
  52. Codebase for fine-tuning / evaluating nougat-based image2latex generation models

    162 stars19 forks
  53. Minimal sharded dataset loaders, decoders, and utils for multi-modal document, image, and text datasets.

    162 stars10 forks
  54. TableNet Implementation on Pytorch

    149 stars43 forks
  55. The official code of CornerTransformer (ECCV 2022, Oral) on top of MMOCR.

    148 stars15 forks
  56. an OCR tool to translate Old Persian cuneiform (Achaemenid language) by AI

    147 stars10 forks
  57. 一个相对完整的文档分析和识别项目

    144 stars38 forks
  58. This repository is the code of our paper "DiffUTE: Universal Text Editing Diffusion Model" (NeurIPS'2023).

    144 stars12 forks
  59. A toolbox for Vietnamese Optical Character Recognition.

    143 stars33 forks
  60. Python3 package for Chinese/English OCR,use paddleocr-v5 onnx model(~20MB), with ultra-fast inference speed. 基于ppocr-v5-onnx模型推理,中英文OCR开源SOTA,推理速度超快。

    134 stars21 forks
  61. WeKnora‑pro是基于原始 WeKnora 的二次开发版本,核心在于提升文档解析能力。 主要改进:1. 支持扫描件通过 (CPU/GPU 自动优化)进行 OCR 与表格提取;且兼容WeKnora多模态增加 2. 文档大小上限提升至 300 MB

    132 stars22 forks
  62. Vehicle-Rear: A New Dataset to Explore Feature Fusion For Vehicle Identification Using Convolutional Neural Networks

    119 stars19 forks
  63. ZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.

    117 stars26 forks
  64. swin-transformer custom for OCR

    116 stars20 forks
  65. 111 stars2 forks
  66. A comic localizer built with python

    108 stars19 forks
  67. The open-source universal adapter for LLMs. Turn messy real-world data into clean, agent-ready context.

    107 stars14 forks
  68. 🔍 Table Extraction Tool: A powerful open-source solution combining OCR and computer vision for extracting structured tabular data from images. Ideal for LLM preprocessing, data analysis, and automation. 🚀

    104 stars26 forks
  69. MTL-TabNet: Multi-task Learning based Model for Image-based Table Recognition

    103 stars12 forks
  70. Checkbox Detection Model for Scanned Documents

    97 stars6 forks
  71. Complex data extraction and orchestration framework designed for processing unstructured documents. It integrates AI-powered document pipelines (GenAI, LLM, VLLM) into your applications, supporting various tasks such as document cleanup, optical character recognition (OCR), classification, splitting, named entity recognition, and form processing

    96 stars13 forks
  72. 基于 BallonsTranslator 的漫画/网漫/韩漫/国漫计算机辅助翻译工具,扩展了 OCR、文本检测、|修复、工作流程、字体和导出选项。Computer-aided manga/comic/manhwa/manhua translation tool based on BallonsTranslator, with expanded OCR, detection, inpainting, workflow, font, and export options.

    96 stars8 forks
  73. Re-implementation of MASTER by mmocr

    89 stars18 forks
  74. Simple Pytorch framework to train OCRs. Supports CRNNs, Attention, CTC and Cross Entropy Loss.

    89 stars17 forks
  75. Automatic Manga Translator

    87 stars25 forks
  76. DeepParseX 是一个强大的多模态文档解析与知识管理平台,支持 PDF、Word、Excel、PPT、图片、视频、音频 等多种文件格式的智能解析,自动提取关键信息,并构建 检索增强生成(RAG) 和 知识图谱(Knowledge Graph) 系统,实现结构化数据的智能检索与推理。

    87 stars17 forks
  77. OCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers

    86 stars2 forks
  78. Zotero AI plugin Research assistant for Zotero 9. Chat with your library, run federated scholarly search, RAG, OCR, systematic reviews, and manage cloud storage. Includes standalone MCP, Agentic capabilities, and skills library.

    83 stars5 forks
  79. (ICCV 2023) ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer

    78 stars7 forks
  80. [AAAI2025] Revisiting Tampered Scene Text Detection in the Era of Generative AI

    77 stars4 forks
  81. Local, privacy-first CAPTCHA recognition for Chrome and Edge. 本地、隐私优先的验证码识别扩展。

    76 stars11 forks
  82. License plate recognition

    73 stars8 forks
  83. [IEEE TPAMI 2025] Privacy-Preserving Biometric Verification With Handwritten Random Digit String

    71 stars0 forks
  84. [CVPR 2026] Official data synthesis code of the paper "What’s Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution".

    71 stars0 forks
  85. [AAAI 2024] SRFormer: Text Detection Transformer with Incorporated Segmentation and Regression

    70 stars9 forks
  86. Unofficial implementation of ''BEDSR-Net: A Deep Shadow Removal from a Single Document Image'' with PyTorch

    67 stars12 forks
  87. A model(ing framework) for sample efficient OCR

    66 stars9 forks
  88. Pytorch Implementation of TableNet

    65 stars23 forks
  89. Classification of KYC documents and OCR extraction

    64 stars20 forks
  90. From a chemical reaction image, detect and classify molecules, text and arrows by using the Vision Transformer DETR. Comparisons with well-established CNNs (RetinaNet and FasterRCNN also provided). The detections are then translated into text "OCR" or into SMILES. The direction of the reaction is learned and preserved into the output files.

    64 stars6 forks
  91. An online document authentication portal used to detect morphed images, handwriting forgeries, fake certificates, ID proofs and all the documents issued by the Government on the go.

    62 stars25 forks
  92. 61 stars14 forks
  93. 💡💡💡awesome compute vision app in gradio

    60 stars9 forks
  94. [IEEE TIFS 2024] Online Writer Retrieval with Chinese Handwritten Phrases: A Synergistic Temporal-Frequency Representation Learning Approach

    60 stars1 forks
  95. Implementation of LaTr: Layout-aware transformer for scene-text VQA,a novel multimodal architecture for Scene Text Visual Question Answering (STVQA)

    56 stars6 forks
  96. Large scale training of Latex formula recognition model, currently being organized and open source

    56 stars4 forks
  97. The official code for the CVPR 2024 paper: Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer

    55 stars5 forks
  98. Unofficial implementation of the paper "Full Page Handwriting Recognition via Image to Sequence Extraction" by Singh et al. (2021).

    54 stars6 forks
  99. An easy-to-run OCR model pipeline based on CRNN and CTC loss

    49 stars20 forks
  100. pytorch implementation of crnn. A sample training of license plate is provided.

    48 stars9 forks