Documents and OCR
GitHub

Albumentations on GitHub in Documents and OCR

Used in 549 public GitHub repositories

Document understanding, OCR, handwriting, scene text, and layout analysis.

Hands writing on a document beside books and a desk lamp
Photo: Kindel Media on Pexels

Public examples

Top public GitHub repositories

Showing the top 100

  1. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

    81,114 stars11,056 forks
  2. pix2tex: Using a ViT to convert images of equations into LaTeX code.

    16,449 stars1,308 forks
  3. Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

    15,181 stars1,277 forks
  4. Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)

    9,999 stars1,035 forks
  5. Effortless data labeling with AI support from Segment Anything and other awesome models.

    9,343 stars1,067 forks
  6. The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

    9,003 stars775 forks
  7. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

    8,132 stars706 forks
  8. All-in-One Development Tool based on PaddlePaddle

    6,160 stars1,208 forks
  9. PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

    6,015 stars417 forks
  10. OpenMMLab Text Detection, Recognition and Understanding Toolbox

    4,736 stars779 forks
  11. CnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】

    3,764 stars539 forks
  12. 🛠 All-in-one web-based IDE specialized for machine learning and data science.

    3,540 stars458 forks
  13. InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)

    3,206 stars233 forks
  14. An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

    3,148 stars275 forks
  15. Optical character recognition for Japanese text, with the main focus being Japanese manga

    2,677 stars135 forks
  16. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

    2,175 stars170 forks
  17. 基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.

    1,804 stars200 forks
  18. 基于 manga-image-translator 的开源漫画翻译工具。支持日/韩/美漫自动翻译,内置 OpenAI、Gemini 等 5 种翻译引擎,并提供可视化编辑器自由调整文本样式。一键安装,开箱即用。如果喜欢,欢迎点亮 ⭐ Star 支持!

    1,748 stars107 forks
  19. OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

    1,360 stars134 forks
  20. A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.

    1,237 stars128 forks
  21. Parsing-free RAG supported by VLMs

    960 stars76 forks
  22. Generate text line images for training deep learning OCR models

    912 stars175 forks
  23. A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。

    826 stars51 forks
  24. CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包

    792 stars115 forks
  25. Transformer OCR

    779 stars252 forks
  26. OCR toolbox from Davar-Lab

    760 stars152 forks
  27. 天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取

    657 stars113 forks
  28. [Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting

    586 stars42 forks
  29. Lightweight & fast OCR models for license plate text recognition.

    578 stars77 forks
  30. Official Code for ECCV 2024 paper — One-Shot Diffusion Mimicker for Handwritten Text Generation

    555 stars56 forks
  31. A scene text recognition toolbox based on PyTorch

    534 stars99 forks
  32. 2nd solution of ICDAR 2021 Competition on Scientific Literature Parsing, Task B.

    470 stars106 forks
  33. PLPR utilizes YOLOv5 and custom models for high-accuracy Persian license plate recognition, featuring real-time processing and an intuitive interface in an open-source framework.

    449 stars126 forks
  34. crnn chinese_plate_recognition

    404 stars80 forks
  35. Code base of 'PAPER2WEB: LET’S MAKE YOUR PAPER ALIVE!' integrating the other work package throughout the entire "paper2present" process.

    377 stars30 forks
  36. MakeACopy is an open-source document scanner app for Android that allows you to digitize paper documents with OCR functionality. The app is designed to be privacy-friendly, working completely offline without any cloud connection or tracking.

    371 stars28 forks
  37. The Infosys Responsible AI toolkit incorporates various features including safety, security, explainability, fairness, bias and hallucination detection to ensure AI solutions are trustworthy and transparent.

    297 stars74 forks
  38. MLEvolve is an open-source autonomous system for end-to-end machine learning algorithm design and optimization powered by progressive search and experience-driven memory.

    293 stars54 forks
  39. The official repo for [CVPR'23] "DeepSolo: Let Transformer Decoder with Explicit Points Solo for Text Spotting" & [ArXiv'23] "DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting"

    290 stars44 forks
  40. Pytorch re-implementation of Paper: SwinTextSpotter: Scene Text Spotting via Better Synergy between Text Detection and Text Recognition (CVPR 2022)

    289 stars41 forks
  41. Automatic Number Plate Detection YOLOv8

    277 stars73 forks
  42. Proceed with text detection only in the selected area of ​​the image

    254 stars48 forks
  43. [NeurIPS2023] This is the official code of the paper "GlyphControl: Glyph Conditional Control for Visual Text Generation"

    238 stars16 forks
  44. [ICCV 2023] Code base for Revisiting Scene Text Recognition: A Data Perspective

    205 stars9 forks
  45. [AAAI'23 Oral] DPText-DETR: Towards Better Scene Text Detection with Dynamic Points in Transformer

    202 stars28 forks
  46. Turning a CLIP Model into a Scene Text Detector (CVPR2023) | Turning a CLIP Model into a Scene Text Spotter (TPAMI)

    201 stars20 forks
  47. [CVPR2023] Towards Robust Tampered Text Detection in Document Image: New Dataset and New Solution

    199 stars29 forks
  48. Automatic Number Plate Recognition pipeline — fast-alpr + fast-plate-ocr + FastAPI. Privacy-aware (HMAC + KVKK + GDPR). pip install anpr-pipeline

    197 stars44 forks
  49. ICDAR 2019: MaskRCNN on PubLayNet datasets. Paragraph detection, table detection, figure detection,...

    183 stars38 forks
  50. Handwritten Text Recognition and Character Detection

    166 stars19 forks
  51. A high-performance, open-source PDF data extraction tool. 一站式开源高性能数据提取工具,将复杂 PDF 文档转换为 Markdown 和 JSON 格式,使用onnx模型。

    162 stars34 forks
  52. Minimal sharded dataset loaders, decoders, and utils for multi-modal document, image, and text datasets.

    161 stars10 forks
  53. Codebase for fine-tuning / evaluating nougat-based image2latex generation models

    160 stars19 forks
  54. YOLO models trained by DocLayNet - power your Document Intelligent by Layout Analysis

    157 stars20 forks
  55. TableNet Implementation on Pytorch

    150 stars43 forks
  56. HyperMotion is a pose guided human image animation framework based on a large-scale video diffusion Transformer.

    147 stars12 forks
  57. The official code of CornerTransformer (ECCV 2022, Oral) on top of MMOCR.

    146 stars15 forks
  58. 一个相对完整的文档分析和识别项目

    144 stars38 forks
  59. This repository is the code of our paper "DiffUTE: Universal Text Editing Diffusion Model" (NeurIPS'2023).

    144 stars11 forks
  60. an OCR tool to translate Old Persian cuneiform (Achaemenid language) by AI

    144 stars10 forks
  61. A convenient way to link, deduplicate, aggregate and cluster data(frames) in Python using deep learning

    139 stars14 forks
  62. A toolbox for Vietnamese Optical Character Recognition.

    138 stars33 forks
  63. Python3 package for Chinese/English OCR,use paddleocr-v5 onnx model(~20MB), with ultra-fast inference speed. 基于ppocr-v5-onnx模型推理,中英文OCR开源SOTA,推理速度超快。

    131 stars21 forks
  64. Vehicle-Rear: A New Dataset to Explore Feature Fusion For Vehicle Identification Using Convolutional Neural Networks

    118 stars19 forks
  65. Reinforcement Learning Framework for Visual Generation

    118 stars7 forks
  66. swin-transformer custom for OCR

    117 stars20 forks
  67. Curated collections of sample applications designed to help you develop optimized AI solutions. Tailored to specific use cases, covering retail, manufacturing, metro, and media & entertainment.

    110 stars147 forks
  68. A comic localizer built with python

    107 stars19 forks
  69. The open-source universal adapter for LLMs. Turn messy real-world data into clean, agent-ready context.

    105 stars14 forks
  70. A computer vision project for image segmentation

    104 stars46 forks
  71. ZerolanCore integrates many open-source, locally deployable AI models, and aims to integrate a series of AI models such as large language model (LLM), automatic speech recognition (ASR), text-to-speech (TTS), image captioning, optical character recognition (OCR), video captioning, etc.

    103 stars26 forks
  72. MTL-TabNet: Multi-task Learning based Model for Image-based Table Recognition

    103 stars12 forks
  73. GCNet, official pytorch implementation of our paper "GCNet: Graph Completion Network for Incomplete Multimodal Learning in Conversation"

    102 stars8 forks
  74. 🔍 Table Extraction Tool: A powerful open-source solution combining OCR and computer vision for extracting structured tabular data from images. Ideal for LLM preprocessing, data analysis, and automation. 🚀

    98 stars26 forks
  75. This is the official code of the paper "A Multi-Agent System Enables Versatile Information Extraction from the Chemical Literature"

    95 stars20 forks
  76. Checkbox Detection Model for Scanned Documents

    94 stars6 forks
  77. MAT: Multi-modal Agent Tuning 🔥 ICLR 2025 (Spotlight)

    94 stars6 forks
  78. 93 stars23 forks
  79. Complex data extraction and orchestration framework designed for processing unstructured documents. It integrates AI-powered document pipelines (GenAI, LLM, VLLM) into your applications, supporting various tasks such as document cleanup, optical character recognition (OCR), classification, splitting, named entity recognition, and form processing

    92 stars11 forks
  80. Re-implementation of MASTER by mmocr

    90 stars18 forks
  81. [ICML 2026]A Large-scale Dataset for training and evaluating model's ability on Dense Text Image Generation

    90 stars0 forks
  82. Simple Pytorch framework to train OCRs. Supports CRNNs, Attention, CTC and Cross Entropy Loss.

    88 stars17 forks
  83. Automatic Manga Translator

    84 stars25 forks
  84. OCR-VQGAN, a discrete image encoder (tokenizer and detokenizer) for figure images in Paper2Fig100k dataset. Implementation of OCR Perceptual loss for clear text-within-image generation. Fork from VQGAN in CompVis/taming-transformers

    83 stars2 forks
  85. Official code for paper Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

    81 stars0 forks
  86. (ICCV 2023) ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer

    78 stars7 forks
  87. License plate recognition

    73 stars8 forks
  88. [AAAI 2024] SRFormer: Text Detection Transformer with Incorporated Segmentation and Regression

    70 stars9 forks
  89. [IEEE TPAMI 2025] Privacy-Preserving Biometric Verification With Handwritten Random Digit String

    69 stars0 forks
  90. [AAAI2025] Revisiting Tampered Scene Text Detection in the Era of Generative AI

    68 stars3 forks
  91. 基于 BallonsTranslator 的漫画/网漫/韩漫/国漫计算机辅助翻译工具,扩展了 OCR、文本检测、|修复、工作流程、字体和导出选项。Computer-aided manga/comic/manhwa/manhua translation tool based on BallonsTranslator, with expanded OCR, detection, inpainting, workflow, font, and export options.

    66 stars8 forks
  92. Pytorch Implementation of TableNet

    65 stars23 forks
  93. WeKnora‑pro是基于原始 WeKnora 的二次开发版本,核心在于提升文档解析能力。 主要改进:1. 支持扫描件通过 (CPU/GPU 自动优化)进行 OCR 与表格提取;且兼容WeKnora多模态增加 2. 文档大小上限提升至 300 MB

    65 stars22 forks
  94. Unofficial implementation of ''BEDSR-Net: A Deep Shadow Removal from a Single Document Image'' with PyTorch

    65 stars12 forks
  95. A model(ing framework) for sample efficient OCR

    65 stars9 forks
  96. Classification of KYC documents and OCR extraction

    63 stars20 forks
  97. From a chemical reaction image, detect and classify molecules, text and arrows by using the Vision Transformer DETR. Comparisons with well-established CNNs (RetinaNet and FasterRCNN also provided). The detections are then translated into text "OCR" or into SMILES. The direction of the reaction is learned and preserved into the output files.

    63 stars6 forks
  98. [CVPR 2026] Official data synthesis code of the paper "What’s Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution".

    63 stars0 forks
  99. 61 stars14 forks
  100. [CVPR 26] MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

    60 stars11 forks