Adoption
Public code

GitHub adoption

Public GitHub repositories using Albumentations, ranked by stars.

Public examples

Top public GitHub repositories

Showing the top 100

  1. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

    81,114 stars11,056 forks
  2. State-of-the-art 2D and 3D Face Analysis Project

    28,784 stars6,062 forks
  3. The fastai deep learning library

    28,011 stars7,653 forks
  4. Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 160,000+ scientists worldwide. 148 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.

    26,687 stars3,130 forks
  5. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

    24,737 stars5,154 forks
  6. A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python

    22,929 stars3,141 forks
  7. 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

    21,620 stars3,304 forks
  8. pix2tex: Using a ViT to convert images of equations into LaTeX code.

    16,449 stars1,308 forks
  9. Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

    15,181 stars1,277 forks
  10. 🐍 Geometric Computer Vision Library for Spatial AI

    11,238 stars1,205 forks
  11. Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)

    9,999 stars1,035 forks
  12. OpenMMLab Semantic Segmentation Toolbox and Benchmark.

    9,832 stars2,852 forks
  13. Easy-to-use image segmentation library with awesome pre-trained model zoo, supporting wide-range of practical tasks in Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, 3D Segmentation, etc.

    9,343 stars1,710 forks
  14. Effortless data labeling with AI support from Segment Anything and other awesome models.

    9,343 stars1,067 forks
  15. The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

    9,003 stars775 forks
  16. AI Toolkit for Healthcare Imaging

    8,273 stars1,567 forks
  17. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

    8,132 stars706 forks
  18. Democratizing Deep-Learning for Drug Discovery, Quantum Chemistry, Materials Science and Biology

    6,761 stars2,263 forks
  19. OpenMMLab's next-generation platform for general 3D object detection.

    6,439 stars1,783 forks
  20. All-in-One Development Tool based on PaddlePaddle

    6,160 stars1,208 forks
  21. PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

    6,015 stars417 forks
  22. An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference.

    5,818 stars962 forks
  23. OpenMMLab Text Detection, Recognition and Understanding Toolbox

    4,736 stars779 forks
  24. TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data

    4,059 stars562 forks
  25. A comprehensive list of Deep Learning / Artificial Intelligence and Machine Learning tutorials - rapidly expanding into areas of AI/Deep Learning / Machine Vision / NLP and industry specific areas such as Climate / Energy, Automotives, Retail, Pharma, Medicine, Healthcare, Policy, Ethics and more.

    3,992 stars1,636 forks
  26. DORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.

    3,768 stars418 forks
  27. CnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】

    3,764 stars539 forks
  28. RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI

    3,587 stars608 forks
  29. 🛠 All-in-one web-based IDE specialized for machine learning and data science.

    3,540 stars458 forks
  30. Open source hardware and software platform to build a small scale self driving car.

    3,447 stars1,366 forks
  31. InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)

    3,206 stars233 forks
  32. An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

    3,148 stars275 forks
  33. GeoAI: Artificial Intelligence for Geospatial Data

    3,091 stars456 forks
  34. Optical character recognition for Japanese text, with the main focus being Japanese manga

    2,677 stars135 forks
  35. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

    2,634 stars414 forks
  36. The largest open-source medical AI skills library for OpenClaw🦞.

    2,575 stars402 forks
  37. Jupyter Notebook tutorials on solving real-world problems with Machine Learning & Deep Learning using PyTorch. Topics: Face detection with Detectron 2, Time Series anomaly detection with LSTM Autoencoders, Object Detection with YOLO v5, Build your first Neural Network, Time Series forecasting for Coronavirus daily cases, Sentiment Analysis with BER

    2,494 stars643 forks
  38. An open source library and framework for deep learning on satellite and aerial imagery.

    2,221 stars399 forks
  39. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

    2,175 stars170 forks
  40. 仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能

    2,076 stars283 forks
  41. yolov5 + csl_label.(Oriented Object Detection)(Rotation Detection)(Rotated BBox)基于yolov5的旋转目标检测

    1,945 stars427 forks
  42. [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.

    1,941 stars94 forks
  43. Large-scale LLM inference engine

    1,808 stars206 forks
  44. 基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.

    1,804 stars200 forks
  45. Code base of the BEVDet series .

    1,788 stars309 forks
  46. 基于 manga-image-translator 的开源漫画翻译工具。支持日/韩/美漫自动翻译,内置 OpenAI、Gemini 等 5 种翻译引擎,并提供可视化编辑器自由调整文本样式。一键安装,开箱即用。如果喜欢,欢迎点亮 ⭐ Star 支持!

    1,748 stars107 forks
  47. 【ICML 2025 Spotlight】 Official Repo for Paper ‘’HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation‘’

    1,624 stars240 forks
  48. AI for GNU Image Manipulation Program

    1,553 stars137 forks
  49. MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

    1,550 stars257 forks
  50. [ICLR'23 Spotlight & ECCV'24 & IJCV'24] MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction

    1,512 stars249 forks
  51. All-in-one training for vision models (YOLO, ViTs, RT-DETR, DINOv3): pretraining, fine-tuning, distillation.

    1,476 stars91 forks
  52. OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

    1,360 stars134 forks
  53. Awesome Person Re-identification

    1,346 stars199 forks
  54. A lightweight adapter bridges SAM with medical imaging [MedIA]

    1,313 stars126 forks
  55. Paint by Example: Exemplar-based Image Editing with Diffusion Models

    1,247 stars113 forks
  56. A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.

    1,237 stars128 forks
  57. MedRAX: Medical Reasoning Agent for Chest X-ray - ICML 2025

    1,176 stars202 forks
  58. Datasets for deep learning with satellite & aerial imagery

    1,153 stars127 forks
  59. a self-hosted webui for 30+ generative ai

    1,134 stars133 forks
  60. Dexbotic: Open-Source Vision-Language-Action Toolbox

    1,099 stars174 forks
  61. UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image and UAV image segmentation.

    1,080 stars152 forks
  62. A comprehensive benchmark of deepfake detection

    1,050 stars184 forks
  63. Offical PyTorch implementation of "BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework"

    966 stars126 forks
  64. Parsing-free RAG supported by VLMs

    960 stars76 forks
  65. OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.

    937 stars127 forks
  66. tracking medical datasets, with a focus on medical imaging

    923 stars120 forks
  67. Generate text line images for training deep learning OCR models

    912 stars175 forks
  68. Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)

    899 stars70 forks
  69. ⚕️GenAI powered multi-agentic medical diagnostics and healthcare research assistance chatbot. 🏥 Designed for healthcare professionals, researchers and patients.

    894 stars207 forks
  70. A prize winning solution for DFDC challenge

    867 stars224 forks
  71. An Agnostic Computer Vision Framework - Pluggable to any Training Library: Fastai, Pytorch-Lightning with more to come

    867 stars148 forks
  72. A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。

    826 stars51 forks
  73. A Python toolkit for fine-tuning Geospatial Foundation Models (GFMs).

    813 stars159 forks
  74. yolov8 face detection with landmark

    805 stars103 forks
  75. Official PyTorch implementation of FB-BEV & FB-OCC - Forward-backward view transformation for vision-centric autonomous driving perception

    798 stars69 forks
  76. CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包

    792 stars115 forks
  77. Transformer OCR

    779 stars252 forks
  78. 国内首个占据栅格网络全栈课程《从BEV到Occupancy Network,算法原理与工程实践》,包含端侧部署。Surrounding Semantic Occupancy Perception Course for Autonomous Driving (docs, ppt and source code) 在线课程主页:http://111.229.117.200:8100/ (作者独立搭建)

    772 stars92 forks
  79. 🔥🔥Official Repository for Anti-UAV🔥🔥

    765 stars131 forks
  80. OCR toolbox from Davar-Lab

    760 stars152 forks
  81. LSD (LiDAR SLAM & Detection) is an open source perception architecture for autonomous vehicle/robotic

    752 stars155 forks
  82. [AAAI 2026] OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

    739 stars80 forks
  83. [ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control”

    720 stars27 forks
  84. Project AirSim is Microsoft's evolution of AirSim, an advanced simulation platform for building, training, and testing autonomous systems in high-fidelity virtual environments

    690 stars137 forks
  85. YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework(Supports RGBT detection for all YOLO series from YOLOv3 to YOLOv13, as well as RTDETR. 【Ultralytics YOLOv3-YOLOv13】

    687 stars95 forks
  86. HybridNets: End-to-End Perception Network

    678 stars134 forks
  87. YOLOPv2: Better, Faster, Stronger for Panoptic driving Perception

    672 stars100 forks
  88. 天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取

    657 stars113 forks
  89. This is the pytorch implement of our paper "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model"

    657 stars44 forks
  90. 🦀 Low-level 3D Computer Vision library in Rust

    655 stars190 forks
  91. Towards deepfake detection that actually works

    650 stars199 forks
  92. ZeroCostDL4Mic: A Google Colab based no-cost toolbox to explore Deep-Learning in Microscopy

    648 stars143 forks
  93. [CoRL 2022] InterFuser: Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer

    647 stars59 forks
  94. Wining solution and its improvement for MICCAI 2017 Robotic Instrument Segmentation Sub-Challenge

    639 stars212 forks
  95. A 3D computer vision development toolkit based on PaddlePaddle. It supports point-cloud object detection, segmentation, and monocular 3D object detection models.

    638 stars149 forks
  96. Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]

    630 stars37 forks
  97. 🔥🔥First-ever hour scale video understanding models

    624 stars42 forks
  98. An open source lane detection toolbox based on PyTorch, including SCNN, RESA, UFLD, LaneATT, CondLane, etc.

    623 stars98 forks
  99. [CVPR 2026] Official implementation of "Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation"

    605 stars70 forks
  100. [Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting

    586 stars42 forks

By vertical

GitHub adoption by field

Choose a computer vision vertical for its own static page and public examples.

Choose a vertical to see its Albumentations adoption pages.