Adoption
Public code

GitHub adoption

Public GitHub repositories using Albumentations, ranked by stars.

Public examples

Top public GitHub repositories

Showing the top 100

  1. Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

    90,495 stars11,434 forks
  2. Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 160,000+ scientists worldwide. 148 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.

    47,184 stars3,130 forks
  3. State-of-the-art 2D and 3D Face Analysis Project

    29,877 stars6,102 forks
  4. 🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning

    27,884 stars5,154 forks
  5. A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python

    23,507 stars3,188 forks
  6. 🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools

    22,024 stars3,304 forks
  7. OCR, layout analysis, reading order, table recognition in 90+ languages

    21,431 stars1,550 forks
  8. pix2tex: Using a ViT to convert images of equations into LaTeX code.

    16,581 stars1,311 forks
  9. Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

    15,181 stars1,277 forks
  10. 🐍 Geometric Computer Vision Library for Spatial AI

    11,391 stars1,205 forks
  11. X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.

    10,579 stars1,171 forks
  12. Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)

    10,464 stars1,035 forks
  13. OpenMMLab Semantic Segmentation Toolbox and Benchmark.

    9,966 stars2,857 forks
  14. Easy-to-use image segmentation library with awesome pre-trained model zoo, supporting wide-range of practical tasks in Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, 3D Segmentation, etc.

    9,399 stars1,711 forks
  15. The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.

    9,059 stars777 forks
  16. AI Toolkit for Healthcare Imaging

    8,739 stars1,567 forks
  17. Unified framework for robot learning with multi-physics/renderer support

    8,266 stars3,932 forks
  18. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

    8,227 stars716 forks
  19. Democratizing Deep-Learning for Drug Discovery, Quantum Chemistry, Materials Science and Biology

    7,025 stars2,263 forks
  20. OpenMMLab's next-generation platform for general 3D object detection.

    6,546 stars1,783 forks
  21. All-in-One Development Tool based on PaddlePaddle

    6,270 stars1,208 forks
  22. An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference.

    6,211 stars962 forks
  23. PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.

    6,015 stars417 forks
  24. RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI

    5,419 stars608 forks
  25. OpenMMLab Text Detection, Recognition and Understanding Toolbox

    4,752 stars779 forks
  26. TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data

    4,192 stars562 forks
  27. DORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.

    3,991 stars448 forks
  28. 仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能

    3,902 stars283 forks
  29. CnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】

    3,764 stars539 forks
  30. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

    3,750 stars494 forks
  31. Open source hardware and software platform to build a small scale self driving car.

    3,512 stars1,368 forks
  32. Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.

    3,484 stars427 forks
  33. GeoAI: Artificial Intelligence for Geospatial Data

    3,406 stars486 forks
  34. An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.

    3,265 stars283 forks
  35. The largest open-source medical AI skills library for OpenClaw🦞.

    3,040 stars402 forks
  36. Dexbotic: Open-Source Vision-Language-Action Toolbox

    2,976 stars280 forks
  37. 基于manga-image-translator 实现的开源漫画AI翻译桌面工具。支持日、韩、英文漫画自动处理,集成OpenAl、Gemini等多翻译引擎;实现OCR文字检测、原文擦除、AI翻译、图像修复、译文排版完整链路,自带可视化编辑器,支持自定义文本样式,一键部署开箱即用。

    2,893 stars138 forks
  38. Optical character recognition for Japanese text, with the main focus being Japanese manga

    2,793 stars142 forks
  39. Jupyter Notebook tutorials on solving real-world problems with Machine Learning & Deep Learning using PyTorch. Topics: Face detection with Detectron 2, Time Series anomaly detection with LSTM Autoencoders, Object Detection with YOLO v5, Build your first Neural Network, Time Series forecasting for Coronavirus daily cases, Sentiment Analysis with BER

    2,510 stars643 forks
  40. DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

    2,286 stars177 forks
  41. An open source library and framework for deep learning on satellite and aerial imagery.

    2,242 stars397 forks
  42. An Open-World Foundation Model for General-Purpose Embodied Intelligence.

    2,176 stars184 forks
  43. yolov5 + csl_label.(Oriented Object Detection)(Rotation Detection)(Rotated BBox)基于yolov5的旋转目标检测

    1,948 stars426 forks
  44. 基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.

    1,874 stars200 forks
  45. Large-scale LLM inference engine

    1,869 stars206 forks
  46. Code base of the BEVDet series .

    1,803 stars309 forks
  47. 【ICML 2025 Spotlight】 Official Repo for Paper ‘’HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation‘’

    1,662 stars243 forks
  48. AI for GNU Image Manipulation Program

    1,556 stars137 forks
  49. [ICLR'23 Spotlight & ECCV'24 & IJCV'24] MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction

    1,555 stars249 forks
  50. OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

    1,464 stars149 forks
  51. Awesome Person Re-identification

    1,363 stars199 forks
  52. A lightweight adapter bridges SAM with medical imaging [MedIA]

    1,326 stars127 forks
  53. Paint by Example: Exemplar-based Image Editing with Diffusion Models

    1,252 stars112 forks
  54. MedRAX: Medical Reasoning Agent for Chest X-ray - ICML 2025

    1,237 stars206 forks
  55. A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.

    1,235 stars127 forks
  56. Datasets for deep learning with satellite & aerial imagery

    1,204 stars127 forks
  57. Everything about ComfyUI, including workflow sharing, resource sharing, knowledge sharing, tutorial sharing, and more.关于ComfyUI的一切,工作流分享、资源分享、知识分享、教程分享等

    1,181 stars168 forks
  58. a self-hosted webui for 30+ generative ai

    1,151 stars135 forks
  59. A comprehensive benchmark of deepfake detection

    1,120 stars191 forks
  60. UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image and UAV image segmentation.

    1,109 stars153 forks
  61. Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)

    1,106 stars82 forks
  62. TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC

    1,098 stars113 forks
  63. PyTorch implementation of UNet++ (Nested U-Net).

    1,051 stars219 forks
  64. Offical PyTorch implementation of "BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework"

    989 stars127 forks
  65. ⚕️GenAI powered multi-agentic medical diagnostics and healthcare research assistance chatbot. 🏥 Designed for healthcare professionals, researchers and patients.

    981 stars207 forks
  66. Parsing-free RAG supported by VLMs

    981 stars76 forks
  67. OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.

    956 stars126 forks
  68. tracking medical datasets, with a focus on medical imaging

    930 stars120 forks
  69. A curated, multilingual library of 182 installable AI agent skills for end-to-end academic research—spanning literature discovery, scientific writing, grant development, bioinformatics, drug discovery, clinical research, machine learning, and data analysis.

    930 stars61 forks
  70. An Open-Source World Model for Action-Conditioned Embodied Intelligence.

    926 stars2 forks
  71. Generate text line images for training deep learning OCR models

    915 stars174 forks
  72. A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。

    903 stars53 forks
  73. Project AirSim is Microsoft's evolution of AirSim, an advanced simulation platform for building, training, and testing autonomous systems in high-fidelity virtual environments

    882 stars170 forks
  74. Neuroimaging analysis and visualization suite

    877 stars301 forks
  75. An Agnostic Computer Vision Framework - Pluggable to any Training Library: Fastai, Pytorch-Lightning with more to come

    868 stars148 forks
  76. A prize winning solution for DFDC challenge

    867 stars220 forks
  77. A Python toolkit for fine-tuning Geospatial Foundation Models (GFMs).

    864 stars163 forks
  78. 🔥🔥Official Repository for Anti-UAV🔥🔥

    859 stars142 forks
  79. yolov8 face detection with landmark

    832 stars103 forks
  80. 天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取

    830 stars126 forks
  81. [AAAI 2026] OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

    819 stars80 forks
  82. Transformer OCR

    817 stars257 forks
  83. 国内首个占据栅格网络全栈课程《从BEV到Occupancy Network,算法原理与工程实践》,包含端侧部署。Surrounding Semantic Occupancy Perception Course for Autonomous Driving (docs, ppt and source code) 在线课程主页:http://111.229.117.200:8100/ (作者独立搭建)

    797 stars92 forks
  84. Official PyTorch implementation of FB-BEV & FB-OCC - Forward-backward view transformation for vision-centric autonomous driving perception

    797 stars69 forks
  85. CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包

    794 stars114 forks
  86. LSD (LiDAR SLAM & Detection) is an open source perception architecture for autonomous vehicle/robotic

    773 stars158 forks
  87. Lightweight & fast OCR models for license plate text recognition.

    766 stars99 forks
  88. OCR toolbox from Davar-Lab

    762 stars152 forks
  89. Text-to-3D Generation within 5 Minutes

    737 stars55 forks
  90. [ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control”

    727 stars27 forks
  91. YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework(Supports RGBT detection for all YOLO series from YOLOv3 to YOLOv13, as well as RTDETR. 【Ultralytics YOLOv3-YOLOv13】

    724 stars95 forks
  92. An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.

    720 stars73 forks
  93. 🦀 Low-level 3D Computer Vision library in Rust

    718 stars203 forks
  94. YOLOPv2: Better, Faster, Stronger for Panoptic driving Perception

    692 stars99 forks
  95. HybridNets: End-to-End Perception Network

    688 stars135 forks
  96. 本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战

    671 stars30 forks
  97. [CVPR 2026] Official implementation of "Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation"

    668 stars70 forks
  98. This is the pytorch implement of our paper "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model"

    668 stars44 forks
  99. Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

    665 stars14 forks
  100. [CoRL 2022] InterFuser: Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer

    652 stars59 forks

By vertical

GitHub adoption by field

Choose a computer vision vertical for its own static page and public examples.

Choose a vertical to see its Albumentations adoption pages.