PexelsUsed in 40,684 public GitHub repositories
Open on GitHubPublic GitHub repositories using Albumentations, ranked by stars.
PexelsUsed in 40,684 public GitHub repositories
Open on GitHubPublic examples
Showing the top 100
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
State-of-the-art 2D and 3D Face Analysis Project
The fastai deep learning library
Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 160,000+ scientists worldwide. 148 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
pix2tex: Using a ViT to convert images of equations into LaTeX code.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
🐍 Geometric Computer Vision Library for Spatial AI
Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)
OpenMMLab Semantic Segmentation Toolbox and Benchmark.
Easy-to-use image segmentation library with awesome pre-trained model zoo, supporting wide-range of practical tasks in Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, 3D Segmentation, etc.
Effortless data labeling with AI support from Segment Anything and other awesome models.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
AI Toolkit for Healthcare Imaging
Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Democratizing Deep-Learning for Drug Discovery, Quantum Chemistry, Materials Science and Biology
OpenMMLab's next-generation platform for general 3D object detection.
All-in-One Development Tool based on PaddlePaddle
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference.
OpenMMLab Text Detection, Recognition and Understanding Toolbox
TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data
A comprehensive list of Deep Learning / Artificial Intelligence and Machine Learning tutorials - rapidly expanding into areas of AI/Deep Learning / Machine Vision / NLP and industry specific areas such as Climate / Energy, Automotives, Retail, Pharma, Medicine, Healthcare, Policy, Ethics and more.
#26
DORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.
CnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】
#28
RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
🛠 All-in-one web-based IDE specialized for machine learning and data science.
Open source hardware and software platform to build a small scale self driving car.
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
GeoAI: Artificial Intelligence for Geospatial Data
Optical character recognition for Japanese text, with the main focus being Japanese manga
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
The largest open-source medical AI skills library for OpenClaw🦞.
Jupyter Notebook tutorials on solving real-world problems with Machine Learning & Deep Learning using PyTorch. Topics: Face detection with Detectron 2, Time Series anomaly detection with LSTM Autoencoders, Object Detection with YOLO v5, Build your first Neural Network, Time Series forecasting for Coronavirus daily cases, Sentiment Analysis with BER
An open source library and framework for deep learning on satellite and aerial imagery.
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
yolov5 + csl_label.(Oriented Object Detection)(Rotation Detection)(Rotated BBox)基于yolov5的旋转目标检测
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
#43
Large-scale LLM inference engine
基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.
Code base of the BEVDet series .
基于 manga-image-translator 的开源漫画翻译工具。支持日/韩/美漫自动翻译,内置 OpenAI、Gemini 等 5 种翻译引擎,并提供可视化编辑器自由调整文本样式。一键安装,开箱即用。如果喜欢,欢迎点亮 ⭐ Star 支持!
【ICML 2025 Spotlight】 Official Repo for Paper ‘’HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation‘’
AI for GNU Image Manipulation Program
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
#50
[ICLR'23 Spotlight & ECCV'24 & IJCV'24] MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
All-in-one training for vision models (YOLO, ViTs, RT-DETR, DINOv3): pretraining, fine-tuning, distillation.
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
Awesome Person Re-identification
A lightweight adapter bridges SAM with medical imaging [MedIA]
Paint by Example: Exemplar-based Image Editing with Diffusion Models
A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.
MedRAX: Medical Reasoning Agent for Chest X-ray - ICML 2025
Datasets for deep learning with satellite & aerial imagery
a self-hosted webui for 30+ generative ai
Dexbotic: Open-Source Vision-Language-Action Toolbox
UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image and UAV image segmentation.
A comprehensive benchmark of deepfake detection
Offical PyTorch implementation of "BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework"
Parsing-free RAG supported by VLMs
OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
tracking medical datasets, with a focus on medical imaging
Generate text line images for training deep learning OCR models
Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
⚕️GenAI powered multi-agentic medical diagnostics and healthcare research assistance chatbot. 🏥 Designed for healthcare professionals, researchers and patients.
A prize winning solution for DFDC challenge
An Agnostic Computer Vision Framework - Pluggable to any Training Library: Fastai, Pytorch-Lightning with more to come
#72
A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。
A Python toolkit for fine-tuning Geospatial Foundation Models (GFMs).
yolov8 face detection with landmark
Official PyTorch implementation of FB-BEV & FB-OCC - Forward-backward view transformation for vision-centric autonomous driving perception
CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包
Transformer OCR
国内首个占据栅格网络全栈课程《从BEV到Occupancy Network,算法原理与工程实践》,包含端侧部署。Surrounding Semantic Occupancy Perception Course for Autonomous Driving (docs, ppt and source code) 在线课程主页:http://111.229.117.200:8100/ (作者独立搭建)
🔥🔥Official Repository for Anti-UAV🔥🔥
OCR toolbox from Davar-Lab
LSD (LiDAR SLAM & Detection) is an open source perception architecture for autonomous vehicle/robotic
[AAAI 2026] OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
[ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control”
Project AirSim is Microsoft's evolution of AirSim, an advanced simulation platform for building, training, and testing autonomous systems in high-fidelity virtual environments
YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework(Supports RGBT detection for all YOLO series from YOLOv3 to YOLOv13, as well as RTDETR. 【Ultralytics YOLOv3-YOLOv13】
HybridNets: End-to-End Perception Network
YOLOPv2: Better, Faster, Stronger for Panoptic driving Perception
天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取
This is the pytorch implement of our paper "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model"
🦀 Low-level 3D Computer Vision library in Rust
Towards deepfake detection that actually works
ZeroCostDL4Mic: A Google Colab based no-cost toolbox to explore Deep-Learning in Microscopy
[CoRL 2022] InterFuser: Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer
Wining solution and its improvement for MICCAI 2017 Robotic Instrument Segmentation Sub-Challenge
A 3D computer vision development toolkit based on PaddlePaddle. It supports point-cloud object detection, segmentation, and monocular 3D object detection models.
Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
🔥🔥First-ever hour scale video understanding models
An open source lane detection toolbox based on PyTorch, including SCNN, RESA, UFLD, LaneATT, CondLane, etc.
[CVPR 2026] Official implementation of "Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation"
[Open-Source Project] Combining MMOCR with Segment Anything & Stable Diffusion. Automatically detect, recognize and segment text instances, with serval downstream tasks, e.g., Text Removal and Text Inpainting
By vertical
Choose a computer vision vertical for its own static page and public examples.
Agriculture and foodUsed in 104 public GitHub repositoriesView adoption Pexels
Autonomous mobilityUsed in 259 public GitHub repositoriesView adoption Pexels
Documents and OCRUsed in 549 public GitHub repositoriesView adoption Pexels
Drones and UAVsUsed in 98 public GitHub repositoriesView adoption Pexels
Energy and utilitiesUsed in 23 public GitHub repositoriesView adoption Pexels
Environment and conservationUsed in 48 public GitHub repositoriesView adoption Pexels
Geospatial and Earth observationUsed in 542 public GitHub repositoriesView adoption Pexels
Industrial and manufacturingUsed in 96 public GitHub repositoriesView adoption Pexels
InfrastructureUsed in 19 public GitHub repositoriesView adoption Pexels
Life sciencesUsed in 1,021 public GitHub repositoriesView adoption Pexels
MaritimeUsed in 96 public GitHub repositoriesView adoption Pexels
Medical imagingUsed in 959 public GitHub repositoriesView adoption Pexels
RoboticsUsed in 196 public GitHub repositoriesView adoption Pexels
Security and biometricsUsed in 313 public GitHub repositoriesView adoption PexelsChoose a vertical to see its Albumentations adoption pages.