PexelsUsed in 40,984 public GitHub repositories
Open on GitHubPublic GitHub repositories using Albumentations, ranked by stars.
PexelsUsed in 40,984 public GitHub repositories
Open on GitHubPublic examples
Showing the top 100
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 160,000+ scientists worldwide. 148 ready-to-use skills plus 100+ scientific databases covering biology, chemistry, medicine, and drug discovery. Compatible with Cursor, Claude Code, Codex, Pi, Antigravity, and the open Agent Skills standard.
State-of-the-art 2D and 3D Face Analysis Project
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
A Lightweight Face Recognition and Facial Attribute Analysis (Age, Gender, Emotion and Race) Library for Python
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
OCR, layout analysis, reading order, table recognition in 90+ languages
pix2tex: Using a ViT to convert images of equations into LaTeX code.
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.
🐍 Geometric Computer Vision Library for Spatial AI
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-art AI models and flexible multi-format export.
Translate manga/image 一键翻译各类图片内文字 https://cotrans.touhou.ai/ (no longer working)
OpenMMLab Semantic Segmentation Toolbox and Benchmark.
Easy-to-use image segmentation library with awesome pre-trained model zoo, supporting wide-range of practical tasks in Semantic Segmentation, Interactive Segmentation, Panoptic Segmentation, Image Matting, 3D Segmentation, etc.
The official repo for “Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting”, ACL, 2025.
AI Toolkit for Healthcare Imaging
Unified framework for robot learning with multi-physics/renderer support
Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
Democratizing Deep-Learning for Drug Discovery, Quantum Chemistry, Materials Science and Biology
OpenMMLab's next-generation platform for general 3D object detection.
All-in-One Development Tool based on PaddlePaddle
An anomaly detection library comprising state-of-the-art algorithms and features such as experiment management, hyper-parameter optimization, and edge inference.
PDF craft can convert PDF files into various other formats. This project will focus on processing PDF files of scanned books.
#24
RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
OpenMMLab Text Detection, Recognition and Understanding Toolbox
TorchGeo: datasets, samplers, transforms, and pre-trained models for geospatial data
#27
DORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.
仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
CnOCR: Awesome Chinese/English OCR Python toolkits based on PyTorch. It comes with 20+ well-trained models for different application scenarios and can be used directly after installation. 【基于 PyTorch/MXNet 的中文/英文 OCR Python 包。】
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
Open source hardware and software platform to build a small scale self driving car.
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
GeoAI: Artificial Intelligence for Geospatial Data
An Open-Source Python3 tool with SMALL models for recognizing layouts, tables, math formulas (LaTeX), and text in images, converting them into Markdown format. A free alternative to Mathpix, empowering seamless conversion of visual content into text-based representations. 80+ languages are supported.
The largest open-source medical AI skills library for OpenClaw🦞.
Dexbotic: Open-Source Vision-Language-Action Toolbox
基于manga-image-translator 实现的开源漫画AI翻译桌面工具。支持日、韩、英文漫画自动处理,集成OpenAl、Gemini等多翻译引擎;实现OCR文字检测、原文擦除、AI翻译、图像修复、译文排版完整链路,自带可视化编辑器,支持自定义文本样式,一键部署开箱即用。
Optical character recognition for Japanese text, with the main focus being Japanese manga
Jupyter Notebook tutorials on solving real-world problems with Machine Learning & Deep Learning using PyTorch. Topics: Face detection with Detectron 2, Time Series anomaly detection with LSTM Autoencoders, Object Detection with YOLO v5, Build your first Neural Network, Time Series forecasting for Coronavirus daily cases, Sentiment Analysis with BER
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
An open source library and framework for deep learning on satellite and aerial imagery.
An Open-World Foundation Model for General-Purpose Embodied Intelligence.
yolov5 + csl_label.(Oriented Object Detection)(Rotation Detection)(Rotated BBox)基于yolov5的旋转目标检测
基于PaddleOCR重构,并且脱离PaddlePaddle深度学习训练框架的轻量级OCR,推理速度超快 —— A lightweight OCR system based on PaddleOCR, decoupled from the PaddlePaddle deep learning training framework, with ultra-fast inference speed.
#45
Large-scale LLM inference engine
Code base of the BEVDet series .
【ICML 2025 Spotlight】 Official Repo for Paper ‘’HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation‘’
AI for GNU Image Manipulation Program
#49
[ICLR'23 Spotlight & ECCV'24 & IJCV'24] MapTR: Structured Modeling and Learning for Online Vectorized HD Map Construction
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.
Awesome Person Re-identification
A lightweight adapter bridges SAM with medical imaging [MedIA]
Paint by Example: Exemplar-based Image Editing with Diffusion Models
MedRAX: Medical Reasoning Agent for Chest X-ray - ICML 2025
A central hub for gathering and showcasing amazing projects that extend OpenMMLab with SAM and other exciting features.
Datasets for deep learning with satellite & aerial imagery
Everything about ComfyUI, including workflow sharing, resource sharing, knowledge sharing, tutorial sharing, and more.关于ComfyUI的一切,工作流分享、资源分享、知识分享、教程分享等
a self-hosted webui for 30+ generative ai
A comprehensive benchmark of deepfake detection
UNetFormer: A UNet-like transformer for efficient semantic segmentation of remote sensing urban scene imagery, ISPRS. Also, including other vision transformers and CNNs for satellite, aerial image and UAV image segmentation.
Official Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
TurboOCR, >200 img/s OmnidocBench. TensorRT FP16, PP-OCRv6, HTTP + gRPC
PyTorch implementation of UNet++ (Nested U-Net).
Offical PyTorch implementation of "BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework"
⚕️GenAI powered multi-agentic medical diagnostics and healthcare research assistance chatbot. 🏥 Designed for healthcare professionals, researchers and patients.
Parsing-free RAG supported by VLMs
OpenEMMA, a permissively licensed open source "reproduction" of Waymo’s EMMA model.
tracking medical datasets, with a focus on medical imaging
A curated, multilingual library of 182 installable AI agent skills for end-to-end academic research—spanning literature discovery, scientific writing, grant development, bioinformatics, drug discovery, clinical research, machine learning, and data analysis.
An Open-Source World Model for Action-Conditioned Embodied Intelligence.
Generate text line images for training deep learning OCR models
#72
A minimalist SOTA LaTeX OCR model with only 20M parameters, running in browser. Full training pipeline available for self-reproduction. | 超轻量SOTA LaTeX公式识别模型,仅20M参数量,可在浏览器中运行。训练全流程代码开源,以便自学复现。
Project AirSim is Microsoft's evolution of AirSim, an advanced simulation platform for building, training, and testing autonomous systems in high-fidelity virtual environments
Neuroimaging analysis and visualization suite
An Agnostic Computer Vision Framework - Pluggable to any Training Library: Fastai, Pytorch-Lightning with more to come
A prize winning solution for DFDC challenge
A Python toolkit for fine-tuning Geospatial Foundation Models (GFMs).
🔥🔥Official Repository for Anti-UAV🔥🔥
yolov8 face detection with landmark
天枢 - 企业级 AI 一站式数据预处理平台 | PDF/Office转Markdown | 支持MCP协议AI助手集成 | Vue3+FastAPI全栈方案 | 文档解析 | 多模态信息提取
[AAAI 2026] OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
Transformer OCR
国内首个占据栅格网络全栈课程《从BEV到Occupancy Network,算法原理与工程实践》,包含端侧部署。Surrounding Semantic Occupancy Perception Course for Autonomous Driving (docs, ppt and source code) 在线课程主页:http://111.229.117.200:8100/ (作者独立搭建)
Official PyTorch implementation of FB-BEV & FB-OCC - Forward-backward view transformation for vision-centric autonomous driving perception
CnSTD: 基于 PyTorch/MXNet 的 中文/英文 场景文字检测(Scene Text Detection)、数学公式检测(Mathematical Formula Detection, MFD)、篇章分析(Layout Analysis)的Python3 包
LSD (LiDAR SLAM & Detection) is an open source perception architecture for autonomous vehicle/robotic
Lightweight & fast OCR models for license plate text recognition.
OCR toolbox from Davar-Lab
Text-to-3D Generation within 5 Minutes
[ICCV 2025] Official implementation of the paper “MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control”
YOLOv11-RGBT: Towards a Comprehensive Single-Stage Multispectral Object Detection Framework(Supports RGBT detection for all YOLO series from YOLOv3 to YOLOv13, as well as RTDETR. 【Ultralytics YOLOv3-YOLOv13】
An all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.
🦀 Low-level 3D Computer Vision library in Rust
YOLOPv2: Better, Faster, Stronger for Panoptic driving Perception
HybridNets: End-to-End Perception Network
本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战
[CVPR 2026] Official implementation of "Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation"
This is the pytorch implement of our paper "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model"
Implementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
[CoRL 2022] InterFuser: Safety-Enhanced Autonomous Driving Using Interpretable Sensor Fusion Transformer
By vertical
Choose a computer vision vertical for its own static page and public examples.
Agriculture and foodUsed in 101 public GitHub repositoriesView adoption Pexels
Autonomous mobilityUsed in 227 public GitHub repositoriesView adoption Pexels
Documents and OCRUsed in 495 public GitHub repositoriesView adoption Pexels
Drones and UAVsUsed in 92 public GitHub repositoriesView adoption Pexels
Energy and utilitiesUsed in 21 public GitHub repositoriesView adoption Pexels
Environment and conservationUsed in 44 public GitHub repositoriesView adoption Pexels
Geospatial and Earth observationUsed in 422 public GitHub repositoriesView adoption Pexels
Industrial and manufacturingUsed in 84 public GitHub repositoriesView adoption Pexels
InfrastructureUsed in 17 public GitHub repositoriesView adoption Pexels
Life sciencesUsed in 784 public GitHub repositoriesView adoption Pexels
MaritimeUsed in 83 public GitHub repositoriesView adoption Pexels
Medical imagingUsed in 726 public GitHub repositoriesView adoption Pexels
RoboticsUsed in 211 public GitHub repositoriesView adoption Pexels
Security and biometricsUsed in 304 public GitHub repositoriesView adoption PexelsChoose a vertical to see its Albumentations adoption pages.