Albumentations on GitHub in Robotics
Used in 211 public GitHub repositories
Robot perception, manipulation, grasping, navigation, and embodied systems.

Public examples
Top public GitHub repositories
Showing the top 100
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
27,884 stars5,154 forks🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
22,024 stars3,304 forks🐍 Geometric Computer Vision Library for Spatial AI
11,391 stars1,205 forksUnified framework for robot learning with multi-physics/renderer support
8,266 stars3,932 forksRLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
5,419 stars608 forksDORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.
3,991 stars448 forks仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
3,902 stars283 forksStarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
3,750 stars494 forksAgent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
3,484 stars427 forksDexbotic: Open-Source Vision-Language-Action Toolbox
2,976 stars280 forksAn Open-World Foundation Model for General-Purpose Embodied Intelligence.
2,176 stars184 forksAI for GNU Image Manipulation Program
1,556 stars137 forksPaint by Example: Exemplar-based Image Editing with Diffusion Models
1,252 stars112 forksOfficial Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
1,106 stars82 forksAn Open-Source World Model for Action-Conditioned Embodied Intelligence.
926 stars2 forksLSD (LiDAR SLAM & Detection) is an open source perception architecture for autonomous vehicle/robotic
773 stars158 forksText-to-3D Generation within 5 Minutes
737 stars55 forksAn all-in-one VLA engineering platform for embodied AI — from data to real-robot deployment.
720 stars73 forks🦀 Low-level 3D Computer Vision library in Rust
718 stars203 forks本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战
671 stars30 forks[CVPR 2026] Official implementation of "Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation"
668 stars70 forksImplementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
665 stars14 forksWining solution and its improvement for MICCAI 2017 Robotic Instrument Segmentation Sub-Challenge
638 stars212 forksOfficial implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
631 stars37 forksThis repo is the official implementation of "τ0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation".
628 stars22 forksDreamGen: Nvidia GEAR Lab's initiative to solve the robotics data problem using world models
616 stars63 forksIsaac Lab - Arena is a robotics simulation framework that enhances NVIDIA Isaac Lab by providing a composable, scalable system for creating diverse simulation environments and evaluating robot learning policies. The framework enables developers to rapidly prototype and test robotic tasks with various robot embodiments, objects, and environments.
592 stars74 forks🔥🔥🔥🔥🔥🔥Docker NVIDIA Docker2 YOLOV5 YOLOX YOLO Deepsort TensorRT ROS Deepstream Jetson Nano TX2 NX for High-performance deployment(高性能部署)
543 stars133 forkslearning-Journey-AI 是一个2025年最新整理推出AI学习资料归纳,包含了基础理论知识,书籍,文章和论文以方便学习,还提供了各种实战项目和工具以方便练习。
525 stars81 forks[TPAMI 2024 & CVPR 2023] PyTorch code for DGM4: Detecting and Grounding Multi-Modal Media Manipulation and beyond
520 stars44 forksThis is the official code repo for DiT4DiT, a Vision-Action-Model (VAM) framework that combines video generation model with flow-matching-based action prediction for generalizable robotic manipulation.
461 stars36 forksInternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
432 stars27 forks💄 Lipstick ain't enough: Beyond Color-Matching for In-the-Wild Makeup Transfer (CVPR 2021)
420 stars63 forks[CVPR 2024] Official implementation of FreeDrag: Feature Dragging for Reliable Point-based Image Editing
420 stars20 forks#35
NVlabs/sage
Official Code Release of SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
417 stars32 forks[IROS 26 ScaleInfra Workshop Best Tool Paper] Involving over 40 Advanced Manipulation Policies
387 stars98 forksHEX is a whole-body vision-language-action framework for full-sized humanoid robots.
335 stars18 forkscode for Image Manipulation Detection by Multi-View Multi-Scale Supervision
334 stars53 forks[NeurIPS'24 Spotlight] A comprehensive benchmark & codebase for Image manipulation detection/localization.
319 stars54 forksOfficial repository of paper “IML-ViT: Benchmarking Image manipulation localization by Vision Transformer”
310 stars38 forksA personal research and development (R&D) lab that facilitates the sharing of knowledge.
299 stars51 forksAutoNode: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
299 stars34 forks[ECCV 2024] ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
283 stars21 forksVideo Face Manipulation Detection Through Ensemble of CNNs
279 stars102 forksThe Comprehensive Toolkit for Embodied AI Models
251 stars46 forks[CVPR 2025] The offical Implementation of "Universal Actions for Enhanced Embodied Foundation Models"
245 stars12 forks[ECCV 2026] Towards Generalizable Robotic Manipulation in Dynamic Environments
234 stars13 forks[CVPR 2024] Official code for "Text-Driven Image Editing via Learnable Regions"
226 stars24 forksA unified inference runtime for VLA models.
213 stars36 forks机械臂定位抓取
210 stars23 forksRobotics Knowledge Base. The Wiki for Robot Builders.
190 stars174 forks#52
NVlabs/DREAM
DREAM: Deep Robot-to-Camera Extrinsics for Articulated Manipulators (ICRA 2020)
190 stars40 forksStable Diffusion-based image manipulation method with a sketch and reference image
183 stars10 forks[ICRA 2026] Re3Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
179 stars10 forksAn All-in-one robot manipulation learning suite for policy models training and evaluation on various datasets and benchmarks.
175 stars10 forksOfficial implementation for BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
167 stars13 forksHigh-Performance Serving Runtime for Robotics Models
157 stars6 forksIndustry leading face manipulation platform
155 stars24 forksA curated collection of sample applications intended for reference in developing optimized AI solutions and testing hardware performance across various industry use cases.
137 stars169 forks[NeurIPS 2026] Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
135 stars7 forksOfficial Repository for DenseMatcher Learning 3D Semantic Correspondence for Category-Level Manipulation from One Demo
124 stars8 forks[CVPR 2026] Official code of "EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding"
124 stars3 forks[CoRL2024] ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter. https://arxiv.org/abs/2407.11298
121 stars10 forksDaily updated resources on AI across various domains including ML, development, education, healthcare, real estate, robotics, crypto, web3 and more, curated by enthusiasts.
119 stars23 forksOfficial Implementation of "WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN"
116 stars0 forks#66
qcf-568/MIML
[CVPR2024] Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel Methods
113 stars4 forksOfficial repository for the AAAI2025 paper (Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization through Spare-Coding Transformer)
111 stars8 forks- 111 stars1 forks
(国服)部落冲突自动化脚本
107 stars13 forks[ICLR’26] Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
107 stars1 forks[CVPR2024 Oral] Learning to Produce Semi-dense Correspondences for Visual Localization
104 stars5 forksA unified VLA post-training framework for human-in-the-loop data collection, fine-tuning, and reinforcement learning.
103 stars20 forksCode and trained models for our paper: K. Triaridis, V. Mezaris, "Exploring Multi-Modal Fusion for Image Manipulation Detection and Localization", Proc. 30th Int. Conf. on MultiMedia Modeling (MMM 2024), Amsterdam, NL, Jan.-Feb. 2024.
103 stars11 forksModular DRL framework for autonomous robot navigation in ROS2. Plug-and-play RL backends (Stable-Baselines3, DreamerV3), composable reward functions, observation spaces & neural architectures - built for research and deployment.
103 stars8 forksRepository associated with paper titled "CLP: Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think"
98 stars1 forks[NeurIPS 2023] Official Code for CycleNet: Rethinking Cycle Consistent in Text‑Guided Diffusion for Image Manipulation
96 stars9 forks[CVPR 25] G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
96 stars6 forks#78
scu-zjz/RITA
[CVPR 26 Findings] Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
95 stars19 forksA world-model-enhanced VLA for steerable dexterous manipulation, along with optimized training infra.
94 stars5 forksSurfEmb (CVPR 2022)
92 stars19 forks[IROS 2024] [ICML 2024 Workshop Differentiable Almost Everything] MonoForce: Learnable Image-conditioned Physics Engine
92 stars8 forks[CoRL 2025] UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
92 stars3 forksCode for "ACG: Action Coherence Guidance for Flow-based Vision-Language-Action Models" (ICRA 2026)
88 stars9 forksSemantic Segmentation of Images and Point Clouds for Traversability Estimation
87 stars13 forks[arxiv 2025] TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
86 stars2 forksPython-native stack for real-life ML robotics
84 stars12 forks[NeurIPS 2023] SiT Dataset: Socially Interactive Pedestrian Trajectory Dataset for Social Navigation Robots
84 stars7 forks[NeurIPS 2023] Fine-Grained Cross-View Geo-Localization Using a Correlation-Aware Homography Estimator
79 stars9 forksSLAM-Supported Semi-Supervised Learning for 6D Object Pose Estimation
77 stars6 forksRepresentation Learning and Representation Fusion for computer vision, semantic scene understanding, and robotics.
74 stars10 forksVersion of the carmen robot framework developed by LCAD for IARA - Intelligent Autonomous Robotic Autonomobile.
70 stars32 forksSPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation
65 stars8 forksA Bimanual Manipulation Benchmark across Dexterous Hands with Simulated Visuo-Tactile Support
65 stars7 forkscode for CoRL2025 "LaDiWM: A Latent Diffusion-based World Model for Predictive Manipulation"
62 stars7 forksAn autonomy stack for the Unitree G1 humanoid, built on ROS 2.
62 stars6 forksOfficial implementation of Chronos: A Physics-Informed Full-History Framework for Non-Markovian Long-Horizon Manipulation.
58 stars3 forks[ICRA'25] DoorBot: Closed-Loop Task Planning and Manipulation for Door Opening in the Wild with Haptic Feedback
57 stars6 forksThis is a lightweight GAN developed for real-time deblurring. The model has a super tiny size and a rapid inference time. The motivation is to boost marker detection in robotic applications, however, you may use it for other applications definitely.
56 stars11 forksScaling up Robot Learning in Simulation with Generative Models
56 stars1 forksOfficial Implementation for “CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World” (RSS 2025).
55 stars1 forks