Albumentations on GitHub in Robotics
Used in 196 public GitHub repositories
Robot perception, manipulation, grasping, navigation, and embodied systems.

Public examples
Top public GitHub repositories
Showing the top 100
🤗 LeRobot: Making AI for Robotics more accessible with end-to-end learning
24,737 stars5,154 forks🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
21,620 stars3,304 forks🐍 Geometric Computer Vision Library for Spatial AI
11,238 stars1,205 forksDORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.
3,768 stars418 forksRLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
3,587 stars608 forksStarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
2,634 stars414 forks仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
2,076 stars283 forksAI for GNU Image Manipulation Program
1,553 stars137 forksPaint by Example: Exemplar-based Image Editing with Diffusion Models
1,247 stars113 forksDexbotic: Open-Source Vision-Language-Action Toolbox
1,099 stars174 forksOfficial Codebase for "DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos" (ICML 2026)
899 stars70 forksLSD (LiDAR SLAM & Detection) is an open source perception architecture for autonomous vehicle/robotic
752 stars155 forks🦀 Low-level 3D Computer Vision library in Rust
655 stars190 forksWining solution and its improvement for MICCAI 2017 Robotic Instrument Segmentation Sub-Challenge
639 stars212 forksOfficial implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
630 stars37 forks[CVPR 2026] Official implementation of "Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation"
605 stars70 forksDreamGen: Nvidia GEAR Lab's initiative to solve the robotics data problem using world models
560 stars60 forks- 547 stars33 forks
🔥🔥🔥🔥🔥🔥Docker NVIDIA Docker2 YOLOV5 YOLOX YOLO Deepsort TensorRT ROS Deepstream Jetson Nano TX2 NX for High-performance deployment(高性能部署)
541 stars132 forks[TPAMI 2024 & CVPR 2023] PyTorch code for DGM4: Detecting and Grounding Multi-Modal Media Manipulation and beyond
514 stars44 forkslearning-Journey-AI 是一个2025年最新整理推出AI学习资料归纳,包含了基础理论知识,书籍,文章和论文以方便学习,还提供了各种实战项目和工具以方便练习。
481 stars81 forksImplementation of Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
448 stars11 forks[CVPR 2024] Official implementation of FreeDrag: Feature Dragging for Reliable Point-based Image Editing
421 stars20 forks💄 Lipstick ain't enough: Beyond Color-Matching for In-the-Wild Makeup Transfer (CVPR 2021)
418 stars63 forksInternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
415 stars27 forksIsaac Lab - Arena is a robotics simulation framework that enhances NVIDIA Isaac Lab by providing a composable, scalable system for creating diverse simulation environments and evaluating robot learning policies. The framework enables developers to rapidly prototype and test robotic tasks with various robot embodiments, objects, and environments.
411 stars74 forkscode for Image Manipulation Detection by Multi-View Multi-Scale Supervision
326 stars54 forksRobotics software featuring legged locomotion algorithms and a momentum-based controller core with optimization. Supporting software for world-class robots including humanoids, running birds, exoskeletons, mechs and more.
315 stars110 forksOfficial repository of paper “IML-ViT: Benchmarking Image manipulation localization by Vision Transformer”
309 stars38 forksThis is the official code repo for DiT4DiT, a Vision-Action-Model (VAM) framework that combines video generation model with flow-matching-based action prediction for generalizable robotic manipulation.
308 stars30 forks#31
NVlabs/sage
Official Code Release of SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
307 stars28 forks[NeurIPS'24 Spotlight] A comprehensive benchmark & codebase for Image manipulation detection/localization.
298 stars50 forksAutoNode: A Neuro-Graphic Self-Learnable Engine for Cognitive GUI Automation
298 stars34 forksA personal research and development (R&D) lab that facilitates the sharing of knowledge.
297 stars51 forks[ECCV 2024] ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
279 stars21 forksVideo Face Manipulation Detection Through Ensemble of CNNs
278 stars102 forks本项目旨在为致力于进入VLA(Vision-Language-Action)领域的算法工程师提供一份全中文、实战导向的学习/面试手册。 不同于通用的 CV/NLP 面试指南,本项目聚焦于 Robotics 特有的挑战
255 stars30 forks[CVPR 2025] The offical Implementation of "Universal Actions for Enhanced Embodied Foundation Models"
241 stars11 forks[CVPR 2024] Official code for "Text-Driven Image Editing via Learnable Regions"
227 stars24 forksHEX is a whole-body vision-language-action framework for full-sized humanoid robots.
208 stars16 forks[ECCV 2026] Towards Generalizable Robotic Manipulation in Dynamic Environments
198 stars9 forksThe Comprehensive Toolkit for Embodied AI Models
191 stars42 forks#43
NVlabs/DREAM
DREAM: Deep Robot-to-Camera Extrinsics for Articulated Manipulators (ICRA 2020)
187 stars39 forks机械臂定位抓取
183 stars20 forksStable Diffusion-based image manipulation method with a sketch and reference image
183 stars10 forksGitHub repo for Jetson AI Lab
180 stars55 forksRobotics Knowledge Base. The Wiki for Robot Builders.
178 stars174 forksAn All-in-one robot manipulation learning suite for policy models training and evaluation on various datasets and benchmarks.
173 stars10 forks[ICRA 2026] Re3Sim: Generating High-Fidelity Simulation Data via 3D-Photorealistic Real-to-Sim for Robotic Manipulation
158 stars8 forksOfficial implementation for BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
154 stars13 forksIndustry leading face manipulation platform
148 stars21 forksMemory-Dependent Manipulation Benchmark based on RoboTwin
132 stars14 forksOfficial Repository for DenseMatcher Learning 3D Semantic Correspondence for Category-Level Manipulation from One Demo
122 stars8 forks#54
qcf-568/MIML
[CVPR2024] Towards Modern Image Manipulation Localization: A Large-Scale Dataset and Novel Methods
118 stars4 forks[CoRL2024] ThinkGrasp: A Vision-Language System for Strategic Part Grasping in Clutter. https://arxiv.org/abs/2407.11298
116 stars10 forksDaily updated resources on AI across various domains including ML, development, education, healthcare, real estate, robotics, crypto, web3 and more, curated by enthusiasts.
112 stars23 forksCurated collections of sample applications designed to help you develop optimized AI solutions. Tailored to specific use cases, covering retail, manufacturing, metro, and media & entertainment.
110 stars147 forks[ICLR’26] Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
107 stars2 forks[CVPR2024 Oral] Learning to Produce Semi-dense Correspondences for Visual Localization
104 stars5 forksCode and trained models for our paper: K. Triaridis, V. Mezaris, "Exploring Multi-Modal Fusion for Image Manipulation Detection and Localization", Proc. 30th Int. Conf. on MultiMedia Modeling (MMM 2024), Amsterdam, NL, Jan.-Feb. 2024.
103 stars11 forksOfficial repository for the AAAI2025 paper (Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization through Spare-Coding Transformer)
102 stars8 forks(国服)部落冲突自动化脚本
97 stars13 forks[NeurIPS 2023] Official Code for CycleNet: Rethinking Cycle Consistent in Text‑Guided Diffusion for Image Manipulation
96 stars9 forks[CVPR 25] G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
96 stars6 forksSurfEmb (CVPR 2022)
91 stars19 forks[IROS 2024] [ICML 2024 Workshop Differentiable Almost Everything] MonoForce: Learnable Image-conditioned Physics Engine
91 stars8 forks[CVPR 2026] Official code of "EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding"
91 stars3 forksSemantic Segmentation of Images and Point Clouds for Traversability Estimation
87 stars13 forks[CoRL 2025] UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
86 stars3 forks[NeurIPS 2023] SiT Dataset: Socially Interactive Pedestrian Trajectory Dataset for Social Navigation Robots
82 stars7 forksCode for "ACG: Action Coherence Guidance for Flow-based Vision-Language-Action Models" (ICRA 2026)
80 stars9 forks[NeurIPS 2023] Fine-Grained Cross-View Geo-Localization Using a Correlation-Aware Homography Estimator
76 stars7 forksRepresentation Learning and Representation Fusion for computer vision, semantic scene understanding, and robotics.
73 stars10 forksVersion of the carmen robot framework developed by LCAD for IARA - Intelligent Autonomous Robotic Autonomobile.
70 stars32 forks[arxiv 2025] TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
69 stars1 forksAn Open-Source World Model for Action-Conditioned Embodied Intelligence.
57 stars2 forkscode for CoRL2025 "LaDiWM: A Latent Diffusion-based World Model for Predictive Manipulation"
56 stars6 forksSPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation
55 stars8 forksScaling up Robot Learning in Simulation with Generative Models
55 stars1 forksThis is a lightweight GAN developed for real-time deblurring. The model has a super tiny size and a rapid inference time. The motivation is to boost marker detection in robotic applications, however, you may use it for other applications definitely.
54 stars11 forksOfficial Implementation for “CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World” (RSS 2025).
52 stars1 forks:robot: PyTorch toolkit for biomedical imaging :heart:
51 stars3 forksCodes of paper "GraspSAM: When Segment Anything Model meets Grasp Detection", ICRA 2025
49 stars9 forks[ICRA'25] DoorBot: Closed-Loop Task Planning and Manipulation for Door Opening in the Wild with Haptic Feedback
49 stars6 forks[CVPR 2025] OmniGuard: Hybrid Manipulation Localization via Augmented Versatile Deep Image Watermarking
49 stars3 forksLiDAR Image Pretraining for Visual Place Recognition
48 stars2 forks[CoRL 2024] ClutterGen: A Cluttered Scene Generator for Robot Learning
48 stars1 forks#88
DrLuo/RTM
The official repository of Real Text Manipulation (RTM)
45 stars5 forksRobust humanoid motion control via history-conditioned reinforcement learning and online distillation.
45 stars4 forksBehAV: Behavioral Rule Guided Autonomy Using VLM for Robot Navigation in Outdoor Scenes (ICRA'25)
44 stars9 forksA frontier exploration module implementied with ROS 2, C++, and Python.
44 stars8 forks[CVPR 2025 highlight] Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
44 stars4 forksICCV 2023 - Neural Collage Transfer: Artistic Reconstruction via Material Manipulation
44 stars4 forks[ICLR 2026] [NeurIPS 2025] ViPRA: Video Prediction for Robot Actions
44 stars1 forks[IROS 2025] ManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World Model
43 stars5 forksPIPE Planner: Pathwise Information Gain with Map Predictions for Indoor Robot Exploration
43 stars2 forksCode repository for paper Instance Segmentation for Autonomous Log Grasping in Forestry Operations
40 stars5 forks[CoRL 2025] Pretraining code for FLOWER VLA on OXE
40 stars4 forksSwap faces with AI from a source image to a destination medium. Img2Img, Img2GIF, & Img2MP4
37 stars12 forks[JFR 2023] - Whole-Body Motion Planning and Tracking of a Mobile Robot with a Gimbal RGB-D Camera for Outdoor 3D Exploration
37 stars4 forks