AlbumentationsX vs PIL/Pillow
Compare AlbumentationsX with PIL/Pillow for image augmentation: API differences, RGB JPEG-to-CUDA benchmark results, and migration examples.
What Is Different?
Pillow is an image toolkit. AlbumentationsX is an augmentation pipeline library. That sounds subtle, but it changes the shape of the code: Pillow gives you image operations; AlbumentationsX gives you randomized, reproducible transforms that keep images, masks, bounding boxes, and keypoints synchronized.
- Pillow works with PIL Image objects; AlbumentationsX works with NumPy arrays and returns a dictionary of augmented targets.
- Pillow is great for loading, saving, drawing, and simple image edits; AlbumentationsX is built for training-time augmentation pipelines.
- AlbumentationsX has first-class random composition, probabilities, target synchronization, and bbox/keypoint parameter handling.
- Pillow code usually becomes manual orchestration when masks or labels must follow the image; AlbumentationsX keeps that in the pipeline contract.
RGB input pipeline results
Every path reads RGB JPEGs, prepares the recipe, and delivers a synchronized CUDA batch. CPU and GPU labels identify where augmentation runs; normalization runs on GPU for every path.
| Measured path | Throughput / AX | GPU memory (MiB) |
|---|---|---|
| AlbumentationsX CPU | 1.00× | 1,852 |
| Pillow CPU | 0.71× | 1,814 |
Each comparison uses its own shared recipe set. Averages from different sets cannot rank all libraries. The table below includes every measured recipe for these paths, including recipes outside the summary set. Recipe names are shortened; hover over a name for its full pipeline.
Higher throughput is better. Values are medians across seeds. Hover for the observed range. A dash means no measured result.
| R01Resize224 | 4,751 | 2,390 |
|---|---|---|
| R02RandomCrop224 | 4,740 | 3,768 |
| R03RandomResizedCrop | 4,785 | 2,721 |
| R04HorizontalFlip | 4,723 | 3,759 |
| R05VerticalFlip | 4,907 | 3,567 |
| R06Pad+RandomCrop224 | 4,397 | 3,659 |
| R07Rotate | 3,352 | 3,555 |
| R08Affine | 3,049 | 2,647 |
| R09Perspective | 2,879 | — |
| R10Elastic | 1,990 | — |
| R11ColorJitter | 3,523 | — |
| R12ChannelShuffle | 5,026 | — |
| R13Grayscale | 5,157 | 3,687 |
| R14RGBShift | 4,348 | — |
| R15GaussianBlur | 4,679 | 2,393 |
| R16GaussianNoise | 3,288 | — |
| R17Invert | 5,076 | 3,540 |
| R18Posterize | 5,110 | 3,493 |
| R19Solarize | 4,607 | 3,506 |
| R20Sharpen | 4,223 | — |
| R21AutoContrast | 4,263 | 3,513 |
| R22Equalize | 3,986 | 3,245 |
| R23Erasing | 4,938 | — |
| R24JpegCompression | 4,232 | 3,183 |
| R25RandomGamma | 4,969 | — |
| R26PlankianJitter | 4,538 | — |
| R27MedianBlur | 3,805 | 241 |
| R28MotionBlur | 4,223 | — |
| R29CLAHE | 2,373 | — |
| R30Brightness | 4,622 | 3,519 |
| R31Contrast | 4,642 | 3,198 |
| R32Blur | 4,821 | 3,048 |
| R33ChannelDropout | 4,972 | — |
| R34LinearIllumination | 3,866 | — |
| R35CornerIllumination | 4,090 | — |
| R36GaussianIllumination | 3,930 | — |
| R37Hue | 4,274 | — |
| R38PlasmaBrightness | 2,461 | — |
| R39PlasmaContrast | 2,162 | — |
| R40PlasmaShadow | 2,489 | — |
| R41Rain | 4,069 | — |
| R42SaltAndPepper | 4,147 | — |
| R43Saturation | 4,136 | 3,470 |
| R44Snow | 3,857 | — |
| R45OpticalDistortion | 3,110 | — |
| R46Shear | 2,658 | 2,550 |
| R47ThinPlateSpline | 858 | — |
| R48PhotoMetricDistort | 3,369 | — |
| R49ColorJiggle | 3,526 | — |
| R50LongestMaxSize+RandomCrop224 | 3,658 | — |
| R51SmallestMaxSize+RandomCrop224 | 3,221 | — |
| R52Transpose | 4,942 | 3,676 |
| R53RandomRotate90 | 5,002 | — |
| R54RandomJigsaw | 4,662 | — |
| R55EnhanceEdge | 4,389 | 2,601 |
| R56EnhanceDetail | 4,720 | 2,700 |
| R57UnsharpMask | 3,120 | 2,132 |
Measurement setup and limits
g2-standard-16, nvidia-l4; 10,000 selected ImageNet JPEGs. Batch size 256, 15 workers, prefetch factor 2; persistent workers enabled. Output: cuda float16, BCHW 256×3×224×224.
Seeds: 137, 138, 139. Each observation follows 1 warm-up batch and times 32 batches, ending with CUDA synchronization. Pipeline construction and worker startup are outside throughput timing; prefetch effects remain. JPEG files are prewarmed, so this measures filesystem reads and decoding with a warm page cache.
NVML samples peak process GPU memory every 50 ms, from pipeline construction through final synchronization and cleanup. Brief peaks can be missed. The measurements include no model and do not establish training speed or augmentation quality. Seeds do not guarantee identical augmentation draws across libraries. Observed ranges describe variation between runs; they are not confidence intervals.
In this published run, DALI Crop includes resizing the short side, and DALI Affine omits rotation and shear.
Run 3f8e2e315710528399b8e82e2359ab85c58c809644595b68a92fb9d83492cc8c · 759 measurements · measured source 5fc35f6 · machine-readable results · paper and methodology. This is the run reported in the paper.
Conversion Guide
The main conversion is to move from calling Pillow methods one at a time to defining an AlbumentationsX Compose pipeline.
- Convert PIL images to NumPy arrays before augmentation.
- Replace manual randomness with transform-level p values.
- Keep image-adjacent targets in the same Compose call instead of transforming them separately.
- Convert back to PIL only if downstream code specifically needs PIL objects.
from PIL import Image, ImageEnhance
import random
image = Image.open("image.jpg").convert("RGB")
if random.random() < 0.5:
image = image.transpose(Image.Transpose.FLIP_LEFT_RIGHT)
if random.random() < 0.5:
image = ImageEnhance.Brightness(image).enhance(1.2)import albumentations as A
import cv2
transform = A.Compose([
A.HorizontalFlip(p=0.5),
A.RandomBrightnessContrast(brightness_limit=0.2, contrast_limit=0.0, p=0.5),
])
image = cv2.cvtColor(cv2.imread("image.jpg"), cv2.COLOR_BGR2RGB)
image = transform(image=image)["image"]Use AlbumentationsX When
- Training-time augmentation where randomness, replayability, and target synchronization matter.
- Segmentation, detection, keypoint, OCR, document, satellite, medical, or any multi-target computer vision workflow.
- CPU data-loader pipelines where augmentation speed can become the training bottleneck.
Use PIL/Pillow When
- Image IO, format conversion, drawing, thumbnails, and lightweight one-off image manipulation.
- Small scripts where you only touch a single image and do not need labels, masks, or reproducible random policies.

