RGB Input Pipeline Benchmark

Compare RGB JPEG-to-CUDA throughput and peak GPU memory across AlbumentationsX, Pillow, TorchVision, Kornia, and NVIDIA DALI. 57 recipes, 7 execution paths, one selected run.

RGB input pipeline results

Every path reads RGB JPEGs, prepares the recipe, and delivers a synchronized CUDA batch. CPU and GPU labels identify where augmentation runs; normalization runs on GPU for every path.

Mean relative throughput on the same 11 recipes. AlbumentationsX = 1×; higher is faster.
Mean relative throughput on the same 11 recipes. AlbumentationsX = 1×; higher is faster. Scroll horizontally to see the full chart. Open the image for full size.
Median peak process GPU memory on the same 11 recipes, in MiB. Lower is better.
Median peak process GPU memory on the same 11 recipes, in MiB. Lower is better. Scroll horizontally to see the full chart. Open the image for full size.
Per-recipe throughput on the 11 recipes shared by all paths. Ratios are relative to AX; its absolute throughput is shown in images/s.
Per-recipe throughput on the 11 recipes shared by all paths. Ratios are relative to AX; its absolute throughput is shown in images/s. Scroll horizontally to see the full chart. Open the image for full size.
Same 11 recipes for every row. Throughput is the arithmetic mean of per-recipe ratios to AlbumentationsX; memory is the median of per-recipe peak-memory medians.
Measured pathThroughput / AXGPU memory (MiB)
DALI GPU1.17×2,086
AlbumentationsX CPU1.00×1,852
TorchVision CPU0.78×1,814
Pillow CPU0.74×1,814
TorchVision GPU0.72×1,972
Kornia GPU0.45×1,900
Kornia CPU0.40×1,776

Each comparison uses its own shared recipe set. Averages from different sets cannot rank all libraries. The table below includes every measured recipe for these paths, including recipes outside the summary set. Recipe names are shortened; hover over a name for its full pipeline.

Benchmark metric

Higher throughput is better. Values are medians across seeds. Hover for the observed range. A dash means no measured result.

R01Resize2244,7512,3903,6713,8522,0631,9905,070
R02RandomCrop2244,7403,7684,4244,3752,0732,1075,029
R03RandomResizedCrop4,7852,7213,5183,5581,7761,7465,091
R04HorizontalFlip4,7233,7594,2074,3742,0112,0844,990
R05VerticalFlip4,9073,5674,3194,5941,9802,1514,992
R06Pad+RandomCrop2244,3973,6593,6853,6264,993
R07Rotate3,3523,5552,5851,5351,4662,1225,025
R08Affine3,0492,6472,3841,4781,4452,1385,031
R09Perspective2,8792,1798861,303
R10Elastic1,99023622102271
R11ColorJitter3,5231,2037311,0062,0475,047
R12ChannelShuffle5,0264,2254,3371,9442,120
R13Grayscale5,1573,6873,8914,3871,8922,113
R14RGBShift4,3481,8892,149
R15GaussianBlur4,6792,3932,4472,8111,1892,1554,990
R16GaussianNoise3,2881,6622,0225,025
R17Invert5,0763,5404,0674,4301,9202,167
R18Posterize5,1103,4934,3364,4531,7512,135
R19Solarize4,6073,5063,6714,5201,6552,142
R20Sharpen4,2232,1173,2391,2412,144
R21AutoContrast4,2633,5132,6873,8591,6972,092
R22Equalize3,9863,2453,0641,5911,2655235,033
R23Erasing4,9384,0102,9071,5585,025
R24JpegCompression4,2323,1833,4486842,1514,997
R25RandomGamma4,9691,6862,114
R26PlankianJitter4,5381,8572,125
R27MedianBlur3,8052411241,015
R28MotionBlur4,2231,2102,127
R29CLAHE2,3736921044,840
R30Brightness4,6223,5193,8234,4341,8722,1015,041
R31Contrast4,6423,1983,3773,3331,8682,0925,025
R32Blur4,8213,0481,2842,147
R33ChannelDropout4,9721,9112,084
R34LinearIllumination3,8661,698
R35CornerIllumination4,0901,481
R36GaussianIllumination3,9301,492247
R37Hue4,2741,1262,1005,047
R38PlasmaBrightness2,4614692,083
R39PlasmaContrast2,1624772,103
R40PlasmaShadow2,4898322,123
R41Rain4,0691,455517
R42SaltAndPepper4,1471,4794375,009
R43Saturation4,1363,4701,1242,0835,047
R44Snow3,8571,1312,059
R45OpticalDistortion3,1101,4332,100
R46Shear2,6582,5501,5085,052
R47ThinPlateSpline8586962,121
R48PhotoMetricDistort3,3691,170670
R49ColorJiggle3,5261,2217427582,1064,953
R50LongestMaxSize+RandomCrop2243,6581,0731,066
R51SmallestMaxSize+RandomCrop2243,221845834
R52Transpose4,9423,676
R53RandomRotate905,0021,4622,106
R54RandomJigsaw4,6621,7512,117
R55EnhanceEdge4,3892,601
R56EnhanceDetail4,7202,700
R57UnsharpMask3,1202,132

Measurement setup and limits

Throughput measures batch consumption and final CUDA synchronization. GPU memory is sampled from pipeline construction through cleanup.
Throughput measures batch consumption and final CUDA synchronization. GPU memory is sampled from pipeline construction through cleanup. Scroll horizontally to see the full chart. Open the image for full size.

g2-standard-16, nvidia-l4; 10,000 selected ImageNet JPEGs. Batch size 256, 15 workers, prefetch factor 2; persistent workers enabled. Output: cuda float16, BCHW 256×3×224×224.

Seeds: 137, 138, 139. Each observation follows 1 warm-up batch and times 32 batches, ending with CUDA synchronization. Pipeline construction and worker startup are outside throughput timing; prefetch effects remain. JPEG files are prewarmed, so this measures filesystem reads and decoding with a warm page cache.

NVML samples peak process GPU memory every 50 ms, from pipeline construction through final synchronization and cleanup. Brief peaks can be missed. The measurements include no model and do not establish training speed or augmentation quality. Seeds do not guarantee identical augmentation draws across libraries. Observed ranges describe variation between runs; they are not confidence intervals.

In this published run, DALI Crop includes resizing the short side, and DALI Affine omits rotation and shear.

Run 3f8e2e315710528399b8e82e2359ab85c58c809644595b68a92fb9d83492cc8c · 759 measurements · measured source 5fc35f6 · machine-readable results · paper and methodology. This is the run reported in the paper.