
Description
Before training vision models, datasets hide duplicates, blurry shots, outliers and wrong labels, impossible to check by hand at scale. fastdup rapidly analyzes image and video datasets to find duplicates, anomalies and label errors.
It handles millions of images on a CPU with visual reports, cutting data cleaning costs.
Dedup:Near-duplicates.
Anomalies:Blur and outliers.
Labels:Mislabels.
Fast:Millions on CPU.
It handles millions of images on a CPU with visual reports, cutting data cleaning costs.
Features
Dedup:Near-duplicates.
Anomalies:Blur and outliers.
Labels:Mislabels.
Fast:Millions on CPU.

