The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

leakaudit

Detects near-duplicate images across train/validation/test splits using perceptual hashing (dHash), reports the resulting leakage, and produces a corrected, leak-free split assignment.

Why

When building an image classification dataset, near-duplicate images (same photo re-saved, resized, lightly cropped, or recompressed) commonly end up split across train and test. This silently inflates test-set performance, since the model has effectively already seen a near-identical copy of the “unseen” example.

Install

# development version
remotes::install_github("anakincodex/leakaudit")

# from CRAN, once published
install.packages("leakaudit")

Usage

library(leakaudit)

hashes  <- compute_hashes(image_paths, split = split_labels)
grouped <- find_duplicate_groups(hashes, threshold = 5)
report  <- dhash_audit(grouped)
print(report)

clean <- clean_splits(grouped, priority = c("train", "val", "test"))

How this differs from other leakage packages

A few CRAN packages deal with “leakage” in machine learning workflows, but none of them work at the image level:

leakaudit is specifically about auditing and fixing near-duplicate contamination across image dataset splits, and its dhash_audit() function is unrelated to bioLeak::audit_leakage(), which audits fitted models on tabular data via permutation testing.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.