The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
Detects near-duplicate images across train/validation/test splits using perceptual hashing (dHash), reports the resulting leakage, and produces a corrected, leak-free split assignment.
When building an image classification dataset, near-duplicate images (same photo re-saved, resized, lightly cropped, or recompressed) commonly end up split across train and test. This silently inflates test-set performance, since the model has effectively already seen a near-identical copy of the “unseen” example.
# development version
remotes::install_github("anakincodex/leakaudit")
# from CRAN, once published
install.packages("leakaudit")library(leakaudit)
hashes <- compute_hashes(image_paths, split = split_labels)
grouped <- find_duplicate_groups(hashes, threshold = 5)
report <- dhash_audit(grouped)
print(report)
clean <- clean_splits(grouped, priority = c("train", "val", "test"))A few CRAN packages deal with “leakage” in machine learning workflows, but none of them work at the image level:
dhash(), phash()) but has no concept
of dataset splits or leakage reporting.leakaudit is specifically about auditing and fixing near-duplicate
contamination across image dataset splits, and its
dhash_audit() function is unrelated to
bioLeak::audit_leakage(), which audits fitted models on
tabular data via permutation testing.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.