The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
Sharing a linelist externally — with collaborators, for CRAN
examples, in a research supplement — requires that directly identifying
information is removed. molting() automates this using a
cryptographic hash: each row gets a unique fingerprint derived from its
identifiers, the identifiers are stripped, and the fingerprint is stored
in a lookup table that only authorised personnel hold.
The name comes from the natural process of moulting: a bird sheds its distinctive, identifiable plumage and temporarily becomes more uniform. The old plumage is not destroyed — it re-grows from the same follicles. The lookup table is those follicles.
molting() uses regular-expression pattern matching to
detect PII columns. The default patterns cover:
name, surname,
firstname, lastnamedob, birthmrn, urn,
medicare, patient_id, subject_id,
_id$address, street,
phone, emailAge category variables (age2cat,
age5cat, age10cat, etc.) are automatically
preserved — they are not directly identifying.
patient_data <- data.frame(
patient_name = c("John Doe", "Jane Smith"),
dob = as.Date(c("1980-01-01", "1975-05-15")),
mrn = c("12345", "67890"),
age5cat = factor(c("18-64", "18-64")), # preserved automatically
diagnosis = c("Condition A", "Condition B"),
lab_value = c(120, 95)
)
result <- suppressMessages(molting(patient_data))
names(result$deidentified) # hash + retained columns
#> [1] "row_hash" "age5cat" "diagnosis" "lab_value"
names(result$lookup) # hash + removed columns
#> [1] "row_hash" "patient_name" "dob" "mrn"
molting() returns a named list with two elements when
return_lookup = TRUE (the default):
$deidentified — the de-identified data frame, with
row_hash as the first column$lookup — the lookup table, with row_hash
plus all removed identifier columnsstr(result, max.level = 1)
#> List of 2
#> $ deidentified: tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
#> $ lookup : tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
head(result$lookup)
#> # A tibble: 2 × 4
#> row_hash patient_name dob mrn
#> <chr> <chr> <chr> <chr>
#> 1 89573bbf928ef324ba95e8d04fd1701dfc5c15265ab21b3dbbb2… John Doe 1980… 12345
#> 2 7b2536bf2d008eb3555f9404d1762dcc529e1a7d5e6880c66341… Jane Smith 1975… 67890
Store $lookup securely and separately from
$deidentified. Consider encrypting the lookup file before
archiving. In a Queensland Health context, the lookup table should
remain within the Health Service network.
The default is SHA-256, which provides strong collision resistance for typical surveillance dataset sizes (tens of thousands of rows). For very large datasets where speed matters more than collision resistance, SHA-1 or MD5 are faster but should not be used where linkage integrity is critical.
# SHA-256 (default, recommended)
result_256 <- molting(patient_data, hash_method = "sha256")
# MD5 (shorter hash, faster, lower collision resistance)
result_md5 <- molting(patient_data, hash_method = "md5")
# Blake3 (fast and cryptographically strong — good for large datasets)
result_b3 <- molting(patient_data, hash_method = "blake3")
By default, molting() hashes all detected PII columns.
Supply id_cols to override this — useful when you want a
shorter, more stable hash based only on a true unique identifier.
# Hash only on MRN and DOB — more stable if name variations exist
result_ids <- suppressMessages(
molting(patient_data, id_cols = c("mrn", "dob"))
)
result_ids$lookup
#> # A tibble: 2 × 4
#> row_hash patient_name dob mrn
#> <chr> <chr> <chr> <chr>
#> 1 255ad8c4452e80121bdee96a35de7cea0c25475a5e9e6f73088c… John Doe 1980… 12345
#> 2 107510d61f22148ec63b636654c41efea411429ace78b9c371f0… Jane Smith 1975… 67890
Use additional_pii_cols for dataset-specific identifiers
that don’t match the default patterns.
patient_data2 <- patient_data
patient_data2$study_code <- c("SC-001","SC-002")
result2 <- suppressMessages(
molting(patient_data2, additional_pii_cols = "study_code")
)
names(result2$deidentified)
#> [1] "row_hash" "age5cat" "diagnosis" "lab_value"
If you genuinely do not need to relink (e.g. producing a public-use
file), set return_lookup = FALSE. This is irreversible —
there is no way to recover the original identifiers.
deidentified_only <- suppressMessages(
molting(patient_data, return_lookup = FALSE)
)
class(deidentified_only) # a data frame, not a list
#> [1] "tbl_df" "tbl" "data.frame"
If two rows produce the same hash (extremely rare with SHA-256 for
realistic dataset sizes but possible with very short hashes like CRC32),
molting() warns you. If a collision is detected, switch to
a stronger algorithm or add more columns to id_cols.
Use [homing()] to relink the de-identified data when authorised (see
vignette("homing")).
For aggregated outputs — monthly counts from roost() —
de-identification may not be necessary at all if the counts are not
small enough to be re-identifying. The ABS cell-suppression threshold of
5 is a useful rule of thumb.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.