The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

Why de-identify?

Sharing a linelist externally — with collaborators, for CRAN examples, in a research supplement — requires that directly identifying information is removed. molting() automates this using a cryptographic hash: each row gets a unique fingerprint derived from its identifiers, the identifiers are stripped, and the fingerprint is stored in a lookup table that only authorised personnel hold.

The name comes from the natural process of moulting: a bird sheds its distinctive, identifiable plumage and temporarily becomes more uniform. The old plumage is not destroyed — it re-grows from the same follicles. The lookup table is those follicles.


What gets removed

molting() uses regular-expression pattern matching to detect PII columns. The default patterns cover:

  • Names: name, surname, firstname, lastname
  • Dates: dob, birth
  • Identifiers: mrn, urn, medicare, patient_id, subject_id, _id$
  • Contact: address, street, phone, email

Age category variables (age2cat, age5cat, age10cat, etc.) are automatically preserved — they are not directly identifying.

patient_data <- data.frame(
  patient_name = c("John Doe", "Jane Smith"),
  dob          = as.Date(c("1980-01-01", "1975-05-15")),
  mrn          = c("12345", "67890"),
  age5cat      = factor(c("18-64", "18-64")),   # preserved automatically
  diagnosis    = c("Condition A", "Condition B"),
  lab_value    = c(120, 95)
)

result <- suppressMessages(molting(patient_data))
names(result$deidentified)   # hash + retained columns
#> [1] "row_hash"  "age5cat"   "diagnosis" "lab_value"
names(result$lookup)         # hash + removed columns
#> [1] "row_hash"     "patient_name" "dob"          "mrn"

The output list

molting() returns a named list with two elements when return_lookup = TRUE (the default):

  • $deidentified — the de-identified data frame, with row_hash as the first column
  • $lookup — the lookup table, with row_hash plus all removed identifier columns
str(result, max.level = 1)
#> List of 2
#>  $ deidentified: tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
#>  $ lookup      : tibble [2 × 4] (S3: tbl_df/tbl/data.frame)
head(result$lookup)
#> # A tibble: 2 × 4
#>   row_hash                                              patient_name dob   mrn  
#>   <chr>                                                 <chr>        <chr> <chr>
#> 1 89573bbf928ef324ba95e8d04fd1701dfc5c15265ab21b3dbbb2… John Doe     1980… 12345
#> 2 7b2536bf2d008eb3555f9404d1762dcc529e1a7d5e6880c66341… Jane Smith   1975… 67890

Store $lookup securely and separately from $deidentified. Consider encrypting the lookup file before archiving. In a Queensland Health context, the lookup table should remain within the Health Service network.


Hash algorithm selection

The default is SHA-256, which provides strong collision resistance for typical surveillance dataset sizes (tens of thousands of rows). For very large datasets where speed matters more than collision resistance, SHA-1 or MD5 are faster but should not be used where linkage integrity is critical.

# SHA-256 (default, recommended)
result_256 <- molting(patient_data, hash_method = "sha256")

# MD5 (shorter hash, faster, lower collision resistance)
result_md5 <- molting(patient_data, hash_method = "md5")

# Blake3 (fast and cryptographically strong — good for large datasets)
result_b3 <- molting(patient_data, hash_method = "blake3")

Controlling which columns are hashed

By default, molting() hashes all detected PII columns. Supply id_cols to override this — useful when you want a shorter, more stable hash based only on a true unique identifier.

# Hash only on MRN and DOB — more stable if name variations exist
result_ids <- suppressMessages(
  molting(patient_data, id_cols = c("mrn", "dob"))
)
result_ids$lookup
#> # A tibble: 2 × 4
#>   row_hash                                              patient_name dob   mrn  
#>   <chr>                                                 <chr>        <chr> <chr>
#> 1 255ad8c4452e80121bdee96a35de7cea0c25475a5e9e6f73088c… John Doe     1980… 12345
#> 2 107510d61f22148ec63b636654c41efea411429ace78b9c371f0… Jane Smith   1975… 67890

Adding columns to the removal list

Use additional_pii_cols for dataset-specific identifiers that don’t match the default patterns.

patient_data2 <- patient_data
patient_data2$study_code <- c("SC-001","SC-002")

result2 <- suppressMessages(
  molting(patient_data2, additional_pii_cols = "study_code")
)
names(result2$deidentified)
#> [1] "row_hash"  "age5cat"   "diagnosis" "lab_value"

Irreversible de-identification

If you genuinely do not need to relink (e.g. producing a public-use file), set return_lookup = FALSE. This is irreversible — there is no way to recover the original identifiers.

deidentified_only <- suppressMessages(
  molting(patient_data, return_lookup = FALSE)
)
class(deidentified_only)   # a data frame, not a list
#> [1] "tbl_df"     "tbl"        "data.frame"

Hash collisions

If two rows produce the same hash (extremely rare with SHA-256 for realistic dataset sizes but possible with very short hashes like CRC32), molting() warns you. If a collision is detected, switch to a stronger algorithm or add more columns to id_cols.


What comes next

Use [homing()] to relink the de-identified data when authorised (see vignette("homing")).

For aggregated outputs — monthly counts from roost() — de-identification may not be necessary at all if the counts are not small enough to be re-identifying. The ABS cell-suppression threshold of 5 is a useful rule of thumb.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.