The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

DataAudit

R-CMD-check Codecov test coverage

Overview

DataAudit is an R package for systematic data-quality auditing and validation.

It provides tools for detecting common data-quality problems, defining reusable validation rules, generating structured audit reports, and calculating an overall data-quality score.

DataAudit supports checks for missing values, blank values, duplicates, infinite values, constant and near-zero variance variables, outliers, ranges, categories, data types, patterns, dates, identifiers, lengths, character case, whitespace, sequences, dependencies, uniqueness, and cross-variable consistency.

Installation

You can install DataAudit from GitHub with:

# install.packages("pak")
pak::pak("vinodhpmd/DataAudit")

Then load the package:

library(DataAudit)

Quick example

Create a dataset containing several data-quality problems:

dat <- data.frame(
  ID = c(
    "A001",
    "A002",
    "A002",
    "BAD"
  ),
  Species = c(
    "Dog",
    "Horse",
    "",
    "Cat"
  ),
  Age = c(
    5,
    40,
    -2,
    NA
  ),
  stringsAsFactors = FALSE
)

dat
##     ID Species Age
## 1 A001     Dog   5
## 2 A002   Horse  40
## 3 A002          -2
## 4  BAD     Cat  NA

General data audit

Use audit_data() for a general first-pass audit:

report <- audit_data(
  dat,
  include_nzv = FALSE
)

report
## DataAudit Report
## ================
## 
## Rows:                  4
## Variables:             3
## Variables with issues: 2
## Duplicated rows:       0
## Total issues:          2
## Rule violations:       0
## Overall issues:        2
## 
## Variables with issues
## ---------------------
##  Variable     Class Missing MissingPercent Blank Infinite Unique Constant
##   Species character       0              0     1        0      4    FALSE
##       Age   numeric       1             25     0        0      3    FALSE
##  NearZeroVariance Issue
##             FALSE  TRUE
##             FALSE  TRUE
## 
## No duplicated rows detected.

Detailed results are available through:

report$overview
##                         Metric Value
## 1                         Rows     4
## 2                    Variables     3
## 3               Missing values     1
## 4                 Blank values     1
## 5              Infinite values     0
## 6              Duplicated rows     0
## 7           Constant variables     0
## 8 Near-zero variance variables     0
## 9        Variables with issues     2
report$variables
##         Variable     Class Missing MissingPercent Blank Infinite Unique
## ID            ID character       0              0     0        0      3
## Species  Species character       0              0     1        0      4
## Age          Age   numeric       1             25     0        0      3
##         Constant NearZeroVariance Issue
## ID         FALSE            FALSE FALSE
## Species    FALSE            FALSE  TRUE
## Age        FALSE            FALSE  TRUE
report$duplicates
##   Row Duplicate
## 1   1     FALSE
## 2   2     FALSE
## 3   3     FALSE
## 4   4     FALSE
report$summary
##   Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
## 1    4         3                   2             0           2              0
##   OverallIssues
## 1             2

Individual checks

Specific checks can also be performed independently.

missing_check(dat)
##   Variable Total Missing Complete MissingPercent CompletePercent
## 1       ID     4       0        4              0             100
## 2  Species     4       0        4              0             100
## 3      Age     4       1        3             25              75
blank_check(dat)
##         Variable TextVariable Missing Blank BlankPercent HasBlank
## ID            ID         TRUE       0     0            0    FALSE
## Species  Species         TRUE       0     1           25     TRUE
## Age          Age        FALSE       1     0            0    FALSE
duplicate_check(dat)
##   Row Duplicate
## 1   1     FALSE
## 2   2     FALSE
## 3   3     FALSE
## 4   4     FALSE
range_check(
  dat,
  variable = "Age",
  min = 0,
  max = 30
)
##   Row Value BelowMinimum AboveMaximum OutOfRange
## 1   1     5        FALSE        FALSE      FALSE
## 2   2    40        FALSE         TRUE       TRUE
## 3   3    -2         TRUE        FALSE       TRUE
## 4   4    NA        FALSE        FALSE      FALSE
category_check(
  dat,
  variable = "Species",
  allowed = c(
    "Dog",
    "Cat"
  )
)
##   Row Value Missing ValidCategory InvalidCategory
## 1   1   Dog   FALSE          TRUE           FALSE
## 2   2 Horse   FALSE         FALSE            TRUE
## 3   3         FALSE         FALSE            TRUE
## 4   4   Cat   FALSE          TRUE           FALSE

Validation rules

Reusable validation rules can be defined with audit_rules():

rules <- audit_rules(
  range = list(
    Age = c(0, 30)
  ),
  category = list(
    Species = c(
      "Dog",
      "Cat"
    )
  ),
  unique = "ID",
  required = "Species",
  pattern = list(
    ID = "^A[0-9]{3}$"
  )
)

Apply the rules directly:

rule_results <- apply_audit_rules(
  dat,
  rules
)

rule_results
## DataAudit Rule Validation
## =========================
## 
## Status:           Violations detected
## Total violations: 8
## 
## Rule summary
## ------------
##      Rule Violations
##     Range          2
##  Category          2
##    Unique          2
##  Required          1
##   Pattern          1
## 
## Variables with rule violations
## ------------------------------
## - Range: Age
## - Category: Species
## - Unique: ID
## - Required: Species
## - Pattern: ID

Integrated audit

Custom validation rules can be incorporated directly into the general audit:

report <- audit_data(
  dat,
  include_nzv = FALSE,
  rules = rules
)

report
## DataAudit Report
## ================
## 
## Rows:                  4
## Variables:             3
## Variables with issues: 2
## Duplicated rows:       0
## Total issues:          2
## Rule violations:       8
## Overall issues:        10
## 
## Variables with issues
## ---------------------
##  Variable     Class Missing MissingPercent Blank Infinite Unique Constant
##   Species character       0              0     1        0      4    FALSE
##       Age   numeric       1             25     0        0      3    FALSE
##  NearZeroVariance Issue
##             FALSE  TRUE
##             FALSE  TRUE
## 
## No duplicated rows detected.
## 
## Custom rule validation
## ----------------------
##      Rule Violations
##     Range          2
##  Category          2
##    Unique          2
##  Required          1
##   Pattern          1
## 
## Total rule violations: 8

The report now combines automatic auditing and custom validation:

report$summary
##   Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
## 1    4         3                   2             0           2              8
##   OverallIssues
## 1            10
report$rules
## DataAudit Rule Validation
## =========================
## 
## Status:           Violations detected
## Total violations: 8
## 
## Rule summary
## ------------
##      Rule Violations
##     Range          2
##  Category          2
##    Unique          2
##  Required          1
##   Pattern          1
## 
## Variables with rule violations
## ------------------------------
## - Range: Age
## - Category: Species
## - Unique: ID
## - Required: Species
## - Pattern: ID

A compact summary can be obtained with:

summary(report)
## $dataset
##   Rows Variables VariablesWithIssues DuplicateRows TotalIssues RuleViolations
## 1    4         3                   2             0           2              8
##   OverallIssues
## 1            10
## 
## $variables_with_issues
##         Variable     Class Missing MissingPercent Blank Infinite Unique
## Species  Species character       0              0     1        0      4
## Age          Age   numeric       1             25     0        0      3
##         Constant NearZeroVariance Issue
## Species    FALSE            FALSE  TRUE
## Age        FALSE            FALSE  TRUE
## 
## $duplicate_rows
## [1] Row       Duplicate
## <0 rows> (or 0-length row.names)
## 
## $rules
##       Rule Violations
## 1    Range          2
## 2 Category          2
## 3   Unique          2
## 4 Required          1
## 5  Pattern          1
## 
## attr(,"class")
## [1] "summary.DataAuditReport" "list"

Data-quality score

Calculate an overall data-quality score with:

score <- audit_score(report)

score
## DataAudit Quality Score
## =======================
## 
## Score:               54.55 / 100
## Quality:             Poor
## 
## Dataset
## -------
## Rows:                4
## Variables:           3
## Data cells:          12
## 
## Issues
## ------
## Automatic issues:    2
## Rule violations:     8
## Overall issues:      10

The score ranges from 0 to 100, with higher scores indicating fewer detected data-quality issues relative to dataset size.

DataAudit classifies scores as:

The score is intended as a compact audit summary and should be interpreted alongside the underlying diagnostics and the intended use of the dataset.

Main functionality

DataAudit currently provides functions for:

Typical workflow

Raw data
   |
   v
audit_data()
   |
   +----> Automatic checks
   |
   +----> audit_rules()
              |
              v
        Rule validation
   |
   v
DataAuditReport
   |
   +----> summary()
   |
   +----> audit_score()

Authors

Vinodh Kumar Obli Rajendran
ORCID: 0000-0002-7232-4122

Keerthi Aaradhana

License

DataAudit is released under the MIT License.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.