The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
Data is synthesised by sampling from a multivariate cumulative
distribution (Copula), using the simstudy package.
Data can be synthesised from marginal distributions using the
synthesise_data() function:
User specified correlations can be added to the synthesised data by
supplying a list of correlations to the correlations
parameter. Each correlation is created with the
correlation() function, giving the names of the two
variables and the correlation (rho) between them:
library(RESIDE)
marginals <- import_marginal_distributions()
simulated_data <- synthesise_data(
marginals,
correlations = list(
correlation("AGE", "RSBP", 0.3)
)
)Correlations should be between -1 and 1 and must be consistent with
each other (the resulting correlation matrix must be positive semi
definite), otherwise synthesise_data() will produce an
error.
Categorical variables are correlated using a single category,
specified with the factor_name.x or
factor_name.y parameters for the first or second variable
respectively. For example to correlate male patients (SEX
of M) with age:
simulated_data <- synthesise_data(
marginals,
correlations = list(
correlation("SEX", "AGE", -0.2, factor_name.x = "M"),
correlation("RATRIAL", "RSBP", 0.1, factor_name.x = "Y")
)
)A categorical variable without a factor name will produce an error.
When the marginals were obtained from multiple data frames, variables
that are present in more than one data frame (other than the common
columns) must specify which data frame they belong to. Use
df_name when both variables belong to the same data frame,
or df_name.x and df_name.y when they belong to
different data frames:
simulated_data <- synthesise_data(
marginals,
correlations = list(
# Both variables in the same data frame
correlation("DOMAIN", "AESTDY", 0.2, df_name = "ae", factor_name.x = "AE"),
# Variables in different data frames
correlation(
"SEX",
"AESEV",
0.3,
df_name.x = "dm",
df_name.y = "ae",
factor_name.x = "M",
factor_name.y = "SEVERE"
)
)
)A data frame name specified for a variable that only appears in a single data frame, but not that data frame, will be ignored with a warning.
Correlated variables are synthesised together, with one row per subject, and joined to each data frame by subject. Correlations between data frames are therefore between subjects, and a correlated variable takes a single value for each subject within a data frame.
NB The correlation specified is that of the underlying multivariate normal distribution (Copula), the correlation observed in the synthesised data will be weaker, particularly for binary and categorical variables. It is not possible to entirely maintain all the marginal distributions when specifying correlations, this is a known limitation and is not likely to change.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.