The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
make_split_plan(group = , batch = )
directly?stratum column mean, and does splitGraph
stratify?make_split_plan(group = , batch = )
directly?For a dataset with one subject column and one batch column you
should: bioLeak’s make_split_plan() reaches the same
grouping in one call, and splitGraph adds a hop. splitGraph earns its
place when the structure is not one clean column per axis:
time_index that contradicts
timepoint_precedes: validate_graph() reports
these before any fold exists. Column-based grouping silently produces
folds from inconsistent metadata.mode = "composite" merges samples into connected components
over all chosen relations. bioLeak’s combined mode
instead picks a primary axis and deletes training rows sharing a
secondary level with the test set. Same metadata, different partitions;
splitGraph makes the choice explicit and reproducible.relatedness_edges_from_kinship() and
spatial_edges_from_coords() threshold them into edges and
the derivation takes connected components. No categorical column can
express this.split_spec records which relations, thresholds and strategy
produced each group, and travels as schema-checked JSON to Python or any
other consumer.Strict composite grouping is transitive closure. If sample A shares a subject with B, and B shares a batch with C, then A, B and C land in one group even though A and C share nothing directly. With a few large batches this collapses most of the dataset into one component:
meta <- data.frame(
sample_id = paste0("S", 1:6),
subject_id = c("P1", "P1", "P2", "P2", "P3", "P3"),
batch_id = c("B1", "B2", "B2", "B3", "B3", "B1"),
stringsAsFactors = FALSE
)
g <- graph_from_metadata(meta)
table(grouping_vector(derive_split_constraints(g, "composite", via = c("subject", "batch"))))
#>
#> component_1
#> 6Every subject bridges two batches, so all six samples form a single group and no split is possible. Three remedies, in order of preference:
spec$block_vars).strategy = "rule_based": each sample is grouped by
the first relation in priority that is available to it, so
relations do not chain.detect_dependency_components() to see the component
sizes before deriving, and summarize_leakage_risks() to see
which leakage paths a given mode actually severs (severed
column).grouping_vector(derive_split_constraints(g, "composite", strategy = "rule_based",
via = c("subject", "batch"),
priority = c("subject", "batch")))
#> S1 S2 S3
#> "composite_subject:P1" "composite_subject:P1" "composite_subject:P2"
#> S4 S5 S6
#> "composite_subject:P2" "composite_subject:P3" "composite_subject:P3"relatedness_edges_from_kinship(pairs, threshold) keeps a
pair when its kinship is at least the threshold;
spatial_edges_from_coords(coords, radius) keeps a pair when
its distance is at most the radius. The derivation then forms
connected components over the kept edges. Two consequences:
graph$metadata$edge_sources and
spec$metadata$threshold, so a reader of the spec can see
which cut produced the groups.pairs <- data.frame(id1 = c("P1", "P2"), id2 = c("P2", "P3"), kinship = c(0.26, 0.13))
meta <- data.frame(sample_id = c("S1", "S2", "S3"), subject_id = c("P1", "P2", "P3"))
build <- function(threshold) {
g <- build_dependency_graph(
list(create_nodes(meta, "Sample", "sample_id"), create_nodes(meta, "Subject", "subject_id")),
list(create_edges(meta, "sample_id", "subject_id", "Sample", "Subject", "sample_belongs_to_subject"),
relatedness_edges_from_kinship(pairs, threshold = threshold))
)
grouping_vector(derive_split_constraints(g, "relatedness"))
}
build(0.25) # only P1~P2 pass: {S1,S2}, {S3}
#> S1 S2 S3
#> "relatedness:component_1" "relatedness:component_1" "relatedness:component_2"
build(0.10) # P2~P3 also passes and chains: {S1,S2,S3}
#> S1 S2 S3
#> "relatedness:component_1" "relatedness:component_1" "relatedness:component_1"stratum column mean, and does splitGraph
stratify?No. stratum is an annotation: the outcome level
attached to each sample (from sample_has_outcome, or the
subject’s outcome via subject_has_outcome). It is exposed
through spec$stratum_var so a consumer such as
scikit-learn’s StratifiedGroupKFold or bioLeak’s
stratify = TRUE can balance folds. Balancing is execution
and belongs downstream; splitGraph only describes.
schema_version (currently 0.3.0) is independent of the
package version. It describes the on-disk JSON contract.NA and ignore unknown ones. A
differing major loads with a warning suggesting
migrate_dependency_graph_json() /
migrate_split_spec_json().inst/schema/<schema_version>/, and every written file
carries a $schema URL pointing at its own version, so the
reference stays valid after later bumps.
validate_graph_json() and
validate_split_spec_json() (or
read_*(validate = TRUE)) check a file against the installed
schema without a JSON Schema engine.The R reader is permissive at that boundary, which is worth seeing rather than taking on trust:
tiny <- data.frame(sample_id = c("S1", "S2"), subject_id = c("P1", "P2"),
stringsAsFactors = FALSE)
g_tiny <- graph_from_metadata(tiny)
p <- tempfile(fileext = ".json")
write_split_spec(as_split_spec(derive_split_constraints(g_tiny, "subject"),
graph = g_tiny), p)
# Pretend the file was written by a future splitGraph with a different major.
raw <- jsonlite::fromJSON(p, simplifyVector = FALSE)
raw$schema_version <- "1.0.0"
writeLines(jsonlite::toJSON(raw, auto_unbox = TRUE, null = "null"), p)
back <- withCallingHandlers(
read_split_spec(p),
warning = function(w) {
message("warning: ", conditionMessage(w))
invokeRestart("muffleWarning")
}
)
#> warning: Reading split_spec: JSON schema_version `1.0.0` differs in major version from installed splitGraph schema_version `0.3.0`. Loading anyway; consider `migrate_split_spec_json()` to upgrade the file.
class(back)
#> [1] "split_spec"
unlink(p)One asymmetry to know about if you consume the format outside R: the
shipped Python reader is stricter. Where R warns and loads,
splitspec raises a ValueError naming the
version and refuses the file, on the grounds that a non-interactive
consumer is better off failing than silently misreading a format it does
not know. Both implementations accept every schema sharing major
0.
Yes. graph_from_metadata() is an S3 generic, and the
SummarizedExperiment method reads colData() as
the metadata table. When colData has no
sample_id column the assay column names are used, so a
typical Bioconductor object needs no preparation at all. Pass
sample_id_col = to name a different column, and
columns = to map your own names onto the canonical ones
exactly as for a data frame.
meta <- data.frame(
sample_id = c("S1", "S2", "S3", "S4"),
subject_id = c("P1", "P1", "P2", "P2"),
batch_id = c("B1", "B2", "B1", "B2"),
stringsAsFactors = FALSE
)
se <- SummarizedExperiment::SummarizedExperiment(
assays = list(counts = matrix(0, nrow = 3, ncol = 4,
dimnames = list(NULL, meta$sample_id))),
colData = meta[, c("subject_id", "batch_id")]
)
g_se <- graph_from_metadata(se, graph_name = "from-se")
grouping_vector(derive_split_constraints(g_se, "subject"))
#> S1 S2 S3 S4
#> "subject:P1" "subject:P1" "subject:P2" "subject:P2"
# identical to building from the data frame directly
identical(
grouping_vector(derive_split_constraints(g_se, "subject")),
grouping_vector(derive_split_constraints(graph_from_metadata(meta), "subject"))
)
#> [1] TRUEEvery error inherits from splitgraph_error and carries a
code; see ?splitgraph_conditions for the
subclasses (splitgraph_schema_error,
splitgraph_reference_error,
splitgraph_ambiguity_error,
splitgraph_validation_error,
splitgraph_io_error).
Every step is linear in nodes plus edges. On a 20,000-sample
synthetic cohort the shipped benchmark
(inst/bench/pipeline.R) builds the graph, validates it,
derives a default composite constraint, enriches a spec and writes both
JSON files in about twelve seconds on a laptop. More than half of that
is writing the graph JSON (~6 s); the specification itself
writes in well under a second, and the derivation steps are hundredths
of a second. If you only need the handoff artifact, skip
write_dependency_graph().
Two caveats about the guard rail.
tests/testthat/test-performance.R is a wall-clock
budget at 5,000 samples with deliberately generous limits, not
a benchmark; only the composite derivation is additionally checked for
scaling, by timing 1,000 against 4,000 samples and failing if the ratio
exceeds 8 (a quadratic step would give roughly 16). Other steps could
therefore degrade somewhat without tripping it.
The one inherently quadratic output is the explicit sample-pair table
of detect_shared_dependencies() and
detect_dependency_components()$metadata$projection_edges,
whose size is the number of pairs sharing a target; grouping itself
never enumerates pairs.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.