The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

Package {datom}


Title: A Unified Framework for Versioned, Traceable Tabular Data
Version: 0.2.0
Description: Provides versioned storage for tabular data without a database or a server. Each table is written as an immutable, content-addressed version – identical content is detected and stored only once – while its version history and metadata are kept as code in a 'git' repository and the data itself in a local filesystem or cloud object storage ('S3'). Any past version can be read back exactly by its identifier, and each table records the sources it was derived from, so a project carries full data lineage. A lightweight reader role retrieves current or historical data from storage alone, without 'git' or write access, giving downstream analyses and pipelines a single versioned source of truth. It targets analytical and scientific data management, such as preparing clinical study datasets, and is designed as a foundation for higher-level governance tooling.
License: MIT + file LICENSE
URL: https://github.com/amashadihossein/datom, https://amashadihossein.github.io/datom/
BugReports: https://github.com/amashadihossein/datom/issues
Depends: R (≥ 4.1.0)
Imports: arrow, cli, digest, fs, glue, httr2, jsonlite, paws.storage, purrr, rlang, utils, yaml
Suggests: covr, git2r, knitr, mockery, rio, rmarkdown, testthat (≥ 3.0.0), withr
Config/testthat/edition: 3
Encoding: UTF-8
RoxygenNote: 7.3.3
VignetteBuilder: knitr
NeedsCompilation: no
Packaged: 2026-10-03 19:26:34 UTC; afshinmashadi-hossein
Author: Afshin Mashadi-Hossein [aut, cre, cph]
Maintainer: Afshin Mashadi-Hossein <amashadihossein@gmail.com>
Repository: CRAN
Date/Publication: 2026-10-03 19:50:02 UTC

datom: A Unified Framework for Versioned, Traceable Tabular Data

Description

Provides versioned storage for tabular data without a database or a server. Each table is written as an immutable, content-addressed version – identical content is detected and stored only once – while its version history and metadata are kept as code in a 'git' repository and the data itself in a local filesystem or cloud object storage ('S3'). Any past version can be read back exactly by its identifier, and each table records the sources it was derived from, so a project carries full data lineage. A lightweight reader role retrieves current or historical data from storage alone, without 'git' or write access, giving downstream analyses and pipelines a single versioned source of truth. It targets analytical and scientific data management, such as preparing clinical study datasets, and is designed as a foundation for higher-level governance tooling.

Key terms

Author(s)

Maintainer: Afshin Mashadi-Hossein amashadihossein@gmail.com [copyright holder]

See Also

Useful links:


Abbreviate SHA Hash

Description

Truncates a SHA-256 hash to a short prefix for display. Accepts character vectors; NA values pass through unchanged.

Usage

.datom_abbreviate_sha(sha, n = 8L)

Arguments

sha

Character vector of SHA hashes.

n

Number of characters to keep. Default 8.

Value

Character vector of abbreviated hashes.


Add This Edit's Rows to Whatever Log the Object Already Carries

Description

One log for every edit verb, appended to rather than replaced, and that is what makes a chain of edits produce one honest commit message. With a verb owning its own attribute, ⁠update |> remove |> write⁠ commits a message naming the repoints and silent about the removal – and a destructive edit is the one a ⁠git log⁠ reader most wants named. So an entry says which action it records, and a third editing verb costs an action value rather than a new attribute.

Usage

.datom_append_edits(x, rows)

Arguments

x

The edited object.

rows

A data frame of new entries, carrying .datom_edit_log_fields().

Details

The log is an attribute rather than a field, following the link's carried member record: the write reads tags and members off a set and nothing else, so an attribute cannot reach the payload by construction.

Repointing a member and then removing it leaves both entries. That is an honest history of the edits and slightly odd in a commit message; collapsing them would mean one verb reasoning about the other's rows.

Three actions are written: repoint by datom_update_members(), remove by datom_remove_members(), and add by datom_add_member().

Value

x, with the log extended.


Build the Relative Key for an Artifact's Current-State Metadata

Description

Build the Relative Key for an Artifact's Current-State Metadata

Usage

.datom_artifact_meta_key(name, which = c("metadata", "version_history"))

Arguments

name

Artifact name (validated).

which

"metadata" for metadata.json (current state) or "version_history" for version_history.json (the version index).

Value

Character relative key, e.g. "dm/.metadata/metadata.json".


Build the Relative Key for an Artifact's Payload

Description

The stored data object for an artifact: parquet for a table, JSON for a set. This is the single place that decision is made.

Usage

.datom_artifact_payload_key(name, sha, kind = c("table", "set"))

Arguments

name

Artifact name (validated).

sha

Content hash addressing the payload – data_sha (validated as 6-64 hex, since it is spliced into a storage key).

kind

"table" (parquet payload) or "set" (JSON payload).

Value

Character relative key, e.g. "dm/9f2a....parquet".


Build the Relative Key for a Versioned Metadata Snapshot

Description

Note this is a different directory from the payload key: the snapshot lives under ⁠.metadata/⁠ and is addressed by metadata_sha (the version), whereas the payload sits beside it addressed by data_sha (the content). Both end in .json for a set, which is exactly why they are easy to confuse.

Usage

.datom_artifact_snapshot_key(name, metadata_sha)

Arguments

name

Artifact name (validated).

metadata_sha

The version (validated as 6-64 hex).

Value

Character relative key, e.g. "dm/.metadata/c3d4....json".


Select the Artifacts of One Kind

Description

The one place the artifact list is filtered by kind. Four counters need it – two in the manifest's stored summary block, plus the numbers datom_summary() and datom_status() count for themselves – and a predicate written out at each of them is a predicate that can differ at one of them.

Usage

.datom_artifacts_of_kind(artifacts, kind)

Arguments

artifacts

The manifest's artifacts list, or NULL.

kind

"table" or "set".

Details

An entry that is not a named list is skipped rather than dereferenced. The upgrade step deliberately passes such an entry through untouched, because it has no shape to convert; without the check here, that preserved entry reaches entry$kind and aborts with "$ operator is invalid for atomic vectors". datom_status() is the one that must not do that: it exists to describe a connection when the manifest cannot be trusted, and this count sits outside the error handling that gives it that tolerance. A hand-edited manifest is exactly the document most likely to reach it.

Skipping is not the same as tolerating a missing kind, which stays deliberately uncounted: a typed entry with no type means the conversion was skipped, and a visibly wrong count is the intended signal for that.

Value

The entries of that kind, names preserved.


Assign One Leaf Into a Nested List, Creating Branches on the Way

Description

Recursive rather than iterative because the depth is the length of the axis vector, and tree[[c("a", "b")]] <- v fails when the intermediate list does not exist yet.

Usage

.datom_assign_leaf(tree, path, value)

Arguments

tree

The list to assign into.

path

A character vector of branch names, innermost last.

value

The leaf value.

Value

tree, with the leaf assigned.


Human Label for a Connection's Storage Backend

Description

The word to put in a user-facing message for the backend a connection uses: "S3", "local", or the raw backend name for anything not in the table.

Usage

.datom_backend_label(conn)

Arguments

conn

A datom_conn object, or anything carrying a backend field.

Details

One table, not five. This mapping was written out at four call sites and was about to be written at a fifth, which is the shape of the defect the artifact-kind predicate produced: the same rule copied per site, until one copy lost a term. There is nothing to disagree about here yet, which is exactly when it is cheap to make disagreement impossible.

Value

A single string.


Build a datom_conn from Store Components

Description

Backend-aware helper that creates the appropriate client (S3 client or NULL) and assembles a datom_conn. Used by datom_init_repo(), .datom_get_conn_developer(), and .datom_get_conn_reader().

Usage

.datom_build_init_conn(
  project_name,
  data_store,
  path,
  role,
  endpoint = NULL,
  gov_store = NULL,
  gov_local_path = NULL,
  data_repo_url = NULL,
  github_pat = NULL,
  github_api_url = NULL
)

Arguments

project_name

Project name string.

data_store

A store component (datom_store_s3 or datom_store_local).

path

Local repo path (NULL for readers).

role

One of "developer" or "reader".

endpoint

Optional S3 endpoint URL.

gov_store

A store component for governance (can be NULL).

Value

A datom_conn object.


Build Metadata Object

Description

Constructs the metadata list for a table write, including auto-computed fields (data_sha, dimensions, colnames, timestamp, datom_version) and any user-supplied custom metadata.

Usage

.datom_build_metadata(
  data,
  data_sha,
  custom = NULL,
  table_type = "derived",
  size_bytes = NULL,
  parents = NULL,
  source_lineage = NULL,
  original_file_sha = NULL,
  original_format = NULL,
  project = NULL
)

Arguments

data

Data frame being written.

data_sha

datom-cv1 canonical content hash of the data.

custom

Optional named list of user-supplied custom metadata.

table_type

"derived" (default, from datom_write) or "imported" (from datom_sync).

size_bytes

Size of the parquet file in bytes. NULL if not yet computed.

parents

Lineage list of parent entries (each with source, table, version), or NULL if no lineage recorded.

source_lineage

Pre-computed transitive source list (each entry with project, table, version_sha), or NULL.

original_file_sha

SHA-256 of the source file, for imported tables. Included in the metadata only when non-NULL; the derived path omits it from the object entirely (not present-with-NULL).

original_format

Extension of the source file ("csv", "parquet", ...), for imported tables. Recorded on the same only-when-non-NULL terms as original_file_sha, and for one reason: it was previously written onto the manifest row and nowhere else, which made it the single field a reconstructed index had to drop. It is not part of the version identity – see .datom_metadata_excluded_fields.

project

The name of the project whose namespace this artifact is being written into, from the writing repo's own .datom/project.yaml. Recorded on the only-when-non-NULL terms original_file_sha uses, and last in the signature to match the order the other optional fields were added in. Note what that does not buy: every existing caller passes data and data_sha positionally and everything else by name, so an argument inserted higher up would shift nothing today – it is a convention here, not a guard.

Why the writer records it at all: a reader connection's project_name is a string the caller passed to datom_get_conn() and nothing compares it against the repo, so anything derived from that label is unverified. A write always has a clone, so the name written here is the repo's own declaration – which is what later lets datom_member() and datom_parent() cite a project without trusting a label.

Value

Named list suitable for writing as metadata.json. Always carries kind = "table" (which artifact kind the document describes), schema_version (the format the document is written in) and hash_algo = "datom-cv1", and declares parquet_sha (left NULL here and populated by datom_write() after change detection, since the stored- object hash is not knowable until then; it is excluded from metadata_sha so this deferred assignment is safe).


Build Full S3 URI

Description

Convenience function that combines bucket and key into an S3 URI.

Usage

.datom_build_s3_uri(bucket, key)

Arguments

bucket

S3 bucket name.

key

S3 object key (from .datom_build_storage_key()).

Value

Character string S3 URI.


Build the Metadata Document for a Set Write

Description

A set's metadata.json is a collapsed version of a table's: schema_version, kind, data_sha, hash_algo, document_sha, project, created_at, datom_version, and no more. Everything a table carries that describes a rectangle (nrow, ncol, colnames), the provenance axis (table_type, parents, source_lineage), the stored-parquet facts (parquet_sha, size_bytes) and the user-metadata channel (custom) are all omitted, not nulled – a set's members and its user metadata both live in the payload as tags, and no counter reads a set's byte size.

Usage

.datom_build_set_metadata(payload, document_sha = NULL, project = NULL)

Arguments

payload

The set payload: a list with members (an unnamed list of member records) and optional set-level tags. Must already be tidied and validated – this builder hashes what it is given.

document_sha

SHA-256 of the stored payload bytes, or NULL. Declared either way, mirroring how .datom_build_metadata() declares parquet_sha for datom_write() to populate: the byte hash is not knowable until the payload has been serialized, and it is excluded from metadata_sha, so the deferred assignment cannot move a version. Nothing computes one until the set write path exists, so today it arrives NULL from every caller.

The write path must populate it before writing the document. jsonlite does not omit a NULL element – it writes {}, which reads back as an empty list rather than an absent key. parquet_sha never hits this because its only two outcomes are a real hash or meta$parquet_sha <- NULL, and assigning NULL removes the element. A field left declared-and-unpopulated through a write would satisfy a names-only field-set check while carrying an empty object, so assert on the written bytes where the field set matters.

project

The name of the project whose namespace this set is being written into, from the writing repo's own .datom/project.yaml. Same field and same only-when-non-NULL treatment as .datom_build_metadata()'s project.

Details

Kept beside .datom_build_metadata() on purpose: the two documents are close enough that a field copied from the wrong one is easy to miss, and two of the values here are exactly that kind of trap.

Value

Named list of exactly the fields a set's metadata.json carries, named in the description above. Deliberately not stated as a count: the count was written down in six places and went stale in all of them the first time a field was added.


Build S3 Object Key

Description

Constructs S3 keys from path components, inserting the ⁠datom/⁠ segment per the storage structure convention.

Usage

.datom_build_storage_key(prefix = NULL, ...)

Arguments

prefix

Optional S3 prefix (e.g., "project-alpha"). NULL if none.

...

Path segments after the ⁠datom/⁠ segment (e.g., table name, file name, ".metadata").

Details

Mapping from arguments to key, for reference:

("proj", "customers", "abc123.parquet")
  -> "proj/datom/customers/abc123.parquet"

("proj", "customers", ".metadata", "metadata.json")
  -> "proj/datom/customers/.metadata/metadata.json"

("proj", ".metadata", "dispatch.json")
  -> "proj/datom/.metadata/dispatch.json"

(NULL, "customers", "abc123.parquet")
  -> "datom/customers/abc123.parquet"

Value

Character string S3 key.


Compute the datom-cv1 Canonical Content Hash

Description

The I/O-free identity engine for datom-cv1. Computes data_sha from the in-memory logical values only – no parquet write, no CSV, no temp files, no as.data.frame() or coercion, and never invokes arrow. Columns are read via data[[i]] / names(data) and dimensions via nrow() / ncol(), so two frames with equal values hash identically regardless of container class (tibble vs data.frame vs grouped_df), row names, or arrow version.

Usage

.datom_canonical_hash(data)

Arguments

data

A data frame with at least one row and one column.

Details

Before encoding, every column is scanned through .datom_hash_recourse(); if any are unsupported the function aborts once, listing every offender with its class and canonical recourse. This fires during data_sha computation (step 1 of datom_write()), before any git or storage mutation, so a refusal leaves no partial state.

The final hash is sha256( "datom-cv1" || f64le(nrow) || f64le(ncol) || concat(col_digest_hex...) ).

The per-column digests are an intermediate only and are never returned or persisted. A per-column digest lets anyone holding metadata confirm a guess about one column's values, and metadata is meant to describe a table's shape without revealing its values.

Value

A list with data_sha (character).


Compute the datom-sv1 Canonical Set-Content Hash

Description

The identity engine for a set artifact: data_sha for the payload's semantic content. data_sha = h(0x06 || utf8("datom-sv1") || set(payload)).

Usage

.datom_canonical_set_hash(payload)

Arguments

payload

The set payload: a list with members (an unnamed list of member records, each an id map plus optional tags map) and optional set-level tags.

Details

The hash domain is the parsed-JSON data model, not the in-memory R object, and the write path – which necessarily starts from an in-memory object – agrees with it by construction rather than through a serialize-and-reparse pass. Each way a JSON round trip could mutate a value is unrepresentable instead of handled: there are no numbers (so integer-versus-double cannot arise), NA aborts and absence is omission (so neither the string "NA" nor null can appear), and a single string hashes equal to a one-element array (so scalar-versus-length-1 is not a question). One structural condition remains: the payload must be parsed with simplifyVector = FALSE, so members[] stays a list of records instead of collapsing into a data frame. Both storage backends already do that.

No I/O, no serializer, and no dependency that carries versioned data – which is what makes the hash stable across R versions, jsonlite versions, platforms, and architectures.

Value

A 64-character SHA-256 hex string.


Carry Unrecognised Top-Level Fields Onto a Rebuilt Document

Description

Copies onto rebuilt every top-level field of prior whose name is not in known, so a field this build cannot place survives being rewritten.

Usage

.datom_carry_unknown_fields(rebuilt, prior, known)

Arguments

rebuilt

The document this build assembled, a named list.

prior

The document that was already on disk, a named list, or NULL / anything unparsed when there was none – in which case there is nothing to carry and rebuilt is returned unchanged.

known

Character vector of field names this build can place.

Details

Only unrecognised fields are carried, and that narrowness is the design. A field datom knows about keeps exactly the behaviour it has today, including disappearing when this write does not set it. original_format is the case that makes the difference concrete: a table first imported from a CSV and later written straight from a data frame has no format to declare, and the row is meant to stop claiming one. Carrying every absent field forward instead of only the unplaceable ones would leave that claim standing against a version it does not describe – a wrong statement, which is worse than a missing one.

Where rebuilt already has a field, rebuilt wins. That cannot happen for a genuinely unrecognised field, since this build only writes names it knows; stating the precedence costs one term and removes the question.

Top-level only, at each level separately. A field nested inside a value datom does understand – inside custom, or inside the manifest's summary block – is not this function's business: custom is carried whole as one recognised field, and summary is a derived aggregate that is meant to be recomputed.

A field whose value is JSON null gets no special handling. Absence in a datom document is spelled by omitting the key, never by nulling it, so such a field is already off-convention; it is carried, but a null re-serialises as an empty object rather than as null.

Value

rebuilt, with the unrecognised fields of prior appended.


Refuse to Write One Kind of Artifact Over Another

Description

One name means one artifact, whatever its kind, because both kinds store everything under ⁠{name}/⁠ – a set named dm beside a table named dm would write the same dm/.metadata/metadata.json and each would clobber the other.

Usage

.datom_check_artifact_kind(
  current,
  name,
  expected,
  operation = c("write", "read")
)

Arguments

current

The artifact's current metadata document, or NULL when the name is free.

name

Artifact name.

expected

"table" or "set" – the kind the caller's verb handles.

operation

What the caller was about to do: "write" (default, so existing call sites and the messages they assert on are unchanged) or "read".

Details

Checked against the metadata document in storage, not against the manifest. The manifest is a projection and can lag behind a write that got partway through, so it can say a name is free when it is not. The document is also the copy .datom_has_changes() has just read, so the comparison costs no extra round trip – which is why the current document is passed in rather than fetched here.

An absent kind reads as "table": every document written before the field existed describes a table, because sets did not exist. The pairing with a format check is not needed here the way it is in datom_member() – a document from a future datom has already been refused at the write entry, and on the read path .datom_read_metadata() has just checked the same document.

Both directions of the same invariant, in one function. A read that meets the other kind needs different words and a different suggested verb from a write that does, which operation selects – following .datom_check_schema_version(), which took exactly that shape for exactly this reason. A separate read-side twin would let the two directions drift, each passing its own tests, while the rule they enforce is single: one name is one artifact. For the same reason both aborts carry one condition class, so no test can key on one direction alone.

The check is made by each verb after its own .datom_read_metadata() call rather than inside that function, because the two verbs want different answers from it.

Value

Invisibly NULL. Aborts on a kind mismatch.


Validate Data Store Reachability

Description

Checks that the data store at the ref-resolved location is reachable. For S3: HeadBucket. For local: dir_exists. Provides actionable error messages when data is unreachable after migration.

Usage

.datom_check_data_reachable(conn, migrated = FALSE)

Arguments

conn

A datom_conn object (already ref-resolved).

migrated

Logical, whether a migration was detected.

Value

Invisible TRUE on success. Warns on network error (offline use ok).


Refuse a Document Carrying a Field This Build Cannot Place

Description

Compares one document's top-level key names against the names this build can classify, and aborts naming any it cannot. The evidence is in the file: no version comparison, no configuration, no network.

Usage

.datom_check_document_vocabulary(doc, known, source)

Arguments

doc

Parsed document. A non-list, or a list with no names, has no top-level keys to classify and passes through: it is not this check's job to report a malformed document, and the write fails on it moments later on its own terms.

known

Character vector of field names this build can place.

source

Path or key of the document, for the message.

Details

Chosen over a declared version floor as the primary mechanism for one reason – it cannot be forgotten. A floor protects a repo only if somebody remembers to raise it; this fires on the evidence whether or not anyone did anything.

Top-level keys only, and that is a scope rather than a shortcut. It means do not descend into a value – custom holds arbitrary user keys by design and is classified as one recognised field, and the manifest's summary block is a derived aggregate that is rebuilt on every write. It does not mean skip the manifest's artifact entries: an entry is its own document for this purpose and gets checked against its own vocabulary.

The check cannot fire on the upgrade path, and no code guards against that: a newer build's vocabulary is a superset of every older build's, so it can never meet a name it does not know. A directional special case would be dead code protecting an unreachable state.

The accepted cost is that any release adding a field to a datom-owned document forces a fleet-wide writer upgrade, cosmetic additions included. Writes are infrequent, done by few people, and they change content – a false refusal costs one person an install, a miss costs corrupted data.

Value

Invisibly NULL. Aborts on an unclassifiable key.


Check git2r Availability

Description

Aborts with a helpful message if git2r is not installed.

Usage

.datom_check_git2r()

Value

Invisible TRUE if available.


Check Local Branch is Current with Remote

Description

Fetches from the remote and compares local HEAD SHA against the upstream HEAD SHA. If the local branch is behind, aborts with a clear message telling the developer to pull first.

Usage

.datom_check_git_current(path, pat = NULL)

Arguments

path

Repository path.

pat

GitHub personal access token. Passed to .datom_git_credentials(). NULL means unauthenticated.

Details

Does NOT auto-pull - lets the developer decide how to resolve.

If the fetch fails (offline, unreachable remote), this warns and returns TRUE without comparing anything: the cached remote-tracking refs may be arbitrarily stale, so acting on them would abort an offline developer for being behind a remote they cannot reach. The backstop is .datom_git_push(), which pulls and aborts if the push is rejected, so a write cannot land on storage from a stale base.

Value

Invisible TRUE if the local branch is up to date.


Validate Git Remote Reachability

Description

Checks that the data git remote URL is reachable and that credentials work. Called at conn-construction time in .datom_get_conn_developer() alongside .datom_check_data_reachable().

Usage

.datom_check_git_reachable(conn)

Arguments

conn

A datom_conn object. Uses conn$data_repo_url and conn$github_pat.

Details

Failure behaviour:

Value

Invisible TRUE on success. Warns on network error (offline use ok).


Refuse a Repo With No Remote, Before git2r Does It Unhelpfully

Description

.datom_git_push() reads git2r::remotes(repo)[[1L]], which subscripts an empty list on a repo with no remote and fails with R's own out-of-bounds error – a message naming nothing the caller can act on. A data repo is required to have a remote, so this is an edge rather than a scenario, but datom_repo_push() is the first thing somebody points at a half-configured repo.

Usage

.datom_check_git_remote(path, verb)

Arguments

path

Repository path.

verb

Name of the calling verb, for the message.

Details

Deliberately not inside .datom_git_push(): putting it there would change what four existing callers do on a repo they have never met in that state. Both new verbs call it, so there is no second copy.

Value

The remote name.


Connection Requirements Shared by the Two Git-Mutation Verbs

Description

Both verbs need the same three things and nothing else: a real connection, a developer role, and a local clone to operate on.

Usage

.datom_check_git_verb_conn(conn, verb)

Arguments

conn

A datom_conn object.

verb

Name of the calling verb, for the message.

Value

Invisible TRUE.


Check the Caller's Extra Paths Before a Set Write Does Anything

Description

include_paths is the only way a commit datom makes on its own initiative may carry a path datom does not own, and it is allowed only because the caller enumerated it (R14.3). What it buys is that checking out a set version's commit yields the data pointers and the code and environment that produced them. So every refusal here is a refusal to produce a commit that would claim more than it holds.

Usage

.datom_check_include_paths(conn, name, include_paths)

Arguments

conn

A datom_conn object with a local path.

name

The set being written, as resolved by .datom_check_set_write_gates(). Named rather than discovered because a first write has no directory to discover.

include_paths

The caller's character vector, or NULL.

Details

Four refusals, in this order, each with its own condition class:

  1. Not a path inside the clone. An absolute path, or one climbing out through .., refused lexically before the filesystem is touched. fs::path() joins an absolute second argument onto the clone path rather than replacing it, so ⁠/etc/passwd⁠ would otherwise be reported as a missing path inside the repo – a correct refusal whose message names the wrong thing.

  2. A datom-owned path. ⁠.datom/⁠, the set being written, and any artifact directory already in the clone. The write stages those itself, so listing one is either a misunderstanding or an attempt to hand-place a datom document into a commit through a caller's argument.

  3. A path that does not exist. An error, never a skip: a joint commit is deterministic or it is refused.

  4. A path git is ignoring. git2r::add() on a gitignored path raises nothing and stages nothing, and .datom_git_commit() cannot notice, because it objects only when the staging area ends up empty and datom's own files are always in it. The commit would therefore succeed while omitting exactly the file the caller named, and the set version would claim a joint commit it does not have. Refused rather than dropped in silence (decided 2026-09-18).

This runs before the first hash and the first local write, the same placement as the two gates above, so a refused joint commit leaves nothing behind. One consequence, stated so nobody later softens it: change detection needs the hashes, so a bad path is an error even when the set is unchanged. The refusal wins over the no-op.

Value

Absolute paths in the clone, or NULL. Absolute, because .datom_commit_and_mirror() relativises what it is given against conn$path, and fs::path_rel() on an already-relative path resolves it against the working directory instead – which aborts with "files do not exist" pointing somewhere the caller never named.


Check Whether a Storage Namespace is Free

Description

Checks for the existence of .metadata/manifest.json in the target namespace. If found, the namespace is occupied by an existing datom project and this aborts with datom_namespace_occupied, naming the occupying project when it can be read.

Usage

.datom_check_namespace_free(conn, overridable = TRUE)

Arguments

conn

A datom_conn object (typically a temporary conn built by datom_init_repo() before the repo is fully initialised).

overridable

Whether the caller honours .force as a way past an occupied namespace. TRUE (the default) adds that route to the refusal's recourse; FALSE says the override does not apply and why.

It exists because this function cannot know its caller's policy, which is the same reason the backend label is an argument's worth of work rather than a constant. A product repo is checked with no opt-out, so a static "pass .force = TRUE to override" bullet sent exactly those users into a flag that changes nothing – a message routing somebody in a circle, which is the failure this function's own backend-neutral wording was fixed for one commit earlier.

Details

Checks for the object first (cheap) and only reads the manifest when the namespace is occupied, to extract the project name for the error message.

A store this connection cannot reach means unknown, and unknown fails closed. It used to warn and continue, which was not a deferral of the check but a silent removal of it: datom_init_repo() went on to push the git repo and then aborted at the manifest upload, and the recovery it pointed at performs no occupancy check of any kind. So the tolerance never produced a working offline init – storage is required to finish one – and its only reachable effect was getting past this check, with the outcome being a manifest written over another project's. Refusing here instead names the real problem at the moment it is known, rather than surfacing later as an unrelated upload failure.

There is no .force advice in that refusal, deliberately: .force skips this check but not the manifest upload, so it cannot rescue an init without storage either. Offering it would be advice that does not work.

The tolerated-failure detection lives here, around the one call that touches storage, rather than in a handler wrapping this whole function – which is what the caller used to do. That shape had two defects worth not reintroducing: it recognised the occupied refusal by matching its message text, so rewording the message would have quietly downgraded a refusal to a warning; and it swallowed anything it could not recognise, so any abort added to this function later would have been downgraded too, with nothing failing to say so.

The condition classes are what callers dispatch on – never the message.

Value

Invisible TRUE when the namespace is free. Aborts with class datom_namespace_occupied when it is occupied, or datom_namespace_unverified when the store could not be reached.


Refuse a Project Name That Cannot Be Cited

Description

The project field a metadata builder records exists to be quoted back by whoever cites the artifact, so a missing value or an empty string there is worse than no field at all: it reads as a project called nothing. Checked in the builders rather than at the call sites, because both builders take the value from the same place and a third caller will eventually appear.

Usage

.datom_check_project_field(project)

Arguments

project

The value passed to a builder's project argument.

Details

NULL passes. It means "not recorded", which is what every document written before the field existed looks like, and what a direct builder call in a test that is not about this field looks like.

Value

Invisibly TRUE.


Check project.yaml's Declared Format

Description

The same reader-side check every other datom-owned document gets, pinned to the config file's own ceiling (.datom_project_schema) rather than the repo-wide one. Absent means v1, which is every repo written so far, so no existing repo changes behaviour.

Usage

.datom_check_project_schema(cfg, source, operation = c("read", "write"))

Arguments

cfg

Parsed project.yaml (a named list).

source

Path of the config file, named in the refusal message.

operation

What the caller was about to do. "read" (the default) is what connection construction passes – opening a connection is neither a read nor a write, and "this build cannot read" is literally true of the config file. "write" is for the set-write gates, which read this file to decide whether a write may proceed.

Details

Why the file needs a declared format at all. project.yaml carries fields a writer must obey, not merely fields it may read: min_writer_version already, and mode / set for a product repo. A build that does not recognise such a field walks past it and acts as though the repo had never asked for anything – so the file needs a way to say "this repo needs a newer datom", and a number is that way.

A number here, a vocabulary check there, and the two are not interchangeable. The vocabulary check that guards the manifest and per-artifact metadata draws its power from those documents being machine-written: an unrecognised key there is evidence a newer datom wrote it. project.yaml is hand-edited – storage migrations, prefixes, descriptions, private notes – so an unrecognised key is as likely a typo, and refusing on one would block every write in the repo until somebody found it. Never point the vocabulary check at this file; an unrecognised key here stays tolerated, and there is a test that says so.

This wrapper exists so the pairing of file and ceiling cannot be forgotten. A bare ⁠supported =⁠ argument at each call site is the same shape as the artifact-kind predicate that was written out at four sites and lost a tolerance at one of them. Callers pass the parsed config; the ceiling is not theirs to choose.

Value

Invisible resolved version as an integer. Aborts otherwise.


Check ref.json Matches Connection (Write-Time Guard)

Description

Re-resolves ref.json from the governance store and compares against the current connection's data location. Errors if they disagree, preventing writes to the wrong location after a migration.

Usage

.datom_check_ref_current(conn)

Arguments

conn

A datom_conn object.

Value

Invisible TRUE if current, or skips silently if no governance fields.


Check a Document's Declared Schema Version

Description

Reader-side compatibility check for one metadata or manifest document. Called wherever such a document enters datom from storage or from the local clone, so that a repo written by a newer datom fails with an actionable message instead of degrading silently – an older reader would otherwise find none of the fields it expects and report an empty repo.

Usage

.datom_check_schema_version(
  meta,
  source,
  operation = c("read", "write"),
  supported = .datom_supported_schema
)

Arguments

meta

Parsed document (a named list). A non-list or NULL is treated as carrying no schema_version, i.e. v1.

source

Path or key of the document, used in the message so the user knows which file is too new.

operation

What the caller was about to do – "read" (default) or "write". It only selects a word in the refusal message. An argument with a default rather than a required one, so that every existing call site and the message text they assert on are unchanged: without it a refused write said the format was one "this build cannot read", which is the wrong verb for a write that was stopped at the door.

supported

Highest version this caller can interpret, defaulting to the repo-wide .datom_supported_schema. It feeds the comparison and the message, so a refusal never says "supports up to v2" while refusing a v2 file.

The rule that predicts an override, so a future caller can derive it rather than remember it: a document datom writes takes the repo-wide ceiling; a document that outlives the build that created it and is then edited by hand gets its own. The shared number holds while every document on it is machine-written by one build in one operation. .datom/project.yaml is not – it is stamped once at init and hand-edited afterwards – so it carries .datom_project_schema and is checked through .datom_check_project_schema(), which is where that pairing lives.

Details

The check is deliberately asymmetric:

A present-but-unusable value (a string, a fraction, NA, a vector) aborts as a corrupt document rather than being coerced. Coercion here would compare garbage against the supported version and could silently read as "supported"; and in R a comparison against NA propagates into ⁠if()⁠ as an opaque "missing value where TRUE/FALSE needed" error rather than anything a user can act on.

Both aborts carry a condition class so every call site is provably the same failure: datom_schema_unsupported for a too-new document, datom_schema_invalid for an unusable value.

Value

Invisible resolved schema version as an integer. Aborts otherwise.


Refuse a Set Whose Outputs Were Built From Versions It Does Not Pin

Description

A table records the parent versions it was derived from. When a set lists that table and one of its parents, the two claims can disagree: the set says "input lb is version A" while the output says "I was built from lb version B". A citation of that set would then describe a product that was never built. This check stops the write and names both versions.

Usage

.datom_check_set_parents(conn, name, members)

Arguments

conn

The set's own developer connection.

name

The set's name, for the message.

members

The validated, ordered member list.

Details

What is checked. Every member in the set's own project – all tables, since a product repo holds one set and a set listing itself is refused first. For each parent its snapshot records, the members naming that parent's project and table are found. None -> not checked (the parent is not part of the product). Some, and one of them at the parent's version -> fine, which is what keeps a live table beside a frozen baseline legal. Some, and none at that version -> a mismatch. All mismatches are collected and reported together.

Members of other projects are not read: the write holds a connection to its own project only.

An unreadable snapshot stops the write rather than being skipped. A version that does not exist means the set would point at data nobody can fetch, and storage that cannot be reached would fail the write anyway – only later, after the local files are written. Metadata only: one small JSON per member, no parquet, so a readable snapshot proves the version was recorded, not that its data is still in storage (that is datom_validate()'s job).

Called after the payload has been validated and before the first hash, so a malformed member gets the validator's message and a refusal leaves nothing behind.

Value

Invisibly TRUE.


Refuse a Payload Only a Whole-Payload View Can Judge

Description

Three refusals that .datom_validate_members() deliberately cannot make, because each needs the whole payload rather than one member:

Usage

.datom_check_set_payload(payload, name, project)

Arguments

payload

A tidied, validated, ordered payload.

name

The set's own name.

project

The set's own project.

Details

The same project and name at two different versions is legal and must stay legal: a product carrying a current table beside a locked baseline is atypical and entirely sensible. So the duplicate check keys on the full id, never on project + name – which is the tightening that looks natural and would break that use silently.

Self-reference is a nonsense check, not cycle detection. Cycles are structurally impossible: a member pins a version that already exists, so a set cannot reference anything containing it. Nothing here may grow into a visited set or a depth limit.

Value

Invisibly TRUE.


Refuse a Set That Belongs to Another Project

Description

A datom_set records the project it belongs to: the one it was read from, or the connection it was assembled on. The write stamps conn$project_name into the stored document, so a set from another project written here would be silently re-homed. The name gate does not catch that on its own, because two product repos may declare the same set name (two studies, each with a set called adam).

Usage

.datom_check_set_project(conn, name, set_project)

Arguments

conn

The product repo's developer connection.

name

The resolved set name, for the message.

set_project

The project the set carries, or NULL when it carries none (a plain member list, or a set whose project was never recorded).

Value

Invisibly NULL; aborts with class datom_set_project_mismatch.


Refuse a Set Write the Repo Has Not Declared

Description

Checks that read the clone's .datom/project.yaml directly, all of them before anything is hashed or written:

Usage

.datom_check_set_write_gates(conn, name = NULL)

Arguments

conn

A datom_conn object with a local path.

name

The set name the caller supplied, or NULL to take the repo's declared one.

Details

  1. The config's declared format must be one this build can read (.datom_check_project_schema()). It runs first because the two checks below read fields out of this file: a future format that renamed or moved ⁠set:⁠ would make check 3 report "this repo declares mode: product but names no set" and send the user to hand-edit a file that is already correct. An actionable-looking message that is wrong is worse than no answer. The connection-time gate does not make this one redundant – the file can be hand-edited or pulled between opening a connection and writing through it, which is the same reason the forward-compatibility door is re-run after a route's own pull.

  2. The repo must declare mode: product. A set written into a repo that does not is a set with no declared owner, which defeats the one below it too.

  3. The set's name must be the one the repo declares under ⁠set:⁠. This is what makes "one repo = one set = one product" true rather than aspirational, and it is the precondition the self-reference refusal depends on – that refusal needs the set's own identity, and this is where it is established.

Read from project.yaml, not from the connection. Only min_writer_version rides on a datom_conn; mode and set are read here, at the one place that needs them, so nothing has to be threaded through connection construction for two fields with one consumer. If a later caller wants them on the conn it is a move with one call site to update, rather than a decision to reopen.

⁠datom_init_repo(mode = "product", set = <name>)⁠ is what declares both fields, so the supported route into this check is a repo created that way. It shipped inert one release earlier, on purpose: a build that can notice the declaration has to exist before anything writes it, or the declaration reaches installs that walk straight past it. Hand-editing the file still works and some fixtures do it, which is also what a repo created before that argument existed needs.

Value

The resolved set name.


Everything a Write Must Clear Before It Starts

Description

The one entry sequence for every write route. Ordered, and the order matters:

Usage

.datom_check_write_entry(conn, artifact = NULL)

Arguments

conn

A datom_conn object.

artifact

Name of the single artifact this write touches, or NULL to check every artifact present in the clone.

Details

  1. The floor – refuse if this repo has declared this build too old.

  2. The manifest, read through the one shared reader, which checks the declared format and then converts an older document in memory. Refusing a format from the future has to happen before the conversion, because there is no conversion step for a version this build has never heard of.

  3. The shape the conversion reached. If the artifact list is still absent after the chain has run, this document belongs to a lineage this build cannot produce – refuse rather than overwrite it in an older shape.

  4. The vocabulary, on the manifest's top level, on each of its artifact entries, and on each per-artifact metadata document this write will touch.

Why step 3 is worded as "still absent after the chain" and not "absent". A current build meeting a pre-rename repo finds no artifact list either, and a rule that refused on that would deadlock the very upgrade it exists to protect: no repo could ever move forward. The discriminator is not the key, it is whether the chain can reach the shape. Note the deliberate asymmetry with the reader, which will one day rebuild on this same condition – same evidence, opposite response, because reads limp and writes stop.

Which per-artifact documents get checked depends on the route, which is why artifact exists. A table write or a metadata-only sync touches one artifact; the mirror-everything route touches all of them, and it is the route with no artifact name in its arguments at all. Both cases enumerate through .datom_clone_artifact_names(), the same helper the mirror route itself uses, so the door cannot end up inspecting a different set than the one that gets written.

Callable more than once, and one route calls it twice. Nothing here mutates anything, so re-running it is free. The metadata-only route pulls from the remote as its first act – after this check has already read the clone – so a collaborator's newer document can arrive in that pull; that route runs the sequence again afterwards. See .datom_sync_metadata().

Value

Invisibly NULL. Aborts on any refusal.


Refuse a Write This Repo Has Declared Too Old

Description

A repo may state the lowest version of datom it accepts writes from. The field is optional and lives in project.yaml; absent means no limit, so no repo written so far changes behaviour.

Usage

.datom_check_writer_floor(conn)

Arguments

conn

A datom_conn object.

Details

It exists for the two cases the vocabulary check structurally cannot see, because neither introduces a new field name: a change in what an existing field means, and a block for a reason that is not about format at all ("0.1.4 wrote bad hashes, do not let it write here"). A version number is the right currency for both – the schema number cannot carry them, since it does not move for a change that is reader-safe, and a package version directly answers the question a refusal raises.

The reading half ships even though nothing sets the field yet, and that ordering is the whole point: a build that does not look for the field can never be bound by it. This is exactly why no released datom can be stopped from writing – the looking has to be inside the build being stopped. Setting the field is a separate, later mechanism, and it owns the guard that whoever raises a floor must already satisfy it.

A value that will not parse as a version aborts rather than being ignored. Treating a malformed floor as no floor would turn a typo in a policy field into a silently disabled policy.

Value

Invisibly NULL. Aborts when the running build is older than the declared floor.


Human-Readable Class Label for a Column

Description

Renders the label shown for a column in the all-offenders abort bullets and in the datom_check_hashable() report: the collapsed class(x) string for an explicitly-classed column, or typeof(x) for an unclassed one (so a list column reads list, a complex column complex, and a units column units).

Usage

.datom_class_label(x)

Arguments

x

A single column (vector) from a data frame.

Value

A single character string.


The Artifacts Present in the Local Clone

Description

Enumerates the artifact directories in the git checkout by the one signal that identifies them: a directory holding a metadata.json. Deliberately independent of the manifest, so it still answers correctly when the manifest is the document under suspicion.

Usage

.datom_clone_artifact_names(conn)

Arguments

conn

A datom_conn object with a local path.

Details

Two callers, and they must agree. The data-side metadata sync mirrors exactly these artifacts to storage, and the write-entry check inspects exactly the documents that route is about to write – so discovering them twice, in two spellings, is how the door ends up checking a different set than the one that gets written.

The directory filter is the pre-existing one: dotfiles out, plus the fixed list of non-artifact directories a joint repo carries (⁠R/⁠, ⁠tests/⁠, ⁠renv/⁠, and so on). It is a convenience rather than the discriminator – metadata.json is what actually decides – which is why a foreign directory not on the list is tolerated rather than misread (R14.2).

Value

Character vector of artifact names, possibly empty.


Compute a Single Column's datom-cv1 Digest

Description

Encodes one column to its per-column SHA-256 hex digest for datom-cv1, as sha256( utf8(tag) || utf8(colname) || 0x00 || payload ). The tag is the kind returned by .datom_column_kind() and the payload is produced by the shared encoders. Labelled columns strip their class and attributes and re-dispatch on the bare underlying vector, so value labels never enter identity.

Usage

.datom_col_digest(name, x)

Arguments

name

Column name (used verbatim, UTF-8, in the digest input).

x

The column vector.

Value

A 64-character SHA-256 hex string.


Classify a Column for Canonical Hashing

Description

The single supported-type classifier underneath the datom-cv1 hash. It returns the dispatch kind for a hashable column or NULL for an unsupported one. Both the all-offenders gate (.datom_hash_recourse()) and the per-column encoder in .datom_canonical_hash() consume this one function, so a column the gate accepts can never be one the encoder cannot encode.

Usage

.datom_column_kind(x)

Arguments

x

A single column (vector) from a data frame.

Details

Dispatch is evaluated in a fixed order (reordering can silently change hashes): bit64::integer64, factor, Date (incl. data.table::IDate), POSIXct, difftime/hms, data.table::ITime, then haven_labelled/labelled/labelled_spss (stripped to their underlying type and re-classified), then any other explicitly-classed column is refused, then unclassed atomics (logical/integer/double as "num", character as "chr"), and finally any other type is refused.

Detection uses inherits() / typeof() / is.object() class-string matching only – it adds no new package dependency (bit64, data.table, haven are recognised by their class strings, not by being loaded).

Value

One of the kind tags "i64", "chr", "date", "time", "drtn", "num" for a supported column, or NULL when unsupported.


Commit, Push, Then Mirror to Storage

Description

The tail of every artifact write, in the one order that is allowed: local files are already on disk, this commits and pushes them, and only then does it touch storage. Git push is the serialization point – a clone that is behind fails to push before it can upload anything, which is what makes the reuse decisions in .datom_resolve_parquet_sha() / .datom_resolve_document_sha() safe against a concurrent writer. Nothing may reorder these two halves.

Usage

.datom_commit_and_mirror(
  conn,
  name,
  meta,
  metadata_sha,
  git_paths,
  message,
  upload = NULL
)

Arguments

conn

A datom_conn object (developer, with a local path).

name

Artifact name.

meta

The metadata document to mirror.

metadata_sha

The version being written.

git_paths

Absolute paths of the files this write produced in the clone. .datom/manifest.json is added here rather than by each caller, since every write updates it.

message

Commit message.

upload

Optional list(path =, key =) naming a payload object to upload after the push – the freshly serialized parquet for a table, the payload JSON for a set. NULL when the object is already stored and must not be rewritten.

Details

Extracted when the set write arrived, and the extraction is the point rather than tidiness: this sequence was previously inline in datom_write(), so a second write verb had to either call it or grow a parallel copy – and a second copy of "git must succeed before storage is touched" is a second place for that rule to be broken by a change that only looks at one of them.

Value

The commit SHA.


Compute the datom-cv1 Content Hash of a Data Frame

Description

Thin wrapper over .datom_canonical_hash() returning only the scalar data_sha. Preserves the scalar-string contract for callers that need just the content hash (for example the datom_sync() self-lineage entry). Row and column order are significant; there is no sort option.

Usage

.datom_compute_data_sha(data)

Arguments

data

Data frame to hash.

Value

Character SHA-256 data_sha.


Compute SHA-256 of Metadata (the datom Version)

Description

Hashes the fields named in .datom_metadata_identity_fields and ignores every other key in the document. See that constant for the field-by-field classification, for why selection is an allowlist rather than an exclusion list, and for the obligation that comes with adding a field to a builder.

Usage

.datom_compute_metadata_sha(metadata)

Arguments

metadata

Named list of metadata fields. An unrecognised field is ignored, not refused – that is what lets this build read a document written by a newer datom without reporting a change on content that did not move. Refusing such a document is a separate, write-side concern.

Details

Hashes a JSON canonical form rather than the R object directly. This ensures that metadata read back from JSON (e.g., from S3) produces the same SHA as metadata built in-memory, despite R type differences (integer vs double, character vector vs list) introduced by JSON round-tripping.

Value

Character SHA-256 hash.


Compute SHA-256 of an Input File's Raw Bytes

Description

Answers "have this input artifact's bytes changed?". This is the original_file_sha of the three-SHA identity model – distinct from data_sha (canonical logical content) and parquet_sha (stored bytes).

Usage

.datom_compute_original_file_sha(path)

Arguments

path

Path to file.

Value

Character SHA-256 hash.


Scope-Selecting Connection Accessor

Description

Returns the connection shaped for either the data or governance store. The storage dispatch layer (⁠.datom_storage_*⁠) reads conn$root, conn$prefix, and conn$client; this accessor swaps those fields when callers need to operate on the governance store.

Usage

.datom_conn_for(conn, scope = c("data", "gov"))

Arguments

conn

A datom_conn object.

scope

Either "data" (default; returns conn unchanged) or "gov" (returns a sub-conn with governance fields swapped in).

Details

Single source of truth for "which store am I talking to right now?" – replaces ad-hoc conn$gov_client peeking and the prior .datom_gov_conn() helper.

Value

A datom_conn object scoped to the requested store.


Copy a Single Storage Object Between Two Connections

Description

Dispatches on the (from_backend, to_backend) pair. For local->local uses fs::file_copy; all other combos transfer raw bytes.

Usage

.datom_copy_one(from_conn, to_conn, rel_key)

Arguments

from_conn

Source datom_conn.

to_conn

Destination datom_conn.

rel_key

Relative storage key (after ⁠{prefix}/datom/⁠).

Value

Named list with key (character) and bytes (numeric).


Create a GitHub Repository

Description

Creates a new GitHub repository via the REST API. Handles both org and personal repos.

Usage

.datom_create_github_repo(
  repo_name,
  pat,
  org = NULL,
  private = TRUE,
  api_url = "https://api.github.com"
)

Arguments

repo_name

Repository name.

pat

GitHub personal access token.

org

GitHub organization. NULL for personal repos.

private

Whether the repo should be private (default TRUE).

api_url

GitHub API base URL (default "https://api.github.com").

Details

Safety guard:

Value

The clone URL of the created/reused repository.


Build governance.json Content

Description

Constructs the governance pointer list that is written to both the local git copy and the data-store mirror.

Usage

.datom_create_governance_json(gov_repo_url, gov_store, attached_at = NULL)

Arguments

gov_repo_url

HTTPS clone URL of the governance git repository.

gov_store

A datom_store_s3 or datom_store_local component representing the governance storage (location + credentials). Only the location fields are persisted; credentials are discarded.

attached_at

Optional ISO 8601 UTC timestamp string. Defaults to the current system time.

Value

Named list suitable for serialisation to JSON.


Create Initial ref.json Content

Description

Builds the initial ref.json structure from the data store component. No previous entries on first creation.

Usage

.datom_create_ref(data_store)

Arguments

data_store

A datom_store_s3 component (the data portion of the store).

Value

A list suitable for JSON serialization.


Each Artifact's Current Version in One Project, in One Read

Description

The manifest carries current_version per artifact, so learning what moved costs one read per project rather than one per member. Only the members that actually move then pay a snapshot read, through datom_member(), which is what keeps the new pointer trustworthy.

Usage

.datom_current_artifact_versions(conn)

Arguments

conn

A connection to the project.

Details

The manifest is read directly rather than through datom_list() for two reasons: datom_list() abbreviates that column to 8 characters by default, which is not a version a member can record, and its abort would name S3 on a local backend.

An unreadable manifest refuses. It is the one answer that cannot be reported: a member whose project could not be read is a member whose state is unknown, which is the same situation as a missing connection.

Value

A named character vector of artifact name to current version, with NA for an entry that records none. Empty when the project has no artifacts.


Every Artifact in One Project, With Its Kind and Current Version, in One Read

Description

The same single manifest read as .datom_current_artifact_versions(), and the same refusal when it cannot be read – that function is built on this one. The set sync preview needs two things the name-to-version vector drops: each entry's kind, which each preview row reports and which drops any kind this build does not know, and the project name the manifest records, which is how a mislabelled source connection is caught before it shows every artifact as new.

Usage

.datom_current_artifacts(conn)

Arguments

conn

A connection to the project.

Value

A list of project_name (the name the manifest records, or NULL when it records none) and artifacts, a data frame of name, kind and current_version, with NA for a field an entry does not record usably.


Which Project an Artifact Belongs To, From the Repo Rather Than a Label

Description

The cascade both pointer constructors use – datom_member() and datom_parent() – to answer "which project is this artifact in" without trusting the connection it was reached through.

Usage

.datom_declared_project(conn, snap, what = "member")

Arguments

conn

The connection the artifact was read through.

snap

The artifact's metadata snapshot, already read and already checked for a format this build understands.

what

What is being declared – "member" or "parent" – used only to word the unverified-fallback warning.

Details

Why a label cannot be trusted. On a connection built from a clone, datom reads project_name out of .datom/project.yaml, so it is the repo's own declaration. On a reader connection it is a string the caller passed to datom_get_conn(): the namespace comes from the store's root and prefix, and nothing compares the label against the repo. Both constructors write the name they settle on into a stored document – a member's id$project is hashed into the set's data_sha and cited afterwards, and a parent's source is part of the declaring table's version – so a label nobody checked would be durable wrong data that no hash and no validator can notice.

Three steps, cheapest and most trustworthy first:

  1. The artifact's own snapshot, which the caller has already read. Free, and it is the writing repo's declaration.

  2. The manifest of the namespace the artifact lives in. One extra read, and only for an artifact written before datom recorded the field – which is every artifact in every existing repo, so this is the common path in this release rather than a rare one. Goes through the gated reader, so a manifest whose format this build cannot read is handled the one way datom handles that anywhere; when the manifest is unusable, that reader can escalate to reconstructing the index from a namespace listing, which is accepted because a repo in that state needs attention regardless.

  3. The connection's label, said out loud to be unverified.

One gap, named rather than guarded. When the manifest has to be reconstructed and the document it replaced recorded no project name, the reconstruction fills that field from the connection (.datom_rebuild_manifest()), so step 2 can hand back the label while looking like the repo's declaration. What is lost there is the warning, not the value: the string is exactly the one step 3 would have returned. Closing it properly means the shared manifest reader reporting whether the document it returned was reconstructed, which is a change to that reader rather than to this cascade.

Value

A single non-empty string.


Delete a GitHub Repository

Description

Deletes a GitHub repository via the REST API. Requires a PAT with the delete_repo scope.

Usage

.datom_delete_github_repo(repo_full, pat, api_url = "https://api.github.com")

Arguments

repo_full

Repository in "owner/repo" form.

pat

GitHub personal access token (must have delete_repo scope).

api_url

GitHub API base URL (default "https://api.github.com").

Value

Invisible TRUE on success; aborts on failure.


Is This Member Already in the Set, and Is It the Same Member?

Description

Answers the question the write answers twice, one step earlier, so a repeat lands on the line that introduced it.

Usage

.datom_draft_member_clash(members, record)

Arguments

members

The set's members so far, with their links stripped.

record

The record about to be added.

Details

Both of the write's rules are here, and they are deliberately different rules. .datom_order_set_members() drops an exact repeat – same id and same tags – silently, because the digest it dedupes on covers tags. The same id with different tags survives that and is then refused by .datom_check_set_payload(), because merging the labels and picking one entry both guess. So an exact repeat is a duplicate to skip, and a same-version disagreement is an error.

The comparison uses the write's own two mechanisms rather than restating them: the project / name / version key the payload check keys on, and the datom-sv1 member digest the dedup keys on. identical() on the two records is the spelling to avoid, and it fails in the direction that refuses working input: the encoder sorts a tag map's keys and encodes each value as a sorted, deduplicated set, so domain = c("a", "b") and c("b", "a") are one member to the write and to the digest, while identical() reads them as a disagreement and aborts.

That is also why nothing needs tidying first. Every spelling the write's tidy step collapses is a spelling the digest is already blind to, so a record can be compared – and stored in the set – exactly as the caller supplied it.

Value

A list with status – "new", "duplicate" or "conflict" – and, for the last two, at: the position of the member already in the set.


Drop Tag Keys Whose Value Is Empty

Description

The one tidy rule this file owns: a key that points at nothing is removed, because a tag with no values is spelled by omitting the key. Note this is about a tag value, not about a member having no tags at all – a member with no tags is the ordinary case and is simply accepted. Nothing is lost here – an empty value states no fact – and leaving it in would be worse than cosmetic: a present key with an empty value hashes differently from an absent key, so the same fact would mint two different data_sha values.

Usage

.datom_drop_empty_tags(tags)

Arguments

tags

A named list, or NULL.

Details

Covers both empty spellings R produces. character(0) is the documented one; NULL is what an absent value looks like when a tag map is composed programmatically (list(domain = f()) where f() returned nothing), and it means exactly the same thing. Note the encoder still refuses a NULL value, and must: there it arrives from a parsed file rather than from a caller, so there is no caller intent to tidy toward.

Full canonicalization – sorting keys, sorting and deduplicating values, unboxing single values, ordering members – is not done here. It belongs to the set write, so that canonical form has exactly one implementation.

Value

tags with empty-valued keys removed; NULL unchanged. A non-list is returned untouched, so the validator reports the type rather than this function failing on it.


The Supplied Connections, Keyed by the Project Each One Claims

Description

One connection or a list of them, because a set legitimately spans projects and access in datom is per project. The key is conn$project_name, which nothing verifies – see this file's header for what confirms the choice afterwards.

Usage

.datom_edit_conns(conn, arg = "conn")

Arguments

conn

A datom_conn, or a list of them.

arg

Argument name for the message.

Details

Two connections claiming one project are refused rather than ordered, since choosing between them would be a guess and the wrong one reads another project's namespace.

Value

A named list of connections.


One Display Line Per Edited Member, Grouped by Project

Description

Grouped by project because that is the axis connections are supplied along, so a surprise in the grouping is a surprise about which connection served what. A removal has no connection behind it, but it keeps the same grouping so one message can hold both kinds of entry.

Usage

.datom_edit_lines(edits, abbreviate = TRUE)

Arguments

edits

The edit log.

abbreviate

Whether to shorten versions to 8 characters (the console) or leave them whole (a commit message, where git is the durable record).

Value

A character vector of lines.


The Columns of the Edit Log, in One Place

Description

The Columns of the Edit Log, in One Place

Usage

.datom_edit_log_fields()

Value

A character vector of column names.


The Member List of a Set, or an Abort Naming What Was Passed

Description

A datom_set however it was made – read back, or assembled with datom_assemble_set() – which is exactly what datom_write_set() takes, so the edit verbs and the write accept the same objects.

Usage

.datom_edit_members(x, arg = "x")

Arguments

x

The value the caller passed.

arg

Argument name for the message.

Value

The member list, possibly empty.


The Zero-Row Shape of a Member Listing

Description

In one place, and carrying every column a populated result carries, so rbind() of an empty listing and a populated one works. datom_list() had exactly this defect twice: a zero-row frame built from zero rows loses its columns, and the failure only shows up when somebody binds two results.

Usage

.datom_empty_member_frame()

Value

A zero-row data frame.


A Set With No Version Yet

Description

What a set is before its first write: a name (or NULL, left for the write to resolve), the project it belongs to, and no members. Spelled list(version = NULL, ...) so the empty fields keep their names, which is the shape datom_get_set() returns for a read set. datom_assemble_set() returns one, and the sync preview starts from one when the repo's set has never been written.

Usage

.datom_empty_set(name, project)

Arguments

name

The set's name, or NULL.

project

The repo's project name.

Value

A datom_set.


Encode a Character Payload for Canonical Hashing

Description

The character encoder for the chr column kind (character and factor columns). Emits a one-byte-per-row NA mask (0x01 where is.na(), 0x00 otherwise) followed by each value re-encoded to UTF-8 via enc2utf8() and NUL-terminated. The leading mask makes NA and the empty string "" distinguishable (both have an empty value section, but NA sets its mask byte). No Unicode normalization is applied, so NFC and NFD forms of the same text encode differently (a documented, benign limitation).

Usage

.datom_encode_character(x)

Arguments

x

A vector coercible to character (character or factor).

Value

A raw vector: length(x) mask bytes followed by the NUL-terminated UTF-8 value bytes.


Encode a Numeric Payload for Canonical Hashing

Description

The single shared numeric encoder used by the num, date, time, and drtn column kinds of datom-cv1. Produces a fixed, platform-independent byte sequence: IEEE-754 doubles written little-endian regardless of host endianness, with three canonicalizations so that logically-equal values encode identically:

Usage

.datom_encode_numeric(x)

Arguments

x

A vector coercible to double (logical, integer, double, or the numeric payload of a Date/POSIXct/difftime column).

Details

No rounding is applied: doubles are encoded bit-exact.

Why the canonical NaN is written as bytes, not assigned as a value. Assigning R's NaN (d[nan_idx] <- NaN) folds NaN payloads but inherits the host's NaN sign bit: R's NaN is ⁠0x7ff8...⁠ on macOS/arm64 and ⁠0xfff8...⁠ on Linux/x86_64, because it comes from a C-level 0.0/0.0. That made data_sha platform-dependent for any table containing a NaN – caught by the CI golden matrix (the macOS job passed, the Linux job did not). Splicing the pinned bytes in directly removes the host from the equation, which is the whole premise of a canonical hash.

Value

A raw vector of 8 * length(x) bytes.


Expand a Member List to One Row Per Member Per Tag Value

Description

The shared expansion both shaping verbs are built on – see point 1 of this file's header for why there is exactly one of these.

Usage

.datom_expand_member_tags(members)

Arguments

members

A non-empty member list.

Details

Carries a .member column holding the member's position, which is what lets datom_structure_members() get back from a row to the record it came from. datom_list_members() drops it, because a position is not a fact about a member.

Value

A data frame of .member, name, project, version, kind, key, value.


Find the One Member a Name Refers To

Description

An ambiguous name aborts and teaches. Two members can legitimately share a name – the same artifact at two versions, for instance a current table beside a locked baseline – so a name is not a key, and answering with the first match would be plausible and wrong. The abort lists the candidates with their versions and tags and names the two ways to narrow: tags, which is the navigation axis people reach for, and version, for exact pinning.

Usage

.datom_find_member(members, name, tags = NULL, version = NULL)

Arguments

members

The set's member list.

name

The name to look up.

tags

Optional label filter.

version

Optional version, or a prefix of one.

Value

One member record.


Forget the Version an Edited Set Was Read As

Description

datom_get_set() fills version and data_sha from the payload it read. Once a member moves, those two describe a payload that no longer exists – and a set exists to be cited, so a stale version is a wrong statement rather than a missing one. Left in place when nothing moved: there the object still describes exactly the stored version, and dropping a true fact would cost the common "refresh found nothing" case its citability for no reason.

Usage

.datom_forget_set_identity(x)

Arguments

x

The edited datom_set.

Details

Spelled x["f"] <- list(NULL), never x$f <- NULL, which would REMOVE the element and change names(x). A read set may legitimately report a NULL version, so the field exists and is empty rather than being absent.

A set never written has both fields empty already, so this changes nothing there.

Value

x, with version and data_sha emptied when it had them.


Format a Tag Map for One Line of Output

Description

key=value pairs, several labels joined by |, - when there are no tags. Tags are open-keyed, so a fixed column layout is impossible – do not try.

Usage

.datom_format_tag_line(tags)

Arguments

tags

A tag map, or NULL.

Value

A single string.


Build Connection from Local Repo + Store (Developer Path)

Description

Reads .datom/project.yaml for project identity and cross-checks against the store config. Uses the store for credentials.

Usage

.datom_get_conn_developer(path, store, endpoint = NULL)

Arguments

path

Path to datom repository.

store

A datom_store object.

endpoint

Optional S3 endpoint URL.

Value

A datom_conn object.


Build Connection from Store (Reader Path)

Description

Constructs a connection from a store object and project_name. Uses the data component of the store for S3 configuration.

Usage

.datom_get_conn_reader(store, project_name, endpoint = NULL)

Arguments

store

A datom_store object.

project_name

Project name string.

endpoint

Optional S3 endpoint URL.

Value

A datom_conn object.


Commits on This Branch the Remote Does Not Have

Description

The ahead half of the count .datom_check_git_current() already computes for its behind half: git2r::ahead_behind() element ⁠[[1]]⁠ is ahead, ⁠[[2]]⁠ is behind. No new git machinery.

Usage

.datom_git_ahead(path)

Arguments

path

Repository path.

Details

Does not fetch. The comparison is against the cached remote-tracking ref, so a stale ref can only cause an unnecessary push – and a push pulls first and is idempotent, so the cost of being wrong in that direction is a round trip. Fetching here would instead make a clean-tree call fail when offline.

Value

Integer count of unpushed commits, or NA_integer_ when it cannot be determined – no upstream tracking ref yet, or the comparison failed. NA means cannot prove there is nothing to publish, so callers push: that is exactly the state of a branch that has never been pushed.


Get Author Info from Git Config

Description

Reads user.name and user.email from the repository's git config.

Usage

.datom_git_author(path)

Arguments

path

Repository path.

Value

Named list with name and email.


Get Current Branch

Description

Returns the name of the currently checked-out branch. Aborts on detached HEAD (datom requires a branch).

Usage

.datom_git_branch(path)

Arguments

path

Repository path.

Value

Branch name as a string.


Commit Changes

Description

Stages the specified files and creates a commit.

Usage

.datom_git_commit(path, files, message, staged_deletions = FALSE)

Arguments

path

Repository path.

files

Character vector of files to add (relative to repo root).

message

Commit message.

staged_deletions

If TRUE, skip the file-existence check and use git2r::add(force = TRUE) so deletions can be staged. Default FALSE.

Value

Commit SHA as a string.


Work Out Which Commit First Produced Each of an Artifact's Versions

Description

Walks the commits that touched {name}/metadata.json, oldest-first, hashing the document as each commit left it. A commit whose document hashes to version V is a commit that produced V, and the first one reached is the one recorded – which is what makes a code-only commit nobody's producer: it leaves that document untouched, so it is not in the walk at all.

Usage

.datom_git_commit_shas_by_version(repo_path, name)

Arguments

repo_path

Path to the local clone.

name

Artifact name.

Details

One version maps to one-or-more commits by design, because a version is content-derived and code-invariant. Taking the oldest is not arbitrary tie-breaking; it answers "where did this version come from".

A repo git cannot answer for – a shallow clone, a rewritten history, a document that will not parse – yields no entry for the versions it lost. Callers omit the field in that case rather than recording a blank.

Every give-up here is silent on purpose, and that is not a house style: a version this cannot attribute is one the caller had no stored value for either, since a stored value is what stops it being asked about. So there is nothing to lose and nothing to report. The asymmetry with reading the stored copy, where a failure does lose something, is spelled out at the top of this file.

Value

Named character vector, commit sha named by version. Empty when nothing could be derived.


Build Git Credentials for HTTPS Remotes

Description

Returns a git2r::cred_user_pass object when the remote URL is HTTPS and a PAT has been supplied. Returns NULL for SSH remotes or when pat is absent.

Usage

.datom_git_credentials(remote_url, pat = NULL)

Arguments

remote_url

Character remote URL.

pat

GitHub personal access token. NULL (default) means no authentication; git2r will attempt unauthenticated or SSH access.

Details

The PAT must be supplied explicitly – datom does not read environment variables internally. Callers obtain the PAT from conn$github_pat, which is populated at conn-construction time from store$github_pat.

Value

A git2r::cred_user_pass object or NULL.


Ensure a Repo Has a Local Git Identity

Description

Sets user.name and user.email on the local config of repo so that git2r::default_signature(repo) succeeds even when the host has no global git identity (e.g. CI runners). Values are taken from global config when present; otherwise fallback constants are used.

Usage

.datom_git_ensure_local_identity(
  repo,
  fallback_name = "datom",
  fallback_email = "datom@noreply"
)

Arguments

repo

A git2r::repository handle.

fallback_name

Identity used when no global user.name is set.

fallback_email

Identity used when no global user.email is set.

Details

Idempotent: re-setting the same values is a no-op from git's perspective.

Value

Invisible repo.


Paths git Is Ignoring in a Clone

Description

git2r::status(ignored = TRUE) is the only route – git2r exposes no check-ignore verb – and it reports an ignored directory with a trailing slash and does not recurse into it, so a caller's path has to be matched against these as prefixes rather than compared for equality. The trailing slash is stripped here so one comparison covers a file entry and a directory entry.

Usage

.datom_git_ignored(path)

Arguments

path

Repository path.

Details

Only ever called from .datom_check_include_paths(), and only when the caller supplied paths, so no existing write gains a git read.

Value

Character vector of ignored paths, possibly empty, without trailing slashes.


Pull from Remote (Fetch + Merge)

Description

Fetches from the remote and merges upstream changes into the current branch. Aborts on merge conflicts - user must resolve manually. This is the primary defense against diverged histories.

Usage

.datom_git_pull(path, pat = NULL)

Arguments

path

Repository path.

pat

GitHub personal access token. Passed directly to .datom_git_credentials(). NULL means unauthenticated.

Value

Invisible TRUE on success.


Push to Remote

Description

Pulls (fetch + merge) first to detect conflicts, then pushes. Aborts on merge conflicts – user must resolve manually per spec.

Usage

.datom_git_push(path, pat = NULL, pull_first = TRUE)

Arguments

path

Repository path.

pat

GitHub personal access token. Passed directly to .datom_git_credentials(). NULL means unauthenticated.

Value

Invisible TRUE on success.


Get GitHub Username from PAT

Description

Calls GET /user to get the authenticated user's login.

Usage

.datom_github_username(pat, api_url = "https://api.github.com")

Arguments

pat

GitHub personal access token.

api_url

GitHub API base URL (default "https://api.github.com").

Value

Username string.


Check Whether a Gov Clone Exists

Description

Returns TRUE if gov_local_path is a directory that looks like a git repository (contains a .git folder). Does not validate the remote URL.

Usage

.datom_gov_clone_exists(gov_local_path)

Arguments

gov_local_path

Absolute path to the governance clone directory.

Value

Logical scalar.


Initialise Gov Clone (Clone If Missing, Reuse If Present)

Description

Ensures a valid gov clone exists at gov_local_path:

Usage

.datom_gov_clone_init(gov_repo_url, gov_local_path, pat = NULL)

Arguments

gov_repo_url

GitHub URL of the governance repo (e.g., "https://github.com/org/acme-gov.git").

gov_local_path

Absolute path where the gov clone should live.

pat

GitHub personal access token, threaded to .datom_git_credentials() so private governance repos can be cloned. NULL (default) means unauthenticated / SSH.

Details

Value

Invisible gov_local_path (character).


Open an Existing Gov Clone

Description

Returns a git2r repository handle for the gov clone at gov_local_path. Aborts if the path is not a valid git repository.

Usage

.datom_gov_clone_open(gov_local_path)

Arguments

gov_local_path

Absolute path to the governance clone directory.

Value

A git2r::repository object.


List Registered Project Names

Description

Returns the set of project names registered in the governance repo. When a local gov clone is available, lists directories under ⁠{gov_local_path}/projects/⁠ (offline-friendly, reflects last gov-clone refresh). Otherwise lists keys under ⁠projects/⁠ via the gov storage client and extracts unique top-level segments.

Usage

.datom_gov_list_projects(gov_conn, gov_local_path = NULL)

Arguments

gov_conn

A gov-scoped datom_conn (from .datom_conn_for(conn, "gov") or .datom_build_gov_resolve_conn()).

gov_local_path

Optional absolute path to a local gov clone. When provided and the clone exists, the filesystem path is preferred.

Details

Skips entries that don't contain a ref.json (corrupt registry rows).

Value

Character vector of project names (sorted, may be empty).


Build Project-Scoped Path Within Gov Clone

Description

Returns ⁠{gov_local_path}/projects/{project_name}/⁠. This is where dispatch.json, ref.json, and migration_history.json live for a given project in the shared governance repo.

Usage

.datom_gov_project_path(gov_local_path, project_name)

Arguments

gov_local_path

Absolute path to the governance clone directory.

project_name

Project name string.

Value

An fs_path character scalar.


Validate Gov Clone Remote URL

Description

Reads the first configured remote from the gov clone and compares it against expected_url. Aborts if they differ. This prevents silently reusing a clone that points at a different governance repo.

Usage

.datom_gov_validate_remote(gov_local_path, expected_url)

Arguments

gov_local_path

Absolute path to the governance clone directory.

expected_url

Expected remote URL (from store$gov_repo_url).

Details

URL comparison is normalised: trailing .git is stripped from both sides before comparison so ⁠https://github.com/org/acme-gov⁠ and ⁠https://github.com/org/acme-gov.git⁠ are treated as equivalent.

Value

Invisible TRUE.


Detect Changes Against Current Metadata

Description

Compares the proposed metadata_sha against the current version in S3. Returns the type of change detected.

Usage

.datom_has_changes(conn, name, new_data_sha, new_metadata_sha)

Arguments

conn

A datom_conn object.

name

Table name.

new_data_sha

SHA of the new data.

new_metadata_sha

SHA of the new metadata (from .datom_compute_metadata_sha()).

Value

Named list with two elements: change_type – "none" (no change), "metadata_only" (data same, metadata changed), or "full" (data changed) – and current, the already-read current metadata (or NULL for a brand-new table). Returning current lets datom_write() reuse it (the metadata_only parquet_sha carry-forward and the revert-to-older history scan) without a second storage read.


Canonical Recourse String for an Unhashable Column

Description

The single source of truth for the remediation advice attached to an unsupported column. Returns NULL when .datom_column_kind(x) classifies the column as hashable, otherwise the canonical recourse string for the first matching offender category. Both datom_check_hashable() and the .datom_canonical_hash() all-offenders abort call this one function, so the checker's advice and the abort's advice can never diverge.

Usage

.datom_hash_recourse(x)

Arguments

x

A single column (vector) from a data frame.

Details

The column name and class are added by the caller (a checker row or an abort bullet); the strings here are type-scoped only. Detection order matters: POSIXlt (a list under the hood) is matched before the generic list rows; the nested-data-frame list row before the generic list row; and the class-specific rows (units, sfc, yearmon/yearqtr/chron) before the "other classed" fallback.

Value

NULL when x is hashable, otherwise a canonical recourse string.


The Versions a History Names

Description

The Versions a History Names

Usage

.datom_history_versions(history)

Arguments

history

Parsed version_history.json, a list of entries.

Value

Character vector, NA for an entry with no usable version.


Add commit_sha to a History on Its Way to Storage

Description

Returns history with a commit_sha on every entry whose producing commit is known, and unchanged entries where it is not. Called by each of the three functions that upload version_history.json.

Usage

.datom_history_with_commit_shas(
  conn,
  name,
  history,
  version = NULL,
  commit_sha = NULL
)

Arguments

conn

A datom_conn object with a local path.

name

Artifact name, of either kind – version_history.json is shared.

history

The clone's parsed history, newest-first, as a list of entries.

version

The version this write produced, or NULL.

commit_sha

The commit that produced version, or NULL. A caller that made no commit passes NULL and every entry is derived.

Details

Two sources, in this order:

  1. What storage already holds. Cheap, and it is the only source for a value git can no longer produce.

  2. Derived from git, for the entries still missing after step 1 – and only then, so a repo whose history is complete pays no git walk.

version / commit_sha are the write path's shortcut: the caller has just made the commit that produced that version, so the walk is not needed for it. They are ignored when storage already records a commit for that version, since the recorded value is the first commit that introduced it and a later re-upload must not repoint it.

When step 1 failed rather than found nothing, and something was lost by it, this says so. The two states are not interchangeable: nothing to merge is the ordinary first write, whereas a stored copy that would not read means the values only storage had are now unknown, and the upload below replaces the file wholesale. The warning is raised only when a version actually ends up with no commit – if git could attribute every one of them, the same values were reconstructed and nothing is degraded.

Value

history with commit_sha filled in where it is known.


One id Field as Text, or NA

Description

A member read by datom_get_set() always has four single-string id fields – the read refuses a payload where one is not. This exists for the other input: a datom_set assembled by hand, which is supported and untrusted. NA rather than an abort so a listing still shows the member; the abort belongs to whoever tries to resolve it.

Usage

.datom_id_text(id, field)

Arguments

id

A member's id map.

field

One of project, name, kind, version.

Value

A single string, or NA_character_.


Is a Value a Single Non-Empty, Non-Missing String?

Description

The field test used by the member validator. Deliberately stricter than .datom_validate_parents()'s equivalent, which accepts NA_character_: that value is character, has length 1, and nzchar(NA_character_) is TRUE, so the obvious three-part test lets it through. A missing value in a member's id would be spliced into a storage key or written into a citable payload, so it is refused here.

Usage

.datom_is_text_scalar(x)

Arguments

x

Value to test.

Value

TRUE or FALSE.


Turn a Vector of Description Lines into cli Bullets

Description

Each line is interpolated as a value rather than embedded as message text, because a tag value may legitimately contain a brace and cli reads {anything} in message text as markup. Embedding the lines directly turns an artifact called ⁠dm{1}⁠ into a cli parse error instead of a message.

Usage

.datom_line_bullets(lines)

Arguments

lines

A character vector. Must be bound to the name lines in the frame that calls cli::cli_abort(), which is what the interpolation refers to.

Value

A character vector of bullets, each named *.


Union and deduplicate source_lineage lists (internal wrapper)

Description

Thin wrapper retained for existing internal callers. Delegates to the exported datom_lineage_union().

Usage

.datom_lineage_union(lineage_lists)

Arguments

lineage_lists

List of source_lineage lists (each a list of entries).

Value

Deduplicated list of source_lineage entries.


Description

The highest-value message in the set design, and it is a hint on failure, never a gate. Access in datom is per project and not conjunctive, so resolving a member of another project through this connection genuinely does not work – but without this bullet it presents as a missing object, which names the wrong problem and sends the reader looking for corruption.

Usage

.datom_link_failure(cnd, name, kind, record, conn)

Arguments

cnd

The condition the resolution raised.

name, kind

The member's name and kind.

record

The member record, which holds its recorded project.

conn

The connection the fetch was attempted through.

Details

Why it cannot be a check that runs first. A connection's project_name is not a verified fact. On a reader connection – the primary consumer of a set – it is a label passed to datom_get_conn(): the namespace comes from the store's root and prefix and nothing compares the label against the repo. So a mismatch is the ordinary case, and refusing on it aborts fetches that resolve correctly. Verified end to end by a test that fetches through a deliberately wrong label.

Recording the writer's own project name in metadata made the member's side of that comparison trustworthy; the connection's side is unchanged. Comparing a verified value against an unverified one still refuses working reads, which is why this stayed a hint. What it did buy is the wording: the message names the project the member's own writer recorded, rather than a project someone typed.

When the two names agree, the original condition is re-signalled untouched – same object, same class – because callers dispatch on those classes and a failure that has nothing to do with projects must not be reworded.

Value

Never returns; always signals.


Delete a File from Local Storage

Description

Delete a File from Local Storage

Usage

.datom_local_delete(conn, key)

Arguments

conn

A datom_conn object with backend = "local".

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

Invisible TRUE on success.


Delete All Files Under a Local Storage Prefix

Description

Removes the directory at root/{prefix}/datom/{prefix_key} and everything inside it. A missing prefix is a no-op.

Usage

.datom_local_delete_prefix(conn, prefix_key = NULL)

Arguments

conn

A datom_conn object with backend = "local".

prefix_key

Relative prefix (after ⁠prefix/datom/⁠).

Value

Invisibly, 1L if the directory was removed, 0L if not found.


Download File from Local Storage

Description

Copies a file from the store directory to a local path. Creates parent directories if needed.

Usage

.datom_local_download(conn, key, local_path)

Arguments

conn

A datom_conn object with backend = "local".

key

Relative storage key (after ⁠prefix/datom/⁠).

local_path

Local file path (destination).

Value

Invisible TRUE on success.


Check if Local Storage Object Exists

Description

Check if Local Storage Object Exists

Usage

.datom_local_exists(conn, key)

Arguments

conn

A datom_conn object with backend = "local".

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

TRUE or FALSE.


List Objects in Local Storage

Description

Lists files under a given prefix in the store.

Usage

.datom_local_list_objects(conn, prefix)

Arguments

conn

A datom_conn object with backend = "local".

prefix

Relative prefix to list under.

Value

Character vector of relative keys (relative to conn$root).


Resolve a Storage Key to a Local Path

Description

Builds the full filesystem path from conn$root, conn$prefix, and the relative key segments.

Usage

.datom_local_path(conn, key)

Arguments

conn

A datom_conn object with backend = "local".

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

An absolute filesystem path.


Read and Parse JSON from Local Storage

Description

Reads a JSON file from the store and parses it. Uses simplifyVector = FALSE to match S3 behavior.

Usage

.datom_local_read_json(conn, key)

Arguments

conn

A datom_conn object with backend = "local".

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

Parsed R list.


Upload File to Local Storage

Description

Copies a local file to the store directory. Creates parent directories if needed.

Usage

.datom_local_upload(conn, local_path, key)

Arguments

conn

A datom_conn object with backend = "local".

local_path

Local file path to upload.

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

Invisible TRUE on success.


Write an R List to Local Storage as JSON

Description

Serializes data to JSON and writes to the store directory. Creates parent directories if needed.

Usage

.datom_local_write_json(conn, key, data)

Arguments

conn

A datom_conn object with backend = "local".

key

Relative storage key (after ⁠prefix/datom/⁠).

data

An R list to serialize to JSON.

Value

Invisible TRUE on success.


Most-recent version_history document_sha for a data_sha

Description

The set half of .datom_lookup_history_object_sha(). Unlike its parquet sibling there is no legacy population to return NULL for: sets record document_sha from their first write, which is what lets a set read treat a missing one as an error rather than a skip.

Usage

.datom_lookup_history_document_sha(conn, name, data_sha)

Arguments

conn

A datom_conn object (developer, with local path).

name

Artifact name.

data_sha

Canonical content hash to match.

Value

Character document_sha, or NULL.


Most-recent version_history Stored-Object Hash for a data_sha

Description

Scans the developer's local version_history.json (newest-first) for the most recent entry whose data_sha matches and that carries a non-empty hash in field. Returns NULL when none is found. Reads the local git clone (offline-friendly); a stale clone is tolerated because the subsequent git push serializes concurrent writers (a behind clone fails to push before it can upload).

Usage

.datom_lookup_history_object_sha(conn, name, data_sha, field)

Arguments

conn

A datom_conn object (developer, with local path).

name

Artifact name.

data_sha

Canonical content hash to match.

field

"parquet_sha" (a table's stored parquet) or "document_sha" (a set's stored JSON payload).

Details

One scan serves both kinds, because the question is identical in each case – has this exact content already been stored, and under which byte hash? – and only the field name differs. Two copies would eventually disagree about what counts as a usable recorded value, and the reuse decision they feed is the one place where getting that wrong records a hash of bytes nobody stored.

Value

The recorded hash, or NULL.


Most-recent version_history parquet_sha for a data_sha

Description

The table half of .datom_lookup_history_object_sha(). Returns NULL for a pre-cv1 history, whose entries predate parquet_sha being recorded.

Usage

.datom_lookup_history_parquet_sha(conn, name, data_sha)

Arguments

conn

A datom_conn object (developer, with local path).

name

Artifact name.

data_sha

Canonical content hash to match.

Value

Character parquet_sha, or NULL.


Empty Manifest Skeleton

Description

The one shape of an empty manifest. Callers that need a manifest when none exists yet build it here rather than inline, so a later change to the manifest's shape has a single place to land.

Usage

.datom_manifest_skeleton(project_name = NULL)

Arguments

project_name

Project name, or NULL to omit the field (callers that only need somewhere to look up entries have no project name to hand).

Details

artifacts is a named empty list on purpose: jsonlite serializes an empty bare list as a JSON array (⁠[]⁠) and an empty named list as an object ({}), and a manifest's artifact block must be an object. Inert today, since nothing writes a manifest that still has zero entries, and correct for the one case where it would.

The skeleton declares schema_version itself, so no repo ever exists in a state that declares no format at all – not even between being created and receiving its first artifact. This covers only the built-from-nothing path: a document read from disk in an older shape gets its version from .datom_manifest_upgrade() instead, because the skeleton is unreachable whenever a manifest file exists.

Value

A list with schema_version, project_name (when supplied), artifacts and summary.


Apply Every Upgrade Step from a Declared Version to Current

Description

The dispatcher. Runs each step from declared up to .datom_supported_schema in order, then records the version it reached.

Usage

.datom_manifest_upgrade(manifest, declared)

Arguments

manifest

Parsed manifest document (a named list).

declared

Declared schema version, as returned by .datom_check_schema_version().

Details

declared is a parameter rather than something read off the document, because the only correct source for it is .datom_check_schema_version(), which returns it after refusing a document this build cannot convert. Taking it as an argument is what makes "check first, then upgrade" structural: there is no way to call this without having obtained the number from the check.

Identity on a document already at the current version – zero steps run. A document declaring a version above current is returned untouched too, since no step exists for it; that state is unreachable through the check, which aborts first.

Value

The document in current shape, declaring the version it reached.


Upgrade a v1 Manifest to v2

Description

v1 is every manifest written before the artifact namespace existed: the artifact list sits under tables and no entry declares what kind of artifact it is. v2 renames that key to artifacts and types every entry with kind = "table", which is what all of them are – sets did not exist.

Usage

.datom_manifest_upgrade_v1_to_v2(manifest)

Arguments

manifest

Parsed manifest document (a named list) declaring v1.

Details

The rename is done in place (names() assignment rather than add-then-remove), so the key keeps its position in the document and any sibling key this build does not recognise survives untouched.

A v1 document with no tables key at all is left with no artifact key. That is deliberate: an absent key and an empty one are different states – a truncated document versus a repo with nothing in it – and flattening them here would destroy the distinction a later self-healing read depends on.

Frozen. See the file header.

Value

The same document in v2 shape. The version is stamped by .datom_manifest_upgrade(), not here, so a step is never mistaken for the thing that records the result.


Mask a Secret for Display

Description

By default shows the first 4 characters followed by ⁠****⁠. That prefix is fine for GitHub PATs (the ghp_/github_pat_ prefix is a public type tag), but for AWS secret access keys and session tokens the first characters are real entropy – pass reveal_prefix = FALSE to mask them fully.

Usage

.datom_mask_secret(secret, reveal_prefix = TRUE)

Arguments

secret

A string.

reveal_prefix

If TRUE (default), reveal the first 4 characters. If FALSE, mask the whole secret (no characters revealed).

Value

Masked string.


The Position of a Member Named by a Record or a Link

Description

Matched on the whole id, which is a member's only unique key: the same name can appear twice, and two projects may both hold a dm.

Usage

.datom_match_member_id(members, record)

Arguments

members

The set's member list.

record

The record the caller passed, or the one a link carries.

Value

An integer vector of positions, normally of length one.


Description

The single route from a record to a callable link, used by datom_fetch_member() and by every leaf datom_structure_members() produces – and it is the same factory datom_get_set() uses for ⁠$fetch⁠. So kind dispatch, and the project hint that lives beside it, have one implementation reached by every route.

Usage

.datom_member_as_link(record, what = "member")

Arguments

record

A member record, with or without a fetch element.

what

Noun for messages about an unusable record.

Details

Any fetch already on the record is dropped first, so a record that came from a read produces a link over pure data rather than a link carrying a link.

Value

A datom_link.


Does a Member Carry All the Labels Asked For?

Description

Every key must be present and every value listed under it must be one the member carries. So tags = list(domain = "safety") matches a member tagged domain = c("safety", "efficacy") – narrowing by one label of a multi-valued tag is the ordinary case, since multi-valued tags are the point.

Usage

.datom_member_has_tags(member, tags)

Arguments

member

A member record.

tags

The filter map.

Details

The member's labels are read through .datom_tag_pairs(), not off the map, so there is genuinely one access path to a member's tag values and the duplicate-key hazard documented there cannot be reintroduced here. A member$tags[[k]] read is the same silent-first-match defect, and it fails in the direction that looks like missing data: the member is reported not found under a label the document says it carries.

The filter side is read by position for the same reason, even though datom_fetch_member() refuses a filter with duplicate keys before this runs.

Value

TRUE or FALSE.


A Member's id, or an Abort Saying the Pointer Cannot Be Resolved

Description

Checks only what resolution needs, and deliberately does not call .datom_validate_members(): that is the write-side contract, and it refuses an id field a newer datom added – which the read deliberately carries. Using it here would make an unknown field readable but unfetchable, which is a reads-limp violation arriving by a side door.

Usage

.datom_member_id(record, what = "member")

Arguments

record

A member record.

what

Noun for the message.

Value

The id map.


One Line Describing a Member, for a Message That Has to Name Several

Description

Name, kind, the first 8 characters of its version, and its tags. The version costs nothing – a read member's id$version is the full recorded string – and it is what the reader needs to narrow an ambiguous name.

Usage

.datom_member_lines(members)

Arguments

members

A member list.

Value

A character vector, one entry per member.


Description

Builds the ⁠$fetch⁠ link every member of a read set carries: call it with a connection to the member's project and it resolves the pointer – a table member to data via datom_read(), a set member to references via datom_get_set().

Usage

.datom_member_link(name, kind, version, record)

Arguments

name, kind, version

The member's pinned identity – the three facts resolution needs. project is deliberately not a parameter: see above.

record

The member record the link describes – pure data, attached as the datom_member attribute, and where project remains readable.

Details

fetch rather than read or get because it is genuinely both. This is the one polymorphic door in the design, and the member level is where the domain forces it: iterating members, the caller cannot know each kind in advance. At the top level they named one artifact they chose, which is why datom_read() and datom_get_set() stay separate verbs.

This function is namespace-level, and that is load-bearing. A factory defined inside datom_get_set() would put that call's frame – which holds conn, and therefore the PAT – on the closure's parent chain, and saveRDS() of the member would write the token into the file. Measured, same code both ways: nested, 2094 bytes with the token present; namespace-level, 1609 bytes without. Every argument is forced so that nothing is left as a promise pointing back at the caller's frame. The guard is a test on the serialized bytes, not on environment(link), because an environment check passes on the broken shape – there the connection sits one frame further up.

The link carries its own pointer as an attribute, so a consumer holding only a projection can still cite what they used. Links built without it cannot be repaired afterwards, which is why it ships with the factory rather than later.

It does not compare the member's project against the connection's, and it must not. That looks free – both names are in hand – and it would refuse working reads. For a reader connection, which is the primary consumer of a set, project_name is a label the caller passes to datom_get_conn(): the namespace comes from the store's bucket and prefix and nothing validates the label against the repo. So a mismatch is the ordinary case rather than the error case, and a gate here would abort a fetch that resolves correctly. Recording the writer's own project name in metadata does not change this. It makes the member's side of the comparison trustworthy; the connection's side is still a label nobody checked, so comparing them still refuses working reads. Pinned by a test that fetches through a deliberately mismatched label. A hint on an already-failed resolution is a different thing and is left to the task that owns that message.

Value

A function of one argument (conn), classed datom_link.


Read the Parents One Table Member Records

Description

The snapshot read behind .datom_check_set_parents(). The format check sits outside the read's handler, so a snapshot from a newer datom keeps its own refusal instead of being reworded as a read failure – the same pairing as .datom_parent_record().

Usage

.datom_member_parents(conn, id)

Arguments

conn

The set's own developer connection.

id

The member's id map.

Value

The snapshot's parents list, or NULL when it records none.


Resolve the Third Argument of datom_fetch_member() to a Member Record

Description

One accessor for the three shapes a caller holds, so a console call and a loop use the same verb: a name, a member record, or a link.

Usage

.datom_member_record(members, member, tags = NULL, version = NULL)

Arguments

members

The set's member list.

member

A name, a member record, or a datom_link.

tags

Optional label filter.

version

Optional version, or a prefix of one.

Details

The shape dispatch itself is .datom_member_shape(), shared with datom_add_member(). Only the name half is here, and it genuinely differs between the two verbs: a name means "a member of this set" here and "an artifact in this project's storage" there, so a shared lookup would search the wrong thing on one of the two routes.

A record with no fetch on it is accepted, and that matters: it is the payload shape – what a caller who built a member with datom_member() holds, and what stripping a read set's links produces. The accessor keys on id and nothing else, which is what keeps this verb and datom_write_set() agreeing about what a member is.

tags and version narrow a name. Supplied beside a record or a link they are refused rather than ignored, because ignoring them would resolve a different version than the one asked for and report success.

Value

One member record.


Which of the Three Shapes a Member Argument Arrived In

Description

A caller naming one member holds one of three things, and every verb that takes a member accepts all three: a name, a member record (what datom_member() returns, and what stripping a read set's links produces), or a link (a member's fetch element, or a leaf of datom_structure_members()).

Usage

.datom_member_shape(member, arg = "member")

Arguments

member

The value the caller passed.

arg

Argument name for the message.

Details

This is the shape dispatch alone, deliberately without the lookup. What a name means differs by verb – to datom_fetch_member() it is a member of the set already in hand, to datom_add_member() it is an artifact to look up in the project's storage – so handing the name back to the caller is what lets one dispatch serve both without either searching the wrong thing.

The refusal of tags / version beside a record or a link is left to each caller too: both refuse, and the reason differs enough to word differently (narrowing a search versus declaring a member twice). shape is the phrase to name it by, so the two messages at least agree on what the caller passed.

Value

A list of shape (a phrase naming what arrived) and record (the member record, or NULL when a name arrived).


Read a Metadata Document as One Commit Left It

Description

revparse_single(repo, "<sha>:<path>") is the whole mechanism: it resolves git's own commit:path syntax straight to the blob and raises when the path is absent at that commit. Indexing the tree object instead returns an empty list for a path that is not there, which reads as a successful lookup.

Usage

.datom_metadata_at_commit(repo, sha, rel)

Arguments

repo

A git2r repository handle.

sha

Commit sha.

rel

Repo-relative path of the document.

Value

The parsed document, or NULL when it cannot be read.


Every Metadata Field Name This Build Knows

Description

The two halves of the classification joined: the fields that make up a version's identity, and the fields datom deliberately keeps out of it. A name in neither half is a name this build cannot place.

Usage

.datom_metadata_known_fields()

Details

A function rather than a stored vector, for two reasons that both bite. ⁠R/⁠ is sourced alphabetically (DESCRIPTION declares no Collate), and this file sorts before R/utils-sha.R where both halves are defined – so a constant built from them here would be built from values that do not exist yet and the package would fail to install. Deriving it at call time also means it cannot fall out of step with either half.

Append-only. A name that has ever been written must keep classifying forever, including names datom no longer writes: a build that forgets one meets an older document, fails to place a field it should know, and starts preserving as unfamiliar something it could have handled – or, once the write-side refusal lands, refuses the document outright and blocks the upgrade direction, which must always work.

Value

Character vector of field names, unsorted.


Hash an Already-Selected Set of Metadata Fields

Description

The canonical-form half of metadata_sha, split from field selection so that each half is testable on its own: this function decides how a chosen set of fields becomes bytes and knows nothing about which fields are identity.

Usage

.datom_metadata_sha_from_fields(fields)

Arguments

fields

Named list of fields to hash, already filtered to the identity set by .datom_compute_metadata_sha().

Details

Sorts field names by C-locale byte order (method = "radix") before hashing so the result is deterministic regardless of field insertion order and regardless of the host's LC_COLLATE (default collation sorts differ between C and e.g. en_US.UTF-8, which would otherwise make the same metadata hash differently on different machines). Sorting here rather than relying on the declared order of .datom_metadata_identity_fields is deliberate: it means hash stability does not depend on how that constant happens to be written, so re-ordering it for readability cannot silently change every recorded version.

Value

Character SHA-256 hash.


Normalize a prefix value to NULL or a non-empty string

Description

A NULL prefix serializes to JSON as an empty object () and reads back as an empty list, not NULL. Empty strings can also creep in. This collapses all empty-ish forms (NULL, list(), "", NA) to NULL so that location equality checks survive a JSON round-trip.

Usage

.datom_normalize_prefix(prefix)

Arguments

prefix

A raw prefix value from a parsed ref.json.

Value

NULL or a single non-empty character string.


Say That a Manifest's Format Was Moved Forward

Description

Called from the two places that persist a converted manifest: the entry updater, which rewrites the git-tracked file, and the data-side metadata sync, which mirrors the converted document to storage. Reads convert too and stay silent, deliberately – a read changes nothing, and a line on every datom_list() call would be noise nobody can act on.

Usage

.datom_notify_manifest_upgraded(declared, where)

Arguments

declared

The version the document declared before conversion, as returned by .datom_check_schema_version().

where

Human-readable name of the copy being written.

Details

Why say anything: conversion is one-way for everybody else. Once this repo's manifest declares the newer format, a collaborator on an older datom no longer finds the artifact list where their build looks for it, and their datom_list() reports an empty repo without erroring. Their datom_read() keeps working, because the data path never touches the manifest. That is a real consequence of a command whose stated job was something else – datom_validate(fix = TRUE) in particular reads as a repair – and an unannounced one is the silent degradation the whole schema contract exists to remove.

No-op when the document was already current, which is every ordinary write.

Value

Invisibly NULL.


Deduplicate and Order a Member List for the File

Description

Drops exact duplicates – same id and same tags – by datom-sv1 member digest, then sorts by project, name, version.

Usage

.datom_order_set_members(members)

Arguments

members

An unnamed list of validated member records.

Details

Two sort keys exist and each has its own reason. The identity hash orders member digests, which is what keeps the encoder from having to know what an id looks like. The file orders by name, which is what keeps an entry in place when its tags change so that ⁠git diff⁠ shows one changed field. version is in the key because two versions of one name are legal members, and would otherwise have no defined relative order.

No tiebreaker is required, and none may be added. The only way two members can share project || name || version is the same id with different tags, which survives dedup because the digest covers tags – and that payload is refused one step later. R's radix sort is stable, so the tie resolves to caller order in the meantime. A defensive tiebreaker would be dead code.

Runs after validation, unlike the rest of canonicalization, because the digest is computed by the identity encoder and the encoder refuses a value it cannot encode. Reaching it first would report a bad tag value in the encoder's words rather than the validator's.

Value

The members, deduplicated and ordered.


Resolve One Parent Record From Its Versioned Snapshot

Description

The body of datom_parent() for one table at one named version, shared by both of its routes so a parent declared by version and one declared from a set are read, checked and shaped by the same code.

Usage

.datom_parent_record(conn, table, version)

Arguments

conn

A datom_conn scoped to the parent's project.

table

Parent table name.

version

Parent version.

Value

One parent record.


Declare Parents at the Versions a Set Pins

Description

The ⁠x =⁠ route of datom_parent(). Each name is resolved by .datom_find_member(), the resolver datom_fetch_member() uses for a name, so the two verbs cannot pick different members for the same name and labels. Tag validation is the same call with the same remedy, for the same reason.

Usage

.datom_parents_from_set(conn, table, x, tags)

Arguments

conn

A datom_conn scoped to the parents' project.

table

Character vector of member names.

x

A datom_set.

tags

Optional label filter.

Value

An unnamed list of parent records, one per table.


Parse a ref.json structure into a location list

Description

Common parsing logic shared by storage-backed and clone-backed ref readers.

Usage

.datom_parse_ref(ref, source)

Arguments

ref

Parsed ref.json content (R list).

source

Identifier for error messages (root, key, or path).

Value

A named list with root, prefix, region.


Parse S3 URI into Components

Description

Extracts bucket and prefix from an ⁠s3://⁠ URI.

Usage

.datom_parse_s3_uri(uri)

Arguments

uri

Character string S3 URI (e.g., "s3://my-bucket/prefix/path").

Details

Mapping from URI to components, for reference:

"s3://my-bucket/data/proj" -> list(bucket = "my-bucket", prefix = "data/proj")
"s3://my-bucket"           -> list(bucket = "my-bucket", prefix = NULL)

Value

Named list with bucket (character) and prefix (character or NULL).


The Metadata Document Already in the Clone, If Any

Description

Reads {name}/metadata.json from the local git checkout, for the one purpose of finding fields to carry forward. Returns NULL when there is no such file or it will not parse – both mean there is nothing to preserve, and neither is this function's business to report: a brand-new artifact legitimately has no prior document, and an unparseable one fails moments later on its own terms.

Usage

.datom_prior_metadata(conn, name)

Arguments

conn

A datom_conn object with a local path.

name

Artifact name.

Details

The clone's copy, not storage's. Three reasons, any one sufficient: it is the file being overwritten, so preserving its own content is the claim being made; it is a local file read rather than a network round trip; and it is where a pull from a collaborator on a newer datom lands. Storage cannot legitimately hold a newer document than the clone, because git is written first and gates the storage mirror – if it does, that is drift, and datom_validate() owns drift.

Value

The parsed document, or NULL.


Push Metadata Files to S3

Description

Uploads metadata.json, version_history.json, and a versioned snapshot to S3. Called AFTER git commit+push succeeds to maintain local → git → S3 ordering.

Usage

.datom_push_metadata_s3(conn, name, metadata, metadata_sha, commit_sha = NULL)

Arguments

conn

A datom_conn object.

name

Table name.

metadata

Named list for metadata.json.

metadata_sha

SHA of the metadata (the datom "version").

commit_sha

The commit that produced metadata_sha, or NULL from a caller that made no commit.

Details

The stored history carries one field the clone's copy cannot, and this is the reason commit_sha exists as an argument here: the clone's version_history.json is inside the commit that would name it, so only a storage-bound copy can say which commit produced a version. The upload sends the clone's file wholesale, so without the merge below the field would survive on the newest version only – the second write of an artifact would erase the first version's commit id.

commit_sha is derived, never authored. It reaches this function as an argument only because the caller one layer up already holds the commit it just made; no exported verb accepts it, and every other entry's value is worked out from git. See R/version-commit.R.

Value

Invisible character vector of S3 keys written.


Read governance.json from Local Git Clone

Description

Reads and validates {path}/.datom/governance.json. Returns NULL when the file is absent (project is not gov-attached). Aborts on malformed JSON or failed schema validation.

Usage

.datom_read_governance_json_local(path)

Arguments

path

Absolute path to the root of the local data git clone.

Value

Parsed list or NULL.


Read a Manifest and Check Its Schema Version

Description

The single manifest read. Every reader that takes a manifest into datom goes through this, so the compatibility check happens once and cannot be softened by a caller's error handling.

Usage

.datom_read_manifest(
  conn,
  scope = c("storage", "clone"),
  operation = c("read", "write")
)

Arguments

conn

A datom_conn object.

scope

"storage" for the copy in data storage (.metadata/manifest.json), "clone" for the git-tracked copy (.datom/manifest.json). Both exist; they can differ, and which one a caller wants is a real choice rather than a default.

operation

What the caller is about to do with the document – "read" (default) or "write". Passed through to .datom_check_schema_version(), where it only selects a word in the refusal message, so that a write stopped at the door does not report the format as one this build "cannot read".

Details

Two kinds of failure, handled deliberately differently:

And one document that is not a failure at all. When the artifact list is missing from where this build looks for it – either because the format is newer than this build knows, or because the key is simply not there after the conversion has run – the index is reconstructed from storage and a warning says so. The manifest summarises documents that each hold the same facts, so it is the one datom-owned file with something to rebuild it from. A writer meeting either condition is refused instead (.datom_check_write_entry()): reads limp, writes stop.

Value

A list with:


Read Table Metadata from S3

Description

Fetches both metadata.json (current state) and version_history.json (version index) for a given table from S3.

Usage

.datom_read_metadata(conn, name)

Arguments

conn

A datom_conn object.

name

Table name (validated).

Value

Named list with current (metadata.json contents) and history (version_history.json contents as a list of entries).


Download and Read Parquet from S3

Description

Downloads ⁠{table}/{data_sha}.parquet⁠ from S3 to a temporary file and reads it via arrow::read_parquet(). When an expected parquet_sha is supplied (non-empty), the downloaded object's SHA-256 is verified against it BEFORE parsing, so corruption or tampering aborts rather than being silently read.

Usage

.datom_read_parquet(conn, name, data_sha, parquet_sha = NULL)

Arguments

conn

A datom_conn object.

name

Table name.

data_sha

SHA identifying the parquet file.

parquet_sha

Expected SHA-256 of the stored parquet object bytes, from the resolved metadata (see .datom_resolve_version()). When non-empty, the downloaded file is verified against it and a mismatch aborts. When NULL or empty – which now happens only for pre-cv1 metadata – the integrity check is skipped and the read succeeds.

Value

Data frame.


Normalize One Member Record Read Back from a Payload

Description

Normalizes representation in id and tags, then makes the one refusal the read owns: an id field that is not a single non-empty string after normalization aborts as a malformed document, naming the member.

Usage

.datom_read_set_member(m, at, name)

Arguments

m

One parsed member record.

at

Position label used in error messages, e.g. "members[[2]]".

name

The set's name, for error messages.

Details

Why id is refused where a tag value is tolerated. id values are spliced into storage keys and compared against project names, and .datom_validate_members() enforces that contract on write only – so the read is the only place a payload's id is ever checked. Normalizing without refusing would silently accept a document datom_write_set() cannot produce, and a caller comparing a list against a string would conclude that a member of this project belongs to another one.

Fields outside the four are left alone rather than refused: a newer datom may have added one, and this build never reads it.

Value

The member record, normalized.


Turn a Parsed Member List into Resolvable Member Records

Description

Turn a Parsed Member List into Resolvable Member Records

Usage

.datom_read_set_members(members, name)

Arguments

members

The payload's parsed member list.

name

The set's name, for error messages.

Value

An unnamed list of member records, each carrying ⁠$fetch⁠.


Download, Verify and Parse a Set's Stored Payload

Description

Download, hash, then parse. The order is the point: a set read must not parse an unverified payload, which is the same gate position .datom_read_parquet() uses for parquet_sha.

Usage

.datom_read_set_payload(conn, name, data_sha, document_sha)

Arguments

conn

A datom_conn object.

name

Set name.

data_sha

The resolved content hash – the payload's storage address.

document_sha

The recorded SHA-256 of the stored payload bytes.

Details

.datom_storage_read_json() cannot be used here, and it would work. It parses, so after calling it there is nothing left to hash but bytes re-serialized locally – a hash of bytes nobody stored, which is exactly the defect the write path guards against, inverted. It returns a structure identical to parsing the downloaded file, so nothing fails if you reach for it; the integrity check simply stops meaning anything.

A missing document_sha is an error, not a skip. parquet_sha's skip-on-absent branch exists purely as a grace for metadata written before that field did. Sets have recorded document_sha since their first write, so there is no legacy population to be lenient about, and reproducing the grace would build a silent-degradation path on purpose.

data_sha is deliberately not recomputed from the parsed payload. It is the address the payload was fetched from, so it catches nothing document_sha did not, and it would refuse a payload a newer datom wrote – the sv1 encoder aborts on a top-level payload key it does not know. Same reason the parsed payload is not re-validated. Reads limp.

Value

The parsed payload, with members kept as a list of records.


Normalize a Parsed JSON String Array to a Character Vector

Description

jsonlite::fromJSON(simplifyVector = FALSE) returns a JSON array of strings as a list of length-1 characters, and auto_unbox = TRUE on the write means a single label was written as a bare string. So one tag key comes back in three shapes – character(1), a list of 1, or a list of n – for what is one value in the document.

Usage

.datom_read_string_array(v)

Arguments

v

A parsed JSON value.

Details

Same strings, same order, same count: this is a representation change, not a content change, which is why order is preserved and duplicates are kept. Sorting or deduplicating here would be the write's canonicalization performed by a reader.

Anything that is not an all-text array is returned untouched. A reader has no caller intent to tidy toward and nothing downstream requires tag values to be text, so an odd value is reported by whoever tries to use it rather than refused here.

Value

A character vector when v was an all-text array, otherwise v.


Normalize a Parsed Tag Map's Values

Description

Applies .datom_read_string_array() to every value and does nothing else: no key sorting, no value sorting, no deduplication, no dropping of an empty-valued key. A map with no names is returned untouched rather than refused, for the same reason a single odd value is.

Usage

.datom_read_tag_map(tags)

Arguments

tags

A parsed tag map, or NULL.

Value

The map with each value normalized, or NULL.


Reconstruct the Whole Artifact Index from Storage

Description

One storage listing plus two reads per artifact. The result is a complete manifest in this build's shape: the artifact rows, the summary counters recomputed from them, and the current format declared.

Usage

.datom_rebuild_manifest(conn, prior = NULL)

Arguments

conn

A datom_conn object.

prior

The document being replaced, or NULL.

Details

project_name and updated_at are carried from the document being replaced when it has them. Both are recorded facts about the repo rather than about the artifacts, so neither is recoverable from a listing – and inventing a fresh updated_at would state that the index was rewritten now, when nothing was written at all.

Aborts rather than returning a partial index. Half an artifact list is indistinguishable from a repo that only has half those artifacts, and the caller's job is to decide what an unreachable store means – see .datom_read_manifest(), which keeps a schema refusal separate from an IO failure on the way back out.

Value

A manifest in current shape.


Rebuild One Artifact's Manifest Row from Its Own Documents

Description

Every field on the row is copied from metadata.json, counted from version_history.json, or – for a set's member count – read from the payload. The row's shape has to match what .datom_update_manifest_entry() writes, field for field, or a rebuilt repo answers differently from a healthy one – so the two are pinned against each other by a test rather than by matching comments.

Usage

.datom_rebuild_manifest_entry(conn, name)

Arguments

conn

A datom_conn object.

name

Artifact name.

Details

last_updated is the one field with no recorded source: the writer stamps the wall clock at the moment it rewrites the row, and that moment is not in any document. The version's own created_at is used instead, which is the closest true statement available – when this artifact's current state was written.

A set's row is built from different fields, and costs a third read. A set carries member_count where a table carries size_bytes, and that count lives in the payload rather than in either document read here – hence .datom_rebuild_member_count(). Putting a size_bytes on a set row instead would be worse than leaving the count out: the default is 0, which has length 1 and therefore survives purrr::compact(), so the row would state that the artifact is zero bytes.

Value

A named list: one manifest artifact row.


How Many Members a Set's Current Payload Holds

Description

The third storage read a set's row costs. Only a set needs it, and only a rebuild pays it: the healthy writer knows the count from the payload it just canonicalized.

Usage

.datom_rebuild_member_count(conn, name, data_sha)

Arguments

conn

A datom_conn object.

name

Set name.

data_sha

The current version's content hash – the payload's address.

Details

Returns NULL for anything that is not a readable payload – an unusable data_sha, a missing object, a document that will not parse. That is the same trade the rest of this file makes: an absent count is a gap datom_validate() owns, while a stand-in count would be a statement about the set's contents that nothing supports.

Value

An integer count, or NULL.


The Set Name a Write Uses, From the Call and From the Set Itself

Description

A datom_set carries its name, and the caller may pass ⁠name =⁠ too. When both are given they must agree: preferring either would write a set under a name one of them did not say. When only one is given it is the one used, and .datom_check_set_write_gates() then checks it against the repo's declared set – so a set named for another repo stops there, with the gate's message. A set with no name (an assembled one, usually) takes the declared one.

Usage

.datom_reconcile_set_name(name, set_name)

Arguments

name

The name argument, or NULL.

set_name

The name the set carries, or NULL (also when members was a plain list of records).

Value

The name to hand to the gate, or NULL.


The Version Storage Recorded for an Artifact's Current State

Description

Picks the version_history.json entry that describes metadata.json, and returns the version recorded on it.

Usage

.datom_recorded_current_version(meta, history)

Arguments

meta

The artifact's parsed metadata.json.

history

The artifact's parsed version_history.json, a list of entries, or NULL.

Details

Never recomputed, and that is the point of this function existing at all. Hashing metadata.json here would reach for the identity code in precisely the scenario a rebuild is for – a repo touched by a build whose field classification differs from this one's – and publish a current_version matching no version in the recorded history. An index pointing at a version that does not exist is worse than the empty list it replaced.

Which entry describes the current state is not simply the newest one. History is prepended newest-first, but a write that reverts to content already in the history appends nothing, so the current state can be an older entry. The selection therefore narrows by recorded fields only:

  1. Entries whose data_sha equals the current document's. One match settles it – this is the revert case, and it is why the newest entry alone is wrong.

  2. Several matches means metadata-only versions of the same content; the one whose timestamp equals the document's created_at is the current one, since a version's history entry copies that field verbatim.

  3. Anything still ambiguous takes the newest candidate. That is a choice between two entries that both describe the current content, so the worst case is naming the wrong one of two versions of the same bytes.

No match at all returns NULL, and it deliberately does not fall back to the newest entry. No match means the history does not record the state metadata.json describes – a truncated or partly-synced history. The newest entry there is a version of different content, so naming it would be a wrong statement rather than a missing one, and a row already tolerates carrying no version. datom_validate() owns the inconsistency. This is the same trade the carry-forward rule makes in R/forward-compat.R: a stale claim that outlives what it described is worse than an absent one.

Value

The recorded version string, or NULL when the history records nothing usable – in which case the rebuilt row simply carries no version, rather than a manufactured one.


Refuse a File-Import Argument on a Product Repo

Description

Silently ignoring it would let a caller believe the argument did something.

Usage

.datom_refuse_file_arg_on_product(arg)

Arguments

arg

The argument that was supplied.

Value

Does not return; aborts with class datom_sync_file_arg_on_product.


Refuse the File-Import Path on a Product Repo

Description

A mode: product repo builds its artifacts: derived tables written from data frames, and one set collecting them. It never onboards source files, so the two import verbs refuse instead of answering. Before this they answered unhelpfully – ⁠input_files/⁠ exists and is empty on such a repo, so the scan reported "no files found" and handed back a zero-row frame, which describes a repo with nothing to import rather than a repo that does not import.

Usage

.datom_refuse_import_on_product(verb, context)

Arguments

verb

Name of the sync verb being refused, for the message.

context

What .datom_sync_context() read, so one call makes one gated parse.

Details

Read from the file, not from the connection, and the rule behind that is worth carrying: a check that authorises a write must see the config as it is now, because a hand edit or a pull can replace it after the connection was built. Only datom_status(), which reports rather than decides, reads the mode off the connection. A future site applies the same test: does it authorise a write? Then it reads the file.

Three steps in one place, and the middle one is easy to leave out. Parsing this file makes this a new gated parse: every site that reads .datom/project.yaml checks its declared format first, or a build that cannot interpret the file acts on fields it has misread. Skipping that step here would reopen exactly that hole, on a path that writes.

Called from both sync verbs, not only the first. datom_sync() takes a manifest data frame, so a caller can hand it rows that a refusing datom_sync_manifest() would never have produced.

Above the input-file scan, never in its empty branch. A product repo with a file dropped into ⁠input_files/⁠ by accident would otherwise be imported, which is the thing this exists to prevent; the unhelpful no-op only happened when the directory was empty.

Both verbs have a set route on a product repo, reached by passing ⁠sources =⁠, so the message names that route rather than only the write verbs.

Value

Invisibly NULL. Aborts with class datom_import_on_product when the repo declares mode: product.


Refuse the Set's Own Project as a Source

Description

The set's own project holds its outputs, which are derived from the inputs and move only once they are re-derived. Checked on the labels, before any read.

Usage

.datom_refuse_own_project_source(own, labels)

Arguments

own

The set's own project name.

labels

The project names of the source connections.

Value

Invisibly NULL; aborts with class datom_sync_own_project_source.


Refuse ⁠sources =⁠ (or Another Set-Path Argument) on an Ordinary Repo

Description

Refuse ⁠sources =⁠ (or Another Set-Path Argument) on an Ordinary Repo

Usage

.datom_refuse_sources_on_ordinary(arg)

Arguments

arg

The argument that was supplied.

Value

Does not return; aborts with class datom_sync_sources_on_ordinary.


Render README.md from Template

Description

Reads the template from inst/templates/README.md and fills in project-specific values using {{{ }}} delimiters.

Usage

.datom_render_readme(
  project_name,
  backend = "s3",
  root,
  prefix,
  region = NULL,
  remote_url,
  gov = NULL
)

Arguments

project_name

Project name string.

backend

Storage backend ("s3" or "local").

root

Storage root (S3 bucket name or local directory path).

prefix

Storage prefix (can be NULL).

region

AWS region string (NULL for local backend).

remote_url

Git remote URL.

gov

Governance store component (e.g. from datom_store_s3()), or NULL for a solo project with no governance attached. Determines whether the rendered store snippets use governance = NULL or a gov-store constructor.

Value

Character string — the rendered README content.


Repoint One Member at One Version

Description

Points 1, 2 and 3 of this file's header all live here: the labels are attached verbatim rather than passed through the constructor, the link is rebuilt through the shared factory, and the rebuilt record's recorded project is compared against the one it replaces.

Usage

.datom_repoint_member(record, conn, version)

Arguments

record

The member record being replaced.

conn

The connection for that member's project.

version

The version to pin.

Value

The new member record, carrying the old labels and, when the old record had one, a link to the new version.


Say What Was Dropped, and That Nothing Was Written

Description

Say What Was Dropped, and That Nothing Was Written

Usage

.datom_report_member_removals(dropped, left, n = 20L)

Arguments

dropped

The edit rows for the removed members.

left

How many members remain.

n

Maximum number of lines to print before truncating.

Value

Invisibly NULL.


Say What Moved, What Did Not, and That Nothing Was Written

Description

The report is the deliverable rather than decoration: nothing is written, so this is the dry run, and it is the only place the caller sees what an inferred "current" resolved to.

Usage

.datom_report_member_updates(changes, gone, skipped_lines, n_selected, n = 20L)

Arguments

changes

The change table, possibly with zero rows.

gone

The selected members whose artifact no longer appears in its project.

skipped_lines

Description lines for members skipped as ambiguous.

n_selected

How many members the call selected.

n

Maximum number of lines to print before truncating.

Details

Lines are emitted with cli::cli_verbatim() because they embed artifact names and label values, and cli reads {anything} in message text as markup – an artifact called ⁠dm{1}⁠ would be a parse error rather than a line.

Value

Invisibly NULL.


Say What Apply Did, and That Nothing Was Written

Description

Say What Apply Did, and That Nothing Was Written

Usage

.datom_report_sync_apply(todo, n = 20L)

Arguments

todo

The new and changed rows that were applied.

n

Maximum number of lines to print before truncating.

Value

Invisibly NULL.


Say What the Set Sync Preview Found

Description

One summary line, then one warning per group of rows or members the caller has to know about, each with its remedy.

Usage

.datom_report_sync_preview(
  result,
  n_sources,
  ambiguous_lines,
  ambiguous_first,
  unpassed,
  gone,
  unversioned
)

Arguments

result

The preview frame.

n_sources

How many sources were mapped.

ambiguous_lines

One line per member behind an ambiguous row.

ambiguous_first

The name of the first ambiguous artifact, for the remedy, or NULL.

unpassed, gone

Member rows (project, name, kind, version) for members of a project not passed, and members whose artifact is no longer listed in their source.

unversioned

"name in project" for artifacts whose manifest entry records no current version.

Value

Invisibly NULL.


Require Governance Attached on a Connection

Description

Guard helper used by gov-only commands (e.g. datom_projects) to fail with a single uniform message when called on a no-governance connection.

Usage

.datom_require_gov(conn, what)

Arguments

conn

A datom_conn object.

what

Character. The user-facing name of the calling function (e.g. "datom_projects()"), used in the error message.

Value

Invisible TRUE when gov is attached. Aborts otherwise.


Resolve Data Location via Ref (Conn-Time Helper)

Description

Called during datom_get_conn() for both readers and developers when a governance store is present. Reads ref.json from governance, detects migration (store$data location != ref location), and returns the ref-resolved location.

Usage

.datom_resolve_data_location(
  store,
  role,
  project_name = NULL,
  path = NULL,
  gov_local_path = NULL,
  endpoint = NULL
)

Arguments

store

A datom_store object with governance component.

role

"developer" or "reader".

project_name

Project name (required when governance is present).

path

Local repo path (developers only; NULL for readers).

gov_local_path

Absolute path to the local gov clone (developers only; NULL for readers or when the clone does not yet exist).

endpoint

Optional S3 endpoint URL.

Details

Read path is role-aware:

Value

A named list with root, prefix, region from the ref, or NULL if no governance store is present (skip ref resolution).


Resolve the document_sha to Record and Whether to Upload

Description

The set analogue of .datom_resolve_parquet_sha(), kept beside it so the two cannot drift: the decision is the same decision, and both are the one place where "these are new bytes, so hash them" is the wrong answer.

Usage

.datom_resolve_document_sha(
  conn,
  name,
  data_sha,
  new_document_sha,
  change_type,
  current
)

Arguments

conn

A datom_conn object.

name

Set name.

data_sha

Canonical content hash (the storage address).

new_document_sha

SHA-256 of the payload bytes just written to the clone.

change_type

"metadata_only" or "full" (never "none").

current

The current metadata (from .datom_has_changes()), or NULL.

Details

Recomputing the hash from freshly emitted bytes while reusing the stored object records a hash of bytes nobody stored. Nothing fails at write time – it surfaces much later as a refused read of a valid version, when the integrity gate compares the stored payload against a hash taken from a different serialization of the same content. Sets reach that state far more easily than tables do: for a table it takes an arrow upgrade, while for a set an ordinary tag-value reorder is enough, because several payload spellings share one data_sha.

Cases, mirroring the parquet ones:

Value

List with document_sha (character or NULL) and upload (logical).


Resolve the Local Path for the Governance Clone

Description

Returns the explicit override if supplied. Otherwise, places the gov clone as a sibling of data_local_path named after the basename of gov_repo_url (stripping a trailing .git suffix). This ensures the gov clone directory name reflects the gov repo's own identity, not any specific data project.

Usage

.datom_resolve_gov_local_path(data_local_path, gov_repo_url, override = NULL)

Arguments

data_local_path

Absolute path to the local data repo directory.

gov_repo_url

GitHub URL of the governance repo (e.g., "https://github.com/org/acme-gov.git").

override

Optional explicit path. If non-NULL, returned as-is.

Value

Absolute path string for the gov clone.


Resolve Gov Clone Path with Store Defaults

Description

Convenience wrapper that derives a gov clone path from a datom_store: returns the store's explicit gov_local_path if set; otherwise derives a sibling-of-data default from gov_repo_url; otherwise returns NULL.

Usage

.datom_resolve_or_default_gov_path(store, data_local_path)

Arguments

store

A datom_store object.

data_local_path

Absolute path to the local data repo (used to compute the sibling default when no override is set).

Details

Centralises the three-arm pattern previously duplicated in datom_init_repo(), datom_clone(), and .datom_get_conn_developer().

Value

Character path string or NULL.


Resolve the parquet_sha to Record and Whether to Upload

Description

For a write that is not a no-op, decides which parquet_sha the new metadata should carry and whether the freshly-serialized parquet bytes need uploading. The caller performs the actual upload AFTER the git push (git push is the serialization point); this function only decides.

Usage

.datom_resolve_parquet_sha(
  conn,
  name,
  data_sha,
  new_parquet_sha,
  change_type,
  current
)

Arguments

conn

A datom_conn object.

name

Table name.

data_sha

Canonical content hash (the storage address).

new_parquet_sha

SHA-256 of the freshly-serialized parquet bytes.

change_type

"metadata_only" or "full" (never "none").

current

The current metadata (from .datom_has_changes()), or NULL.

Details

Cases:

This refines the design's literal step 7 (which gated on .datom_storage_exists()): a recorded parquet_sha is the precise thing we must not clobber, and its presence implies the object exists, so the history lookup subsumes the existence check with identical behavior and one fewer storage round-trip.

Value

List with parquet_sha (character or NULL) and upload (logical).


Resolve Data Location from Governance Store

Description

Reads projects/{project_name}/ref.json from the governance store and returns the current data location as a named list. Single read, no recursion, no chain-walking.

Usage

.datom_resolve_ref(gov_conn, project_name = NULL)

Arguments

gov_conn

A datom_conn-like object scoped to the governance store (i.e., root, prefix, client point to the governance store). Typically produced by .datom_conn_for(conn, "gov").

project_name

Project name string. Used to build the project-scoped storage key projects/{project_name}/ref.json.

Details

If the ref has previous entries, a deprecation-style warning is emitted to alert users that a migration occurred and old locations may sunset.

Value

A named list with root, prefix, region for the current data location.


Resolve Data Location from Local Gov Clone

Description

Reads projects/{project_name}/ref.json directly from a local gov clone on disk. Faster than storage reads, works offline, and reflects the last gov-clone refresh. Used for developer connections.

Usage

.datom_resolve_ref_from_clone(gov_local_path, project_name)

Arguments

gov_local_path

Absolute path to the local gov clone.

project_name

Project name string.

Value

A named list with root, prefix, region for the current data location.


Resolve Version to data_sha, Stored-Object Hash and Recorded Version

Description

Given metadata from .datom_read_metadata(), resolves a version spec to the corresponding data_sha (the storage address), the recorded stored-object integrity hash, and the version string as recorded. If version is NULL, resolves from the current metadata.json; if a metadata_sha string (or a prefix of one), looks it up in version_history.json.

Usage

.datom_resolve_version(
  metadata_list,
  version = NULL,
  name = "table",
  field = "parquet_sha"
)

Arguments

metadata_list

Return value of .datom_read_metadata().

version

NULL (current) or a metadata_sha string / prefix.

name

Artifact name (for error messages).

field

Which recorded stored-object hash to resolve: "parquet_sha" (a table's parquet) or "document_sha" (a set's payload).

Details

One function, two kinds, one field argument. A table's stored object is a parquet file pinned by parquet_sha; a set's is a JSON payload pinned by document_sha. The question is identical either way – which recorded hash pins the version I just resolved – so a second copy of this lookup would eventually disagree with this one about prefix matching or about what an absent hash means. The resolved hash comes back as object_sha regardless, because the caller already knows which field it asked for.

The object_sha may be NULL/"", and what that means is the caller's to decide, not this function's. For a table it is pre-cv1 metadata and tells .datom_read_parquet() to skip the integrity check – a grace for legacy metadata, not a gap in the current writer. For a set there is no legacy population, so the set read treats it as an error.

version is the version recorded for the resolved state, never recomputed: a pinned read echoes the matched history entry's own version string (so a caller who passed an 8-character prefix gets the full one back), and an unpinned read takes the current state's recorded version via .datom_recorded_current_version(). That helper returns NULL when the history records nothing matching the current document, which is a gap datom_validate() owns – a manufactured version would be a wrong statement rather than a missing one.

Value

Named list with data_sha (character), object_sha (character or NULL) and version (character or NULL) for the resolved version.


Create an S3 Client from Credentials

Description

Constructs a paws.storage::s3() client from credential values. Never stores raw credentials beyond the paws client object.

Usage

.datom_s3_client(
  access_key,
  secret_key,
  region = "us-east-1",
  endpoint = NULL,
  session_token = NULL
)

Arguments

access_key

AWS access key ID string.

secret_key

AWS secret access key string.

region

AWS region string (e.g. "us-east-1").

endpoint

Optional S3 endpoint URL. NULL for default AWS endpoint.

session_token

Optional AWS session token for temporary credentials.

Value

A paws.storage S3 client.


Delete All S3 Objects Under a Prefix

Description

Lists every key under {prefix}/datom/{prefix_key} and deletes in batches of up to 1000. A missing prefix is a no-op.

Usage

.datom_s3_delete_prefix(conn, prefix_key = NULL)

Arguments

conn

A datom_conn object.

prefix_key

Relative prefix (after ⁠prefix/datom/⁠).

Value

Invisibly, the count of deleted objects.


Download File from S3

Description

Downloads an S3 object and writes it to a local path. Creates parent directories if needed.

Usage

.datom_s3_download(conn, s3_key, local_path)

Arguments

conn

A datom_conn object.

s3_key

Relative S3 key (after ⁠prefix/datom/⁠).

local_path

Local file path (destination).

Value

Invisible TRUE on success.


Check if S3 Object Exists

Description

Uses a HEAD request for efficiency. Returns TRUE if the object exists, FALSE on 404/NoSuchKey. Any other error (403, network) is re-thrown.

Usage

.datom_s3_exists(conn, s3_key)

Arguments

conn

A datom_conn object.

s3_key

Relative S3 key (after ⁠prefix/datom/⁠).

Value

TRUE or FALSE.


List S3 Objects Under a Prefix

Description

Lists every key under {prefix}/datom/{prefix_key} and returns relative keys (relative to the datom namespace, i.e. with the ⁠prefix/datom/⁠ part stripped). Paginates via ContinuationToken.

Usage

.datom_s3_list_objects(conn, prefix)

Arguments

conn

A datom_conn object.

prefix

Relative prefix (after ⁠prefix/datom/⁠).

Value

Character vector of relative keys (may be empty).


Read and Parse JSON from S3

Description

Downloads an S3 object, reads it as text, and parses as JSON. Uses simplifyVector = FALSE to keep lists as lists (matching how .datom_s3_write_json() writes them).

Usage

.datom_s3_read_json(conn, s3_key)

Arguments

conn

A datom_conn object.

s3_key

Relative S3 key (after ⁠prefix/datom/⁠).

Value

Parsed R list.


Upload File to S3

Description

Reads a local file as raw bytes and uploads via put_object().

Usage

.datom_s3_upload(conn, local_path, s3_key)

Arguments

conn

A datom_conn object.

local_path

Local file path to upload.

s3_key

Relative S3 key (after ⁠prefix/datom/⁠).

Value

Invisible TRUE on success.


Write an R List to S3 as JSON

Description

Serializes data to JSON via jsonlite::toJSON() and uploads to S3.

Usage

.datom_s3_write_json(conn, s3_key, data)

Arguments

conn

A datom_conn object.

s3_key

Relative S3 key (after ⁠prefix/datom/⁠).

data

An R list to serialize to JSON.

Value

Invisible TRUE on success.


Which Members Does This Call Refer To?

Description

The plural selector both edit verbs share. .datom_find_member() resolves exactly one member and aborts on an ambiguous name, which is right for a fetch and only half of what an edit needs: an edit legitimately acts on many.

Usage

.datom_select_members(
  members,
  member = NULL,
  tags = NULL,
  version = NULL,
  version_arg = "version"
)

Arguments

members

The set's member list.

member

A name, a member record, a datom_link, or NULL for all.

tags

Optional label filter.

version

Optional version, or a prefix of one.

version_arg

The calling verb's name for version, for messages.

Details

Three routes, and the difference between them is how many members they can return:

What arrives What comes back
no member every member, narrowed by tags and version
a name exactly one, aborting when the name is ambiguous
a record or a link exactly the member carrying that id

An explicitly named member that is ambiguous aborts, and that is not in tension with the caller who sweeps: a sweep can honour "refresh everything" while skipping a name it cannot choose between, whereas a caller who named one member asked for something that cannot be done, so it is a user error and the narrowing arguments are what resolve it. The abort comes from .datom_find_member() rather than from a second copy of that message.

tags and version narrow a name or a sweep. Supplied beside a record or a link they are refused rather than ignored, because ignoring them would act on a different member than the one asked for and report success.

Value

An integer vector of positions in members, never empty.


Append One Member Record to a Set, as an Edit

Description

The three steps every addition takes – a ⁠$fetch⁠ link, the set's version forgotten, an add row in the edit log – in one place, so datom_add_member() and the set route of datom_sync() cannot drift apart in what they record. See points 5 and 6 of this file's header. It prints nothing: each caller says "nothing has been written" once, which for sync means once per call rather than once per added member.

Usage

.datom_set_add_record(x, record)

Arguments

x

A datom_set.

record

A member record, already validated and checked for clashes.

Details

The link goes through the shared factory, never inline, so no frame holding a connection lands on its parent chain.

Value

x, one member longer.


The Commit Message a Set Write Uses

Description

A set write commits ⁠Update {name}⁠, which says nothing in ⁠git log⁠. When the object being written carries an edit log and the caller passed no message, the default names what changed instead – every action in the log, so a chained ⁠update |> remove⁠ produces one message describing both.

Usage

.datom_set_commit_messages(name, message, edits)

Arguments

name

The set's name.

message

The caller's message, or NULL.

edits

The edit log carried by the object being written, or NULL.

Details

Two messages, because they go to two places. The subject is recorded as the version's commit_message, where one line is what datom_history() can show. The commit gets the subject plus the full list, with whole versions rather than prefixes: git is the durable record, so completeness belongs there rather than on screen.

An explicit message always wins, and a log of the wrong shape is ignored rather than trusted – it is an attribute, so a caller can put anything there.

Value

A list of history (a single line, or NULL to leave the existing default in place) and commit.


The Member List of a Set, or an Abort Naming What Was Passed

Description

Every verb in this file starts here, so "this is not a set" is reported once and identically rather than surfacing as a $ on a data frame returning NULL.

Usage

.datom_set_members(x, arg = "x")

Arguments

x

The value the caller passed.

arg

Argument name for the message.

Details

A set with no members is not an error. The writer refuses an empty member list, but the reader does not – a hand-built payload, or one from a newer datom, reads back with none – so every verb below has to have an answer for zero members.

Value

The member list, possibly empty.


Which Project a Set Reports Itself As Belonging To

Description

Two steps, not the three .datom_declared_project() uses: the set's own metadata document, then the connection's name. See the call site in datom_get_set() for why the manifest step is deliberately absent here.

Usage

.datom_set_project(current, conn)

Arguments

current

The set's metadata.json, as already read by the set read.

conn

The connection the set was read through.

Value

A single string, or whatever the connection carries.


Artifact Names Present in Storage

Description

Enumerates artifacts from a storage listing, by the one signal that identifies one: a {name}/.metadata/metadata.json object. Deliberately independent of the manifest, because the manifest is the document under suspicion whenever this is called.

Usage

.datom_storage_artifact_names(conn)

Arguments

conn

A datom_conn object.

Details

The clone-side equivalent is .datom_clone_artifact_names(). They are not interchangeable and neither can stand in for the other: a storage-only reader has no clone at all, and the clone can hold an artifact whose upload has not happened yet.

The listing returns FULL keys – including the ⁠{prefix}/datom/⁠ portion – while every other part of datom's business logic speaks in keys relative to the datom namespace root. Mixing the two shapes double-prefixes silently and does not error, so the root is stripped here, once, against the same builder the backends use.

Value

Character vector of artifact names, possibly empty. One storage listing, recursive.


Get Byte Size of a Single Storage Object

Description

Returns the byte size of the object at rel_key without reading its content. For S3 uses HEAD; for local uses fs::file_size(). Errors if the object is not found.

Usage

.datom_storage_byte_size(conn, rel_key)

Arguments

conn

A datom_conn object.

rel_key

Relative storage key (after ⁠{prefix}/datom/⁠).

Value

Numeric byte count.


Compute SHA-256 Hash of a Storage Object's Content

Description

For S3, downloads the raw bytes and hashes in memory. For local, hashes the file directly. Used by datom_storage_verify() in content mode.

Usage

.datom_storage_content_hash(conn, rel_key)

Arguments

conn

A datom_conn object.

rel_key

Relative storage key (after ⁠{prefix}/datom/⁠).

Value

Character SHA-256 hex string.


Delete governance.json Mirror from Data Storage

Description

Removes the governance.json mirror during project teardown. No-ops silently when the key is absent. Deletion is implemented via prefix-delete on the exact key path.

Usage

.datom_storage_delete_governance_json(conn)

Arguments

conn

A datom_conn for the data store.

Value

Invisible NULL.


Delete All Objects Under a Storage Prefix

Description

Removes every file under prefix/datom/{prefix_key} from storage. For S3 this lists then batch-deletes. For local it removes the directory. A missing prefix is a no-op (returns 0L). Pass prefix_key = NULL to delete the entire datom namespace for this connection.

Usage

.datom_storage_delete_prefix(conn, prefix_key = NULL)

Arguments

conn

A datom_conn object.

prefix_key

Relative prefix to delete under (after ⁠prefix/datom/⁠). NULL deletes the entire datom namespace root.

Value

Invisibly, the count of deleted objects.


Download File from Storage

Description

Download File from Storage

Usage

.datom_storage_download(conn, key, local_path)

Arguments

conn

A datom_conn object.

key

Relative storage key (after ⁠prefix/datom/⁠).

local_path

Local file path (destination).

Value

Invisible TRUE on success.


Check if Storage Object Exists

Description

Check if Storage Object Exists

Usage

.datom_storage_exists(conn, key)

Arguments

conn

A datom_conn object.

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

TRUE or FALSE.


List Objects Under a Storage Prefix

Description

Returns the keys of every object under {prefix}/datom/{prefix_arg}. Keys are returned in their full storage-key form (i.e. including the ⁠{prefix}/datom/⁠ portion), matching what .datom_local_list_objects() and .datom_s3_list_objects() return.

Usage

.datom_storage_list_objects(conn, prefix)

Arguments

conn

A datom_conn object.

prefix

Relative prefix to list under (after ⁠prefix/datom/⁠).

Value

Character vector of full storage keys (may be empty).


Read governance.json Mirror from Data Storage

Description

Returns the parsed list, or NULL when the key is absent. Aborts on any non-not-found storage error or on failed schema validation.

Usage

.datom_storage_read_governance_json(conn)

Arguments

conn

A datom_conn for the data store.

Value

Parsed list or NULL.


Read and Parse JSON from Storage

Description

Read and Parse JSON from Storage

Usage

.datom_storage_read_json(conn, key)

Arguments

conn

A datom_conn object.

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

Parsed R list.


Strip datom Namespace Prefix from a Full Storage Key

Description

Converts a full storage key (as returned by .datom_storage_list_objects()) to a relative key suitable for upload/download helpers (after ⁠{prefix}/datom/⁠).

Usage

.datom_storage_rel_key(full_key, conn)

Arguments

full_key

Full storage key string.

conn

The source datom_conn (provides prefix for stripping).

Value

Relative key string.


Upload File to Storage

Description

Upload File to Storage

Usage

.datom_storage_upload(conn, local_path, key)

Arguments

conn

A datom_conn object.

local_path

Local file path to upload.

key

Relative storage key (after ⁠prefix/datom/⁠).

Value

Invisible TRUE on success.


Write governance.json Mirror to Data Storage

Description

Writes content to .metadata/governance.json in the data store. Uses .datom_storage_write_json() dispatch (backend-neutral).

Usage

.datom_storage_write_governance_json(conn, content)

Arguments

conn

A datom_conn for the data store.

content

Named list from .datom_create_governance_json().

Value

Invisible NULL.


Write an R List to Storage as JSON

Description

Write an R List to Storage as JSON

Usage

.datom_storage_write_json(conn, key, data)

Arguments

conn

A datom_conn object.

key

Relative storage key (after ⁠prefix/datom/⁠).

data

An R list to serialize to JSON.

Value

Invisible TRUE on success.


Get Backend Type from Store Component

Description

Get Backend Type from Store Component

Usage

.datom_store_backend(component)

Arguments

component

A store component object.

Value

"s3" or "local".


Build a Store-Constructor Snippet for a Component

Description

Renders a copy/paste datom_store_local(...) or datom_store_s3(...) call string for a store component, for embedding in a generated README. Secrets are shown as placeholders.

Usage

.datom_store_constructor_snippet(component)

Arguments

component

A store component (datom_store_local, datom_store_s3, or datom_store_s3_creds).

Value

Character scalar — an R constructor call as text.


Get Region from Store Component

Description

Returns the AWS region for S3, NULL for local.

Usage

.datom_store_region(component)

Arguments

component

A store component object.

Value

Region string or NULL.


Get Root from Store Component

Description

Returns the storage root: bucket name for S3, directory path for local.

Usage

.datom_store_root(component)

Arguments

component

A store component object.

Value

Root string.


The commit_sha Storage Already Holds, by Version

Description

Nothing there and could not look are separate answers, and only one of them is safe to pass over in silence. The first write of an artifact has no stored history, which is ordinary and silent. A stored copy that exists and will not read is the opposite: the values only storage had are now unknown, and the caller is about to replace that file wholesale – so an entry git cannot attribute loses a good value. Collapsing the two into "no known values" makes the loss invisible in exactly the case where it is unrecoverable, which is why this returns the distinction rather than just a map.

Usage

.datom_stored_commit_shas(conn, name)

Arguments

conn

A datom_conn object.

name

Artifact name.

Details

The existence probe is what separates them. Its own failure counts as could not look, never as absence: an unreachable store cannot report that a file is missing.

A stored copy that reads but holds no usable pair is an absence, not a failure – the document was inspected and had nothing to contribute.

Value

A list with shas (named character vector, commit_sha named by version, empty when there are none) and unreadable (TRUE when storage holds a copy this call could not read).


Description

The one step that makes read-modify-write possible. A member record is payload-shaped – exactly id plus optional tags – and datom_get_set() adds a callable fetch to each one, which both .datom_validate_members() and the sv1 encoder refuse. Removing it here means neither of them needs a carve-out for a field that must never reach a payload.

Usage

.datom_strip_member_links(members)

Arguments

members

A member list.

Details

Only a function is dropped. A hand-built fetch = "junk" is left in place so the validator reports it; stripping by name would turn a typo into a silent success.

Value

The member list with any callable fetch element removed.


Coerce a Parsed-JSON String Set to a Character Vector

Description

A tag value arrives in one of three spellings, all meaning the same thing: a length-1 character vector ("output"), a longer character vector (c("safety", "efficacy")), or – after a JSON round trip with simplifyVector = FALSE – a list of length-1 strings. All three normalise to a character vector here, which is what makes a single string and a one-element array hash identically.

Usage

.datom_sv1_as_strings(v, what)

Arguments

v

The value to normalise.

what

Key path used in error messages (e.g. "tags$domain").

Details

character(0) and list() (the parsed form of ⁠[]⁠) both normalise to the empty set. Upstream, canonicalization drops a key whose value is empty – "no labels" is spelled by omitting the key – so the encoder should never meet one. It must not depend on that: an encoder whose correctness rests on an upstream rule breaks silently the day that rule moves.

Value

A character vector, possibly of length zero.


Hash Bytes for datom-sv1

Description

The single SHA-256 call of the datom-sv1 regime. Returns raw bytes rather than hex because every intermediate digest is concatenated into the next hash input; hex would double the width and put a text encoding in the identity path.

Usage

.datom_sv1_h(bytes)

Arguments

bytes

A raw vector.

Value

A raw vector of 32 bytes.


Render a Digest as Lowercase Hex

Description

Used for the two places the specification names hex: the collation key for member digests, and the final data_sha string. Byte order and lowercase-hex C-locale order agree (00-09 before ⁠0a⁠-⁠0f⁠, digits before letters in ASCII), so sorting either representation gives the same result – hex is named in the spec because it is what a reader can compare by eye.

Usage

.datom_sv1_hex(x)

Arguments

x

A raw vector.

Value

A character string of 2 * length(x) lowercase hex digits.


Encode a Map for datom-sv1

Description

⁠map(m) = h(0x03 || concat(str(k) || strset(m[k]) for k in sort(keys(m), radix)))⁠.

Usage

.datom_sv1_map(m, what = "map")

Arguments

m

A named list, or NULL.

what

Key path used in error messages.

Details

One encoder serves both slots of a member record – the id and the tags – so a fifth id field added later is just another key: no positional convention to maintain, and no absent-versus-empty question. id values are single strings, encoded as one-element string sets; enforcing "exactly these four keys, each single-valued" is validation's job, not the encoder's.

An absent map (NULL) and an empty map both encode as h(0x03). Writers never emit an empty map, but the encoder must not depend on that.

Value

A raw vector of 32 bytes.


Encode a Member Record for datom-sv1

Description

member(x) = h(0x04 || map(x.id) || map(x.tags)). Both slots are maps, so swapping content between them cannot collide, and a member with no tags encodes its tags slot as the empty map.

Usage

.datom_sv1_member(x, what = "member")

Arguments

x

A member record: a list with id and optionally tags.

what

Position label used in error messages.

Details

An unexpected field aborts. That is not grammar validation creeping in: a field the encoder ignored would be content that does not enter identity, so two payloads differing in it would share one data_sha and one storage address.

Value

A raw vector of 32 bytes.


Encode a Set Payload for datom-sv1

Description

set(p) = h(0x05 || map(p.tags) || concat(sort(unique(member(m)), radix))).

Usage

.datom_sv1_set(payload, what = "payload")

Arguments

payload

A list with members and optional set-level tags.

what

Position label used in error messages.

Details

Member digests are deduped and sorted, exactly like tag values: arrangement is presentation, not content. The producer of a member list is normally a script, so an insertion-order refactor must not mint a new version of a citable artifact.

A zero-member payload aborts, mirroring .datom_canonical_hash()'s refusal of a zero-row or zero-column table.

Value

A raw vector of 32 bytes.


Encode a String for datom-sv1

Description

str(s) = h(0x01 || utf8(s)). No length prefix and no terminator: the string is the entire hash input, so nothing follows it to be confused with.

Usage

.datom_sv1_str(s, what = "value")

Arguments

s

A length-1 character vector.

what

Key path used in error messages.

Value

A raw vector of 32 bytes.


Encode a String Set for datom-sv1

Description

⁠strset(v) = h(0x02 || concat(str(e) for e in sort(unique(v), radix)))⁠. Order and multiplicity are not identity: a multi-valued tag models simultaneous membership in several categories, which has no order and no notion of a repeated element.

Usage

.datom_sv1_strset(v, what = "value")

Arguments

v

A character vector, or a list of length-1 strings.

what

Key path used in error messages.

Details

The empty set is h(0x02) over an empty concatenation – a pinned golden.

Value

A raw vector of 32 bytes.


A Preview Frame Checked for Shape and Values, as Plain Text Columns

Description

Columns and values only, never where the frame came from: a subset or a hand-built frame is as good as the preview itself. Every check runs before any read.

Usage

.datom_sync_apply_frame(manifest)

Arguments

manifest

What the caller passed.

Details

Values are checked only on the rows apply acts on (new, changed); the others do nothing, so a hand-trimmed not_checked row is harmless. A short version_from would otherwise fail the exact comparison later and stop as stale, naming the wrong problem.

Value

A data frame of the six preview columns, each character.


Refuse a Row Whose Kind Disagrees With the Artifact

Description

Refuse a Row Whose Kind Disagrees With the Artifact

Usage

.datom_sync_check_kind(row, found)

Arguments

row

One preview row.

found

The kind the member or the snapshot records.

Value

Invisibly NULL; aborts with class datom_sync_kind_mismatch.


Which Context a Sync Call Is In: an Ordinary Repo or a Product Repo

Description

The two sync verbs do different jobs depending on the repo: an ordinary repo imports source files, a product repo maps its one set against source projects. This reads which, once per call, so every branch below it acts on one answer.

Usage

.datom_sync_context(conn)

Arguments

conn

A datom_conn object with a local path.

Details

Read from .datom/project.yaml, not from the connection, for the reason .datom_refuse_import_on_product() gives: the answer can authorise a write, and a hand edit or a pull can change the file after the connection was built. And the parse is gated – the file's declared format is checked before mode or set is read out of it.

Value

A list of product (TRUE for a mode: product repo) and set (the declared set name as written, possibly NULL). A repo with no config is reported as ordinary: the file path then fails with its own message about an uninitialised repo.


Sync Data-Side Metadata to Storage

Description

Mirrors the data repo's metadata to the data store so readers see current state: the manifest (.metadata/manifest.json) and each artifact's metadata ({name}/.metadata/metadata.json, version_history.json). Every artifact of either kind, not tables only – discovery is .datom_clone_artifact_names() and has been kind-agnostic since sets existed.

Usage

.datom_sync_data_metadata(conn, .confirm = TRUE)

Arguments

conn

A datom_conn object from datom_get_conn().

.confirm

If TRUE (default), requires interactive confirmation before proceeding. Set to FALSE for non-interactive use.

Details

It is not metadata-only, and the name understates it. For a set, a payload missing from storage is restored from the clone – see .datom_restore_set_payload() in this file for the three conditions on that. So this function can put content into storage, not just documents about content.

Data-only: governance files (dispatch.json, ref.json, migration_history.json) are not touched here. Governance sync is owned by the governance layer (gov_sync_dispatch()).

Two public routes reach this, and both get the restore: datom_write(conn) with no data and no name (the mirror-everything route), and datom_validate(fix = TRUE). Describing the restore as repair-only would leave a reader surprised to see it fire under a write verb.

Used after a failed upload, or by datom_validate(fix = TRUE), to bring storage back in line with the local data clone. Requires a developer connection with a local repo path.

Value

Invisibly, a list with repo_files (character vector of synced keys) and tables (list of per-table sync results).


Sync governance.json Storage Mirror from Git Copy

Description

Reads the git-canonical copy and overwrites the storage mirror. Call after a partial failure to repair a missing or stale storage mirror.

Usage

.datom_sync_governance_json(conn)

Arguments

conn

A datom_conn with path set to the local data git clone.

Value

Invisible NULL.


Sync Single Table Metadata to S3

Description

Sync Single Table Metadata to S3

Usage

.datom_sync_metadata(conn, name)

Arguments

conn

Connection object.

name

Table name.

Value

Summary of sync operation.


Does a Name Match a Sync Glob?

Description

The same glob rule the file scan applies to file names.

Usage

.datom_sync_name_matches(names, pattern)

Arguments

names

Artifact names.

pattern

A glob, "*" for everything.

Value

A logical vector.


Build the Member a new Row Adds, From Its Source

Description

Read through datom_member(), which confirms the version exists and records the project the artifact's own metadata declares. That project and the kind are then checked against the row: the connection's label is what routed the read, and nothing verifies a label.

Usage

.datom_sync_new_member(row, conn, tags)

Arguments

row

One new row.

conn

The connection for the row's project.

tags

Labels for the new member.

Value

A member record.


The Columns a Set Sync Preview Carries, in Order

Description

The Columns a Set Sync Preview Carries, in Order

Usage

.datom_sync_preview_cols()

Value

A character vector.


The Repo's Set As Stored, or an Empty One When It Has Never Been Written

Description

See point 1 of this file's header. The probe is on the set's current-state document, which is what datom_get_set() reads first. FALSE means the set has never been written; TRUE means read it, and any error from that read is the caller's to see. An error from the probe itself is never absence: an unreachable store cannot report that a file is missing.

Usage

.datom_sync_read_set(conn, name)

Arguments

conn

The product repo's developer connection.

name

The set's name.

Value

A datom_set.


Refuse Two Applied Rows for One Artifact

Description

A preview never produces them. Without this the second row would stop as stale, which names the wrong problem.

Usage

.datom_sync_refuse_duplicate_rows(todo)

Arguments

todo

The new and changed rows.

Value

Invisibly NULL; aborts with class datom_sync_manifest_duplicate_row.


Refuse Applied Rows Whose Project Has No Connection

Description

Only new and changed rows: they are the only ones apply acts on, and a full preview's not_checked rows belong by definition to projects not in sources, so checking every row would make an unedited preview impossible to apply.

Usage

.datom_sync_refuse_missing_source(todo, labels)

Arguments

todo

The new and changed rows.

labels

The project names of the source connections.

Value

Invisibly NULL; aborts with class datom_sync_source_missing.


Refuse a Preview the Set Has Moved Away From

Description

A changed row must find exactly one member for its artifact, at version_from; a new row must find none. Anything else means the set was edited after the preview was built, and acting anyway would leave a removed member removed, move a member the preview never showed, or add a second member for one artifact.

Usage

.datom_sync_refuse_stale(todo, hits, members, x_given)

Arguments

todo

The new and changed rows.

hits

For each row of todo, the positions of the set's members with that project and name.

members

The set's member list.

x_given

Whether the caller passed the set, which changes the remedy: the preview always compares against the stored set.

Value

Invisibly NULL; aborts with class datom_sync_manifest_stale.


Apply a Set Sync Preview to a Product Repo's Set

Description

The product-repo route of datom_sync(). See point 5 of this file's header: every check that needs no read comes first, then the set is read (when not passed), then the stale and kind checks against it, and only then one snapshot read per applied row.

Usage

.datom_sync_set_apply(conn, set_name, manifest, sources, tags, x)

Arguments

conn

The product repo's developer connection.

set_name

The set name .datom/project.yaml declares.

manifest

The preview, or any frame with its columns.

sources

One datom_conn or a list of them.

tags

Labels for members added by new rows.

x

A datom_set, or NULL to read the stored one.

Value

The edited datom_set.


The Product Repo's Declared Set Name, or an Abort Saying There Is None

Description

The Product Repo's Declared Set Name, or an Abort Saying There Is None

Usage

.datom_sync_set_name(set_name)

Arguments

set_name

What .datom/project.yaml declares under set.

Value

The set name.


Map a Product Repo's Set Against Its Sources

Description

The product-repo route of datom_sync_manifest(). Three kinds of read: the stored set (or none), one manifest per source, nothing per artifact.

Usage

.datom_sync_set_preview(conn, set_name, sources, pattern)

Arguments

conn

The product repo's developer connection.

set_name

The set name .datom/project.yaml declares.

sources

One datom_conn or a list of them.

pattern

Glob filtering source artifact names.

Value

The preview data frame; see datom_sync_manifest().


One Source's Artifacts, After Checking Its Label Against Its Own Manifest

Description

See point 3 of this file's header.

Usage

.datom_sync_source_artifacts(conn)

Arguments

conn

A source connection.

Value

The artifacts frame from .datom_current_artifacts().


One Member's Tags as Key/Value Pairs

Description

A member with no tags yields one pair of NA / NA, which is what keeps an untagged member visible in a listing rather than absent from it.

Usage

.datom_tag_pairs(tags)

Arguments

tags

A tag map, or NULL.

Details

A key whose value is empty also yields NA rather than no row. The writer drops such a key, so this reaches only a hand-built payload – and there the key IS in the document, so reporting the key with no value states what is there while dropping the row would not.

THE MAP IS READ BY POSITION, NEVER BY NAME, AND THAT IS THE WHOLE POINT OF THE FUNCTION. A tag map can carry the same key twice – jsonlite parses ⁠{"type": "output", "type": "baseline"}⁠ into two same-named elements, and a caller can write list(type = "a", type = "b") – and nothing on the read side refuses it, because a reader does not validate a tag map. tags[["type"]] returns the first match every time, so a by-name read reports one label twice and loses the other: the member lists a value it does not have, vanishes from a branch it belongs under, and cannot be found by the label the document says it carries. Verified in all three verbs before this was positional.

That is the file header's one-expander rule reappearing on the key axis. Having one expander closed the silent-first-value spelling on the value axis; reading that expander's own input by name reopened the identical failure one level up.

Duplicate keys are therefore treated exactly as one multi-valued key would be, which is also what they mean. Identical pairs are not collapsed across duplicate keys: the read reports what the document holds, and deduplicating here would be a reader canonicalizing.

Value

A data frame of key and value, at least one row.


A Tag Value as a Plain Character Vector

Description

This is only possible because the tag grammar is text-only. .datom_validate_tag_map() refuses numbers, booleans, null and nesting, so a tag value is a character vector and the value column of a member listing is a plain character column with no list-column anywhere. If the grammar ever widened, this function and the long format above it are what would have to change.

Usage

.datom_tag_values(v)

Arguments

v

A tag value.

Details

The two shapes handled are the two a value legitimately arrives in: a character vector, and the list-of-length-1-characters a JSON array parses as. A missing value is dropped rather than carried, because it states no label and no datom write can produce one – carried through, it would become a branch named NA.

Value

A character vector, possibly empty.


Tidy a Set Payload

Description

The silent half of canonicalization: every spelling that states the same fact is reduced to one, and nothing here is an error. Covers the set-level tag map, each member's id key order, and each member's tag map.

Usage

.datom_tidy_set_payload(payload)

Arguments

payload

A list with members and optional set-level tags.

Details

Member order and member deduplication are not here, because they need the identity encoder and so can only run once validation has established that every value is encodable – see .datom_order_set_members().

Key order is canonicalized at three levels, not two: the set's own tag map, each member's id, and each member record's own id / tags pair. Stopping at the second leaves one spelling uncanonical for no reason – the encoder reaches both member slots by name, so the two orders hash identically and serialise differently.

An empty tag map has its key removed rather than set to NULL, at both levels. jsonlite writes a NULL element as {}, and "tags": {} is the one spelling a writer must never emit: the hash cannot tell it from an absent map, so nothing would fail, and the stored file would carry an empty object in every untagged member forever.

Value

The tidied payload.


Tidy a Tag Map

Description

Radix-sorts the keys, drops a key whose value is empty, and tidies each value. Never aborts: a malformed map is passed through for the validator to report.

Usage

.datom_tidy_tag_map(tags)

Arguments

tags

A named list, or NULL.

Details

Radix sort throughout, i.e. C-locale byte order, so the canonical form does not depend on the machine's collation – the same reason the identity hash sorts that way.

Value

The tidied map, or NULL when nothing is left.


Tidy One Tag Value

Description

Sorts and dedupes a tag value, and normalises the three spellings of a string set into one character vector so that auto_unbox = TRUE writes a single label as a bare string and only a genuine multi-label value as an array.

Usage

.datom_tidy_tag_value(v)

Arguments

v

A tag value.

Details

Anything this build does not recognise as text is returned untouched. That is what keeps tidying from aborting: sort() on a list or a function fails with a base-R message that names nothing, whereas leaving the value alone hands it to the validator, whose message names the key and the allowed types. Tidy what you can, refuse the rest – in that order, and never the reverse.

A missing value is left alone for the same reason: NA has no text meaning, so it is a refusal rather than a tidy case.

Value

A sorted, deduplicated character vector, or v unchanged.


Validate GitHub PAT

Description

Calls GitHub GET /user to verify the PAT is valid.

Usage

.datom_validate_github_pat(pat, api_url = "https://api.github.com")

Arguments

pat

GitHub personal access token.

api_url

GitHub API base URL (default "https://api.github.com").

Value

A list with login and id.


Validate a Member List

Description

Checks that members is a list of member records, each an id of exactly project, name, kind, version – all single non-empty strings, with kind one of "table" or "set" – plus an optional tags map. Aborts naming the first offending member, with a remedy pointing at datom_member().

Usage

.datom_validate_members(x)

Arguments

x

Value to validate: a list of member records, or NULL.

Details

This validator sees one member at a time, so two payload-level cases are deliberately not here and belong to the set write, which is the only place that sees a whole payload:

Set-level tags never pass through here at all; the write validates those with .datom_validate_tag_map() directly.

Value

Invisibly TRUE.


Validate a datom Table Name

Description

Checks that a table name is filesystem-safe and S3-safe. Returns the name invisibly on success, errors with a clear message on failure.

Usage

.datom_validate_name(name)

Arguments

name

Character string to validate as a table name.

Value

Invisible name on success.


Validate parents Field Structure

Description

Checks that parents is either NULL or a list of entries each containing non-empty string fields source, table, version, and data_sha. WHERE an entry carries a non-NULL, non-empty source_lineage field, it is validated via .datom_validate_source_lineage(). Aborts with a cli error pointing to the first invalid entry.

Usage

.datom_validate_parents(x)

Arguments

x

Value to validate.

Value

Invisibly TRUE if valid.


Validate a Caller-Supplied Relative Storage Key

Description

Guards a whole key string that a caller composed, as opposed to one datom built itself from validated parts. The internal key builders in R/utils-path.R need no such check: .datom_validate_name() admits only ⁠[a-zA-Z0-9_ ()-]⁠ and .datom_validate_sha() only hex, so their output cannot contain a .. segment or a ⁠datom/⁠ segment. A key arriving through a public export has had no such filtering.

Usage

.datom_validate_rel_key(key, arg = "key")

Arguments

key

Value to validate as a relative storage key.

arg

Name of the calling argument, used in the error message.

Details

Two distinct failures are caught:

Value

Invisible key on success. Aborts otherwise.


Validate S3 Store Connectivity

Description

Checks bucket access via HeadBucket. This validates both credentials and bucket existence/permissions in a single call.

Usage

.datom_validate_s3_store(access_key, secret_key, session_token, region, bucket)

Arguments

access_key

AWS access key ID.

secret_key

AWS secret access key.

session_token

Optional session token.

region

AWS region.

bucket

Bucket name.

Value

Invisible TRUE on success.


Validate a SHA-Like Input (Version / data_sha)

Description

Ensures a user-supplied SHA-like string is 6-64 lowercase hex characters. Used to guard values that get spliced into a storage key ({table}/{sha}) – on the local backend an unvalidated value like "../../x" would escape the namespace via fs::path(). The 6-char minimum still covers the short prefixes .datom_resolve_version() intentionally accepts.

Usage

.datom_validate_sha(x, arg = "version")

Arguments

x

Value to validate.

arg

Name of the calling argument, used in the error message.

Value

Invisible x on success. Aborts otherwise.


Validate source_lineage Field Structure

Description

Checks that source_lineage is either NULL or a list of entries each containing non-empty string fields project, table, and version_sha. Extra fields are allowed (pass-through). Aborts with a cli error pointing to the first invalid entry.

Usage

.datom_validate_source_lineage(x)

Arguments

x

Value to validate.

Value

Invisibly TRUE if valid.


Validate a Tag Map

Description

The tag grammar, in one place, shared by datom_member() (its tags argument), the member validator (each member's tags), and the set write (set-level tags). A tag map is a named list whose values are UTF-8 strings or arrays of them – no numbers, booleans, null, or nesting.

Usage

.datom_validate_tag_map(tags, what = "tags", remedy = NULL)

Arguments

tags

A named list, or NULL (no tags).

what

Label used in error messages, e.g. "tags" or "members[[2]]$tags".

remedy

Optional cli bullet appended to every abort.

Details

Per-value type checking delegates to .datom_sv1_as_strings(), the same coercion the hash encoder uses, rather than restating its rules. That is deliberate: two copies of "what counts as text here" would eventually disagree, and the encoder's messages already name the offending key and the allowed types. What this function adds on top is the empty-label refusal, which the encoder does not make – there, "" hashes as an ordinary label.

An empty value is not refused, because it is a tidy case rather than an error: call .datom_drop_empty_tags() first, which every caller does.

Value

Invisibly TRUE.


Verify a Single Storage Object

Description

Checks that the object at rel_key in to_conn matches the one in from_conn. Returns a named list with key, ok (logical), and issue (character or NA_character_).

Usage

.datom_verify_one(from_conn, to_conn, rel_key, mode)

Arguments

from_conn

Source datom_conn.

to_conn

Destination datom_conn.

rel_key

Relative key (after ⁠{prefix}/datom/⁠).

mode

"structural" or "content".

Value

Named list: key, ok, issue.


Say Which Commit Links Went Unrecorded, and Why

Description

Split out so the wording lives next to the reasoning rather than inside a branch. A warning rather than a refusal, and the reason is that refusing would deadlock the only route out: the repair verb goes through the same helper, so a stored history that will not parse could never be replaced. That file is a projection for git-less readers and rebuilding it is exactly what the repair is for – what must not happen is rebuilding it in silence.

Usage

.datom_warn_commit_shas_lost(name, lost)

Arguments

name

Artifact name.

lost

Versions left with no commit recorded.

Details

It says that the unreadable copy is being replaced, because of which cause is the likelier one. The two are indistinguishable here, but a reachable store holding bad bytes is more plausible than one that refuses a read and accepts a write – and in that case this very operation overwrites the evidence. Somebody who would have gone looking should be told it will not be there. Worded as what this write does rather than as a completed fact: the message is raised before the upload, so a write that then fails leaves the bad copy in place.

Value

Invisibly NULL.


Say That the Artifact Index Was Reconstructed

Description

One warning per rebuild, carrying a condition class so a caller – or a test – can count them rather than match on wording.

Usage

.datom_warn_manifest_rebuilt(source, reason, declared, n)

Arguments

source

Which copy of the manifest was rebuilt.

reason

"schema" (the document declares a format above this build) or "shape" (no artifact list this build can reach).

declared

The version the document declared, for the schema reason.

n

How many artifacts the rebuild found.

Details

It names what happened, why, and what to do, in that order. The "why" is the part a user cannot work out for themselves: a manifest whose artifact list has moved somewhere this build cannot see looks exactly like an empty repo, and the whole point of warning is that this session's answers came from a reconstruction rather than from the recorded index.

Value

Invisibly NULL.


Write governance.json to Local Git Clone

Description

Writes content to {path}/.datom/governance.json. The directory must already exist (created during datom_init_repo() or datom_repo_attach_governance()).

Usage

.datom_write_governance_json_local(path, content)

Arguments

path

Absolute path to the root of the local data git clone.

content

Named list from .datom_create_governance_json().

Value

Invisible NULL.


Write Metadata Files to Git and S3 (Legacy Wrapper)

Description

Calls .datom_write_metadata_local() then .datom_push_metadata_s3(). Kept for backward compatibility. Does NOT commit or push.

Usage

.datom_write_metadata(conn, name, metadata, metadata_sha, message = NULL)

Arguments

conn

A datom_conn object (must be developer with path).

name

Table name.

metadata

Named list for metadata.json.

metadata_sha

SHA of the metadata (the datom "version").

message

Commit message (stored in version_history entry).

Details

It makes no commit, so it has no commit_sha to hand on, and that matters for anything asserted about the history it produces: the stored entries carry a commit only where one can be worked out from git. A test that means "every stored entry names its commit" has to drive a real write.

Value

Invisible list with metadata_sha, git_paths, and s3_keys.


Write Metadata Files Locally

Description

Writes metadata.json and appends to version_history.json in the local git repo. Does NOT commit, push, or touch S3 — the caller handles those.

Usage

.datom_write_metadata_local(
  conn,
  name,
  metadata,
  metadata_sha,
  message = NULL,
  original_file_sha = NULL
)

Arguments

conn

A datom_conn object (must be developer with path).

name

Table name.

metadata

Named list for metadata.json.

metadata_sha

SHA of the metadata (the datom "version").

message

Commit message (stored in version_history entry).

original_file_sha

SHA of the source file for imported tables; NULL for derived.

Value

Invisible list with metadata_sha and local paths written.


Check if Object is a Store Component

Description

Returns TRUE for any datom store component type (datom_store_s3, future datom_store_local, etc.).

Usage

.is_datom_store_component(x)

Arguments

x

Object to test.

Value

TRUE or FALSE.


Add One Member to a Set

Description

Declares one member and appends it to a set – an empty one from datom_assemble_set() or one read back with datom_get_set() – validating it immediately: the artifact must exist at the version given, and its labels must be well formed. The member record is built through the same path datom_member() uses, so a set assembled this way is byte-identical to the same set passed as a list.

Usage

datom_add_member(x, member, version = NULL, tags = NULL, conn = NULL)

Arguments

x

A datom_set, from datom_assemble_set() or datom_get_set().

member

The member to add: an artifact name, a member record, or a link.

version

The version to pin, when member is a name. Required there; refused beside a record or a link.

tags

Optional named list of text labels for this member, when member is a name. Refused beside a record or a link, which carry their own.

conn

A datom_conn from datom_get_conn() for the project a name is looked up in. Required for a name; not used for a record or a link.

Value

The set, one member longer, with its version and data_sha emptied and the addition appended to its datom_edits attribute – or unchanged, when the member was already in it with the same labels.

Naming a member

member accepts the three shapes a caller holds, and the second and third are not merely convenient:

What you pass What it means
a name look this artifact up through conn
a member record use it as given -- from datom_member(), or from a set read back
a link use the member it points at -- x$members[[i]]$fetch, or a leaf of datom_structure_members()

A name is looked up in one project's storage: the project conn is for. A set holds no connection, so a name always needs conn, whichever project it is in. A record or a link carries its own resolved pointer and needs no connection at all:

datom_assemble_set(conn_a) |>
  datom_add_member("dm", v1, conn = conn_a) |>      # this project, by name
  datom_add_member("ae", v2, conn = conn_b) |>      # another project, by name
  datom_add_member(datom_member(conn_b, "vs", v3))  # another project, as a record

A link is how a consumer cites what they used. Someone holding only a projection of a set – a leaf of datom_structure_members() – can add exactly the version they read to a new set, without reconstructing the pointer.

version and tags describe a member given by name. Beside a record or a link they are refused rather than ignored, because a record already carries its own version and labels and a second set of them could only disagree.

Why the version is required

A member pins one exact version, and there is no way to ask for "whatever is current". Inferring current would make a build script produce a different set on each run from byte-identical source. Pinning is what makes the artifact immutable; requiring the pin is what makes the code reproducible. List versions with datom_history().

Every add is an edit

Adding a member behaves like datom_update_members() and datom_remove_members(), whether the set was just assembled or read back:

Adding the same member twice

The two cases differ, and they differ the same way they differ at the write:

So the member count a set reports is the count the write will produce. Two different versions of one artifact are two members, and both are kept.

See Also

datom_assemble_set() to start a set, datom_get_set() to read one back, datom_member() to build a record on another connection, datom_update_members() and datom_remove_members() for the other edits.

Examples

# Adding by name needs a live connection, so the runnable example lives on
# datom_assemble_set(), which shows the whole pipe.
print(names(formals(datom_add_member)))

Start Assembling a Set

Description

Returns an empty set, to be filled in with datom_add_member() and written with datom_write_set():

Usage

datom_assemble_set(conn, name = NULL, tags = NULL)

Arguments

conn

A datom_conn from datom_get_conn() for the product repo. Its project name is recorded as the set's project; nothing else is kept.

name

The set's name. NULL (the default) leaves it to the write, which takes the name the repo declares under ⁠set:⁠ in .datom/project.yaml – the usual case, since one repo holds one set.

tags

Optional named list of set-level text labels, e.g. a description.

Details

datom_assemble_set(conn, tags = list(description = "ADaM datasets")) |>
  datom_add_member("adsl", v_adsl, tags = list(type = "output"),
                   conn = conn) |>
  datom_add_member("dm", v_dm, tags = list(type = "input"),
                   conn = conn_src) |>
  datom_write_set(conn = conn)

The equivalent single call – a list() of datom_member() results passed to datom_write_set() – remains fully supported and is the better fit for a build script. What this path adds is where an error surfaces: a malformed member aborts on the line that declared it and names that member, instead of aborting once the whole list has been assembled and indexed.

Value

A datom_set with no version and no members.

One kind of set

What comes back is a datom_set, the same kind of object datom_get_set() returns, only with no version yet. So every verb that takes a set takes this one: datom_list_members(), datom_update_members(), datom_write_set() and the rest.

It holds no connection. A connection may carry a credential, so it is passed on each call that needs one rather than kept in a value that can be printed or saved: ⁠conn =⁠ on datom_add_member() for a member given by name, and conn on datom_write_set(). The connection given here is used only to record which project the set belongs to; the write checks that it is the project it is written into.

Set-level tags

Supplied here rather than by a third verb, because they are facts about the collection rather than about any member. Editing them later is plain R – x$tags$description <- "..." – and the same grammar applies as to a member's tags: text only, one label or several.

See Also

datom_add_member() to add one member, datom_write_set() to write the result, datom_member() for the single-call form.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  # A product repo declares itself as one and names the single set it owns.
  datom_init_repo(file.path(tmp, "repo"), "example_project", store,
                  mode = "product", set = "example_product")

  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  datom_write(conn, data = datom_example_data("lb"), name = "lb")

  v_dm <- datom_history(conn, "dm")$version[1]
  v_lb <- datom_history(conn, "lb")$version[1]

  x <- datom_assemble_set(
    conn,
    tags = list(description = "Example product for STUDY-001")
  ) |>
    datom_add_member("dm", v_dm, tags = list(type = "input"), conn = conn) |>
    datom_add_member("lb", v_lb, tags = list(type = "output"), conn = conn)

  print(datom_list_members(x))
  x |> datom_write_set(conn = conn)

  unlink(tmp, recursive = TRUE)
}

Check Whether a Table Can Be Hashed by datom

Description

Pre-flight check for the datom table contract. Reports, per column, whether datom_write() can hash it and – when it cannot – exactly what to do about it. Run this before a write to fix a table in one pass instead of discovering offenders one error at a time.

Usage

datom_check_hashable(data)

Arguments

data

A data frame to check.

Details

datom identifies a table version by a canonical hash of its contents (data_sha), which requires every column to be a supported type: logical, integer, double, character, factor, Date, POSIXct, difftime/hms, data.table::ITime/IDate, bit64::integer64, or a labelled vector over one of those. List columns (including nested data frames, blobs, and POSIXlt), complex, raw, sf geometry, units, and zoo/chron columns are refused with specific advice.

The advice printed here is the same single-source recourse text datom_write() would abort with, so the two can never disagree.

Value

Invisibly, a data frame with one row per column of data and columns:

column

Column name.

class

Collapsed class string, or typeof() when unclassed.

status

"ok" or "unsupported".

recourse

NA when ok, otherwise how to make the column hashable.

See Also

datom_write()

Examples

# A clean table: every column is a supported type
clean <- data.frame(
  id = 1:3,
  score = c(1.5, 2.5, 3.5),
  label = c("a", "b", "c"),
  grp = factor(c("x", "y", "x")),
  day = as.Date(c("2026-01-01", "2026-01-02", "2026-01-03"))
)
datom_check_hashable(clean)

# An offending table: a list column and a complex column
messy <- data.frame(id = 1:2)
messy$notes <- list(c("a", "b"), "c")
messy$z <- c(1 + 2i, 3 + 4i)
report <- datom_check_hashable(messy)
report[report$status == "unsupported", c("column", "recourse")]


Clone a datom Repository

Description

Clones a remote datom repository and returns a connection. This is the recommended way for teammates to join an existing datom project – it wraps git2r::clone() and immediately returns a ready-to-use datom_conn.

Usage

datom_clone(path, store, ...)

Arguments

path

Local path to clone into.

store

A datom_store object (from datom_store()). Must have data_repo_url set and role "developer" (i.e., github_pat provided).

...

Additional arguments passed to git2r::clone().

Details

When store$gov_repo_url is set the governance repo is also cloned (or verified if it already exists locally). An existing clone with uncommitted changes causes an error to avoid surprising state.

Value

A datom_conn object (developer role).

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)

  # A teammate joins the project from the remote alone.
  conn <- datom_clone(path = file.path(tmp, "teammate"), store = store)
  print(datom_list(conn))

  unlink(tmp, recursive = TRUE)
}

Monthly Cutoff Dates for Example Study

Description

Returns a named vector of monthly cutoff dates for STUDY-001, useful for simulating EDC data evolution in examples.

Usage

datom_example_cutoffs()

Value

Named character vector with entries month_1 through month_6.

Examples

datom_example_cutoffs()
# month_1    month_2    month_3    month_4    month_5    month_6
# "2026-01-28" "2026-02-28" ...


Load Example Clinical Trial Data

Description

Returns one of five small, made-up tables from a simulated clinical trial of 48 subjects: demographics, exposure (dosing), lab results, adverse events or vital signs. Set cutoff_date to get the data as it stood on that date, which mimics a new data delivery each month; datom_example_cutoffs() lists the dates the examples use.

Usage

datom_example_data(
  domain = c("dm", "ex", "lb", "ae", "vs"),
  cutoff_date = NULL
)

Arguments

domain

One of "dm" (demographics, 48 rows), "ex" (exposure, 48 rows), "lb" (labs, 720 rows: 3 visits x 5 tests per subject), "ae" (adverse events, ~80 rows), or "vs" (vital signs, 432 rows: 3 visits x 3 tests per subject, taken on the same dates as the labs).

cutoff_date

Optional date string ("YYYY-MM-DD") to filter rows whose primary date column is on or before this date, simulating a point-in-time EDC extract. The date column used per domain: RFSTDTC (dm), EXSTDTC (ex), LBDTC (lb), AESTDTC (ae), VSDTC (vs).

Details

The data simulates STUDY-001, a Phase II study enrolling over six months; table and column names loosely follow SDTM.

Value

A data frame.

Examples

# Full demographics
dm <- datom_example_data("dm")

# Month-3 snapshot (subjects enrolled by 2026-03-28)
dm_m3 <- datom_example_data("dm", cutoff_date = "2026-03-28")

# Labs collected through Month 3
lb_m3 <- datom_example_data("lb", cutoff_date = "2026-03-28")


Get the Data Behind One Member of a Set

Description

Returns what one member of a set points at, at the exact version the set records: a data frame for a table, another set for a set. Name the member, as in datom_fetch_member(conn, x, "dm"), and pass a connection to that member's own project.

Usage

datom_fetch_member(conn, x, member, tags = NULL, version = NULL)

Arguments

conn

A datom_conn from datom_get_conn(), scoped to the member's project.

x

A datom_set from datom_get_set().

member

The member to fetch: its name, a member record, or a link.

tags

Optional named list of labels narrowing an ambiguous name, e.g. list(release = "baseline"). A member matches when it carries every label listed.

version

Optional version, or a prefix of one, narrowing an ambiguous name.

Details

x$members[[i]]$fetch(conn) does the same thing.

This is kind dispatch at the member level, which is the only level it belongs at. Iterating members, a caller cannot know each one's kind in advance; at the top level they named one artifact they chose, which is why datom_read() and datom_get_set() stay separate verbs.

Value

Whatever the member points at: a data frame for a table member, a datom_set for a set member.

Naming a member

member accepts the three shapes a caller actually holds, so a console call and a loop use one verb:

What you pass Where it comes from
a name you read the set and know what you want
a member record datom_member(), or x$members[[i]]
a link x$members[[i]]$fetch, or a leaf of datom_structure_members()

A name is not a key. The same artifact at two versions is a legal pair of members – a current table beside a locked baseline, say – so an ambiguous name aborts and lists the candidates rather than answering with the first. Narrow with tags, which is the navigation axis, or pin one exactly with version. Both narrow a name only; supplied beside a record or a link they are refused rather than quietly ignored.

Which connection to pass

The one for the member's project. Access in datom is per project and not conjunctive: reading a set needs the set's project only, and resolving a member is a separate, deliberate step. Same-project members resolve through the connection you already have.

A member's project is not checked against the connection's before the fetch, and that is deliberate: a connection's project name is a label the caller supplied and nothing compares it against the repo, so a mismatch is ordinary rather than wrong. When a fetch fails and the two names differ, the error says which project the member's own writer recorded, so the ordinary cause is named instead of presenting as a missing object.

See Also

datom_list_members() to see every member and its labels, datom_structure_members() for a navigable view.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  # A product repo declares itself as one and names the single set it owns.
  datom_init_repo(file.path(tmp, "repo"), "example_project", store,
                  mode = "product", set = "example_product")

  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  datom_write(conn, data = datom_example_data("lb"), name = "lb")

  members <- list(
    datom_member(conn, "dm", datom_history(conn, "dm")$version[1],
                 tags = list(type = "input")),
    datom_member(conn, "lb", datom_history(conn, "lb")$version[1],
                 tags = list(type = "output", domain = c("safety", "labs")))
  )
  datom_write_set(conn, members)

  x <- datom_get_set(conn, "example_product")

  # By name, and the same fetch by the link the read already put on it.
  print(head(datom_fetch_member(conn, x, "dm")))
  print(head(datom_fetch_member(conn, x, x$members[[1]]$fetch)))

  unlink(tmp, recursive = TRUE)
}

Get a Pointer to a datom Project

Description

Returns a pointer to the project, called a connection (conn): a record of which project you are working on, where its data is kept, and whether you can write. Almost every other datom function takes it as its first argument. Nothing stays open; it only checks once that the storage (and, for a developer, the GitHub repository) can be reached. Developers pass path (their local copy) and store; readers, who have no local copy, pass store and project_name.

Usage

datom_get_conn(path = NULL, store = NULL, project_name = NULL, endpoint = NULL)

Arguments

path

Path to datom repository. If provided, reads config from .datom/project.yaml.

store

A datom_store object. Required for all connections. The data component provides bucket, prefix, region, and credentials.

project_name

Project name. Required for readers (no local repo). Ignored when path is provided (read from yaml).

endpoint

Optional S3 endpoint URL (e.g., for S3 access points). NULL for default.

Details

Developer (local repo + store): provide path and store. Reads project identity from .datom/project.yaml; uses store for credentials and S3 config. Cross-checks bucket/prefix between yaml and store.

Reader (no local repo): provide store and project_name. Store provides everything.

Value

A datom_conn object.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)

  # Developer: local repo plus store.
  conn <- datom_get_conn(path = file.path(tmp, "repo"), store = store)
  print(conn)

  # Reader: store plus project name, no local repo.
  reader_store <- datom_store(data = datom_store_local(file.path(tmp, "storage")))
  print(datom_get_conn(store = reader_store, project_name = "example_project"))

  unlink(tmp, recursive = TRUE)
}

Show a Table's Original Sources or Direct Inputs

Description

Answers "where did this table come from?" from the table's own record. By default (depth = "source") it lists the original imported tables at the start of the chain, skipping the tables in between; depth = "parents" lists only its direct inputs, one step back. Works for any version, and with reader connections.

Usage

datom_get_lineage(conn, name, version = NULL, depth = c("source", "parents"))

Arguments

conn

A datom_conn object from datom_get_conn().

name

Table name.

version

Optional metadata_sha (datom version). If NULL, reads current metadata. If provided, fetches the versioned metadata snapshot.

depth

One of "source" (default) or "parents".

Details

It needs access to this table's project only, not to the projects its sources live in.

The two fields answer different questions:

Value

For depth = "source": the table's recorded source_lineage – a list of source-table descriptors (each with project, table, version_sha), or NULL if the field is absent. For depth = "parents": list of parent entries (each with source, table, version, data_sha), or NULL if no lineage is recorded.

See Also

datom_get_parents() for a direct shorthand for the "parents" depth.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  print(datom_get_lineage(conn, "dm", depth = "parents"))
  print(datom_get_lineage(conn, "dm", depth = "source"))

  unlink(tmp, recursive = TRUE)
}

Get Parent Lineage for a Table

Description

Reads the parents field from a table's metadata. Returns the lineage entries recorded at write time by datom_write(). For imported tables or derived tables with no recorded lineage, returns NULL.

Usage

datom_get_parents(conn, name, version = NULL)

Arguments

conn

A datom_conn object from datom_get_conn().

name

Table name.

version

Optional metadata_sha (datom version). If NULL, reads current metadata. If provided, fetches the versioned metadata snapshot from S3.

Value

List of parent entries (each with source, table, version, data_sha), or NULL if no lineage is recorded. The data_sha field is the parent's authoritative data SHA recorded via datom_parent(), and together with source and version is sufficient to select the parent's project connection and its pinned version.

See Also

datom_get_lineage() for a unified interface that also exposes the transitive source_lineage field via depth = "source".

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  dm <- datom_example_data("dm")
  datom_write(conn, data = dm, name = "dm")
  datom_write(
    conn,
    data    = dm[dm$SEX == "F", ],
    name    = "dm_female",
    parents = list(datom_parent(conn, "dm", datom_history(conn, "dm")$version[1]))
  )
  print(datom_get_parents(conn, "dm_female"))

  unlink(tmp, recursive = TRUE)
}

Read a datom Set

Description

Returns a set's members and labels, at its current version or a past one. It reads no table data: to get the data behind a member, use datom_fetch_member(). Works with reader connections.

Usage

datom_get_set(conn, name, version = NULL)

Arguments

conn

A datom_conn object from datom_get_conn(), scoped to the set's project. A storage-only connection with no git clone is enough.

name

The set's name.

version

Optional version (metadata_sha, or a prefix of one). NULL reads the current version.

Details

Reading a set requires access to the set's own project only. A member is a pointer, and resolving it is a separate, deliberate step – so a 50-member product is readable by someone entitled to none of its members.

Value

A datom_set: a list of name, project, version (possibly NULL), data_sha, tags and members.

What comes back

References and labels, and no data at all – which is why the verb is get rather than read.

A datom_set: name, project, version, data_sha, tags and members. The four identifying facts are there so that a caller who passed version = NULL can still say which version they got, because a set exists to be cited. version is the version recorded in the history, so an 8-character prefix goes in and the full version comes back.

members is a flat, unnamed list in payload order. Not name-keyed, and the reason is not style: the same artifact at two different versions is a legal pair of members, two projects may both hold a dm, and R's $ partial-matches on lists – so a name-keyed list would answer plausibly and wrongly. The unique key is the full id.

One of the four identifying facts has a limit worth knowing before you cite it:

project is the name the set's own metadata records – the declaration of the repo that wrote it, not the name on your connection. It falls back to the connection's name only for a set written by a datom that predates the field, which no released build ever was. Each member carries its own recorded id$project for the same reason, resolved through a slightly longer route because that value is durable and hashed rather than displayed.

Resolving a member

Each member is id (project, name, kind, version), its optional tags, and fetch:

x <- datom_get_set(conn, "study001-adam")
dm <- x$members[[1]]$fetch(conn)

fetch resolves whatever the pointer points at: a table member yields data, a set member yields another datom_set. A link pins the version it was read at – it is a citation, not a subscription, so it never drifts to the latest. Pass a connection scoped to the member's own project; same-project members resolve through the connection you already have.

Two reads of the same set are not identical(), because closures compare by environment. Compare m[c("id", "tags")] instead, or use identical(a, b, ignore.environment = TRUE).

Integrity, and what is not rechecked

The stored payload is verified against the recorded document_sha before it is parsed, and a version that records no document_sha is an error rather than a skipped check. data_sha is not recomputed: it is the address the payload was fetched from, so it would catch nothing the byte hash did not, and it would refuse a payload written by a newer datom. Nothing in the payload is re-canonicalized – what you are shown is what was cited.

See Also

datom_write_set() to write one, datom_member() to declare a member, datom_read() for tables.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  # A product repo declares itself as one and names the single set it owns.
  datom_init_repo(file.path(tmp, "repo"), "example_project", store,
                  mode = "product", set = "example_product")

  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  member <- datom_member(
    conn, "dm", datom_history(conn, "dm")$version[1],
    tags = list(type = "input")
  )
  datom_write_set(conn, list(member),
                  tags = list(description = "Example product"))

  x <- datom_get_set(conn, "example_product")
  print(x)

  # Resolve one member to its data. The link pins the version it was read at.
  print(head(x$members[[1]]$fetch(conn)))

  unlink(tmp, recursive = TRUE)
}

Show Version History

Description

Returns the versions of a table or set, newest first (the 10 most recent by default): when each was saved, by whom, and with what message. Pass a value from the version column to datom_read(), or to datom_get_set() for a set, to read that version back.

Usage

datom_history(conn, name, n = 10, short_hash = FALSE)

Arguments

conn

A datom_conn object from datom_get_conn().

name

Table name.

n

Maximum number of versions to return. Default 10.

short_hash

If TRUE, truncates version and data SHA columns to 8 characters for readability. Default FALSE, so the version column can be passed straight to datom_read().

Value

Data frame with columns: version, data_sha, timestamp, author, commit_message, commit_sha.

A version is content, not code

A datom version answers one question: is this the same content and declared metadata? Nothing code-derived enters it. So a change that alters no content mints no new version, and this is the behaviour most often reported as a bug.

Concretely: you refactor your build script, re-run it, and get byte-identical data. The write is a no-op, datom_history() shows the same version it showed before, and its commit_sha still points at the earlier commit – the one that first produced that content, which does not contain the code you are looking at. That is the recorded value doing its job. It names a commit that provably produces the version; it does not name every commit that could.

The commit is deliberately not part of the version. A set exists to be cited, and if a comment fix minted a new product version, "v47" would stop meaning anything.

Where commit_sha comes from

It is derived, never authored – no argument anywhere sets it. The copy in your clone does not carry it and cannot: version_history.json is committed inside the commit that would name it. Only the copy in storage has it, and it is there for readers who have no clone. With a clone, ⁠git log -p {name}/set.json⁠ answers the same question directly.

NA means the value is not recorded and could not be worked out from this repo's git history – a shallow clone or rewritten history, typically, or a version written by a datom too old to record it.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  print(datom_history(conn, "dm"))

  unlink(tmp, recursive = TRUE)
}

Create a New datom Project

Description

Run once to start a new project. It creates the project folder (including an ⁠input_files/⁠ folder for files you want to bring in), sets up git, pushes a first commit to GitHub – creating the GitHub repository if create_repo = TRUE – and records the new, empty project in storage. The store must carry a GitHub token; to join a project that already exists, use datom_clone() instead.

Usage

datom_init_repo(
  path = ".",
  project_name,
  store,
  create_repo = FALSE,
  repo_name = project_name,
  max_file_size_gb = 1000,
  mode = NULL,
  set = NULL,
  git_ignore = c(".Rprofile", ".Renviron", ".Rhistory", ".Rapp.history", ".Rproj.user/",
    ".DS_Store", "*.csv", "*.tsv", "*.rds", "*.txt", "*.parquet", "*.sas7bdat", ".RData",
    ".RDataTmp", "*.html", "*.png", "*.pdf", ".vscode/", "rsconnect/"),
  .force = FALSE
)

Arguments

path

Path to the project folder. Defaults to current directory.

project_name

Project name, used for S3 namespace and git repo.

store

A datom_store object (from datom_store()). Must have role "developer" (i.e., github_pat provided).

create_repo

If TRUE, create a GitHub repo via API. Mutually exclusive with providing data_repo_url on the store.

repo_name

GitHub repo name when create_repo = TRUE. Defaults to project_name. Useful when the project name (e.g., "STUDY_001") isn't a good GitHub repo name.

max_file_size_gb

Maximum file size limit in GB. Default 1000 (1TB).

mode

Project mode, or NULL (the default) for an ordinary data repo that onboards source files. The only other accepted value is "product", which declares a repo that builds its artifacts instead: it holds one set, writes derived tables, and refuses the file-import path. Absent is not a missing value here – it is "ordinary data repo", which is why nothing is written to project.yaml unless you ask for a product.

set

Name of the set a mode = "product" repo owns. Required with mode = "product" and refused without it: one repo holds one set, and a product repo that names none passes the mode check and then fails every set write, which is a repo that looks initialised and is not.

git_ignore

Character vector of patterns to add to .gitignore.

.force

If TRUE, skip the storage namespace safety check. Use only for intentional takeover of an existing namespace. Default FALSE. Two cases it does not cover, both refusals that stand:

  • a namespace that cannot be reached, because the manifest upload later in this function needs the same storage, so skipping the check buys nothing;

  • a mode = "product" repo, whose namespace check has no override at all – passing .force there is an error rather than a no-op, since a dropped override leaves you believing you took a namespace over when you did not.

Details

Initializes the data repository only. The project is left as a solo project: project.yaml is the location authority, no governance.json / dispatch.json / ref.json is written, and project.yaml omits the storage.governance and repos.governance blocks. A governance store component on store, if present, is ignored here. Governance is attached later via the governance layer (gov_attach()).

Value

Invisible TRUE on success.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )

  datom_init_repo(
    path = file.path(tmp, "repo"),
    project_name = "example_project",
    store = store
  )
  print(list.files(file.path(tmp, "repo"), all.files = TRUE, no.. = TRUE))

  unlink(tmp, recursive = TRUE)
}

Union and Deduplicate source_lineage Lists

Description

Takes a list of zero or more source_lineage lists and returns their deduplicated union. Each entry is a list with project, table, and version_sha. Deduplication uses the composite key paste(project, table, version_sha, sep = "\t"), so each distinct entry appears exactly once and retained entries are returned unchanged.

Usage

datom_lineage_union(lineages)

Arguments

lineages

A list of source_lineage lists (each itself a list of entries with project, table, version_sha). NULL members are treated as empty.

Details

NULL members are tolerated (a parent may carry source_lineage = NULL) and treated as an empty contribution. Empty input, or a list containing only empty lineage lists, returns an empty list.

This helper is the building block of the composable lineage recompute recipe. To check that a derived table's recorded source_lineage matches its parents, read each parent through a connection scoped to that parent's project and union their lineages:

# conn_c is scoped to the derived table's project.
parents <- datom_get_parents(conn_c, "c")

# One connection per project, keyed by each parent's `source`. Never
# reach across project stores with a single connection.
conns <- list(project_a = conn_a, project_b = conn_b)

# Read each parent's lineage through its own project connection.
parent_sls <- lapply(parents, function(p) {
  datom_get_lineage(conns[[p$source]], p$table, version = p$version,
                    depth = "source")
})

recomputed <- datom_lineage_union(parent_sls)
recorded   <- datom_get_lineage(conn_c, "c", depth = "source")
identical(recomputed, recorded)

Value

A deduplicated list of source_lineage entries, or an empty list when there is nothing to union.

Examples

sl1 <- list(list(project = "p", table = "t", version_sha = "a"))
sl2 <- list(list(project = "p", table = "t", version_sha = "a"))
datom_lineage_union(list(sl1, sl2))

List the Tables and Sets in a Project

Description

Returns one row per table and set in the project, with its current version and when it was last updated. Works with both developer and reader connections.

Usage

datom_list(conn, pattern = NULL, include_versions = FALSE, short_hash = TRUE)

Arguments

conn

A datom_conn object from datom_get_conn().

pattern

Optional glob pattern for filtering table names.

include_versions

If TRUE, includes version count info.

short_hash

If TRUE (default), truncates version and data SHA columns to 8 characters for readability. Set to FALSE for full hashes.

Value

Data frame with artifact info (name, kind, current_version, last_updated, etc.).

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  print(datom_list(conn))

  unlink(tmp, recursive = TRUE)
}

List a Set's Members and Their Labels

Description

A data frame with one row per member per label value – long format, not wide. Tags are open-keyed and multi-valued, so a wide frame would need a list-column and a column set that changes from one set to the next; long is a plain frame with fixed columns whatever the set holds.

Usage

datom_list_members(x)

Arguments

x

A datom_set from datom_get_set().

Details

Filtering is therefore ordinary R – subset(), dplyr::filter() – and datom grows no query vocabulary of its own.

Value

A data frame of name, project, version, kind, key, value. Zero rows, with those columns, for a set with no members.

The columns

name, project, version and kind identify the member; key and value are one label. An untagged member still gets a row, with NA for both, so unique(m$name) is the complete member list rather than the tagged part of it.

value is a plain character column, never a list-column, because the tag grammar is text only.

See Also

datom_structure_members() for a navigable view of the same labels, datom_fetch_member() to resolve one member.

Examples

# These two facts about a set -- its members and their labels -- are all this
# verb reads, so a set built by hand shows the shape. In practice `x` comes
# from datom_get_set().
x <- structure(
  list(
    name = "study001-adam", project = "study001", version = NULL,
    data_sha = NULL, tags = list(description = "ADaM datasets"),
    members = list(
      list(
        id = list(project = "study001", name = "adsl", kind = "table",
                  version = strrep("a", 64)),
        tags = list(type = "output", domain = c("safety", "efficacy"))
      ),
      list(
        id = list(project = "study001", name = "dm", kind = "table",
                  version = strrep("b", 64))
      )
    )
  ),
  class = "datom_set"
)

# adsl appears three times, once per label; the untagged dm appears once.
datom_list_members(x)

# Filtering is plain R.
subset(datom_list_members(x), key == "domain" & value == "safety")

Declare a Member of a Set

Description

Resolves one artifact version against a single project connection and returns a pure-data member record to pass to a set write. The record is a pointer: it names the project, artifact, kind, and version, and carries no copy of the data. Reading the artifact's versioned metadata snapshot is what makes the pointer trustworthy – a member can only point at something that already exists, which is also why a set cannot contain itself at any depth.

Usage

datom_member(conn, name, version, tags = NULL)

Arguments

conn

A datom_conn scoped to the member's project store, from datom_get_conn().

name

Artifact name (single validated string).

version

The artifact version (metadata_sha) to pin, e.g. from datom_history().

tags

Optional named list of text labels for this member. Omitted from the record when absent or empty.

Details

Same-project and cross-project members are declared identically; the only difference is which connection is passed. kind comes from the snapshot (defaulting to "table" for a snapshot written before datom recorded the field), and project comes from the repo rather than from the connection: the artifact's own metadata, else the project manifest, else the connection's name with a warning saying it is unverified. That matters because a reader connection's project name is a label the caller supplied and nothing checks it against the repo, while this value is hashed into the set's identity and cited afterwards.

Unlike datom_parent(), a member carries no data_sha: the version already pins the content, and a second copy of that fact would be a second thing to keep consistent.

Value

A list with id (a list of exactly project, name, kind, version) and, when tags were supplied, tags. Pure data: it retains no connection and is serializable.

Tags

tags is an optional named list of text labels describing this member's role in the set – what folder structure would otherwise express. A value may be a single string or several, because the whole point of labels over folders is that an item can be in more than one category at once: list(type = "output", domain = c("safety", "efficacy")).

Values are text only: no numbers, booleans, or nesting. Write a numeric label as a string ("500") and parse it downstream, exactly as you would a folder name.

Three outcomes, and the difference is whether anything is actually there:

What you pass What happens
no tags accepted; the record carries no tags
list(domain = character(0)) or list(domain = NULL) the key is dropped, as if never mentioned
list(domain = "") refused -- a label with no name is almost always an accident

See Also

datom_parent() for the lineage equivalent.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")

  # Pin the version just written and label its role in the set.
  version <- datom_history(conn, "dm")$version[1]
  print(datom_member(conn, "dm", version, tags = list(type = "input")))

  unlink(tmp, recursive = TRUE)
}

Name an Input for a Table You Are About to Write

Description

Use before datom_write() when the table you are writing was made from other datom tables. Each call names one input – one table at one exact version – and returns a note that datom_write() saves with the new table, so its lineage records what it was made from. For several inputs, make one call each and pass them together as a list to parents. If the inputs are members of a set, pass the set as x and name several tables at once; each gets the version the set pins.

Usage

datom_parent(conn, table, version = NULL, x = NULL, tags = NULL)

Arguments

conn

A datom_conn scoped to the parent's project store, from datom_get_conn().

table

Parent table name (single non-empty validated string). With x, a character vector of one or more member names.

version

Parent version (metadata_sha; single non-empty string). Exactly one of version and x is required.

x

Optional datom_set from datom_get_set() (or one being built) whose members supply the versions. Exactly one of version and x is required.

tags

Optional named list of labels narrowing a member name the set holds more than once, e.g. list(type = "input"). Only with x.

Details

The parent's data fingerprint is read from the parent's own saved record; you cannot supply it, so a lineage entry cannot claim data the parent never had. The record is plain data with no connection inside, so it can be saved and reused.

Same-project and cross-project parents are declared identically – the only difference is which connection is passed. source is always derived from the connection's project_name.

Value

Without x, a list with exactly source, table, version, data_sha, and source_lineage. source is the project the parent's own metadata says it belongs to, falling back to the project manifest and then to the connection's name (see .datom_declared_project()) – not simply the name on conn, which on a reader connection is an unverified label and which source cannot afford, since it is part of the declaring table's version. source_lineage is NULL when the snapshot carries none. With x, an unnamed list of such records, one per table, in the order given.

Taking the version from a set

When the inputs of a derivation are members of a set, pass the set as x instead of a version: each table is looked up among the set's members and declared at the version the set pins. That keeps the parents a derived table records in step with the set it is derived through, without looking each version up by hand.

The member is chosen exactly as datom_fetch_member() chooses it. A name held by two members – a current table beside a locked baseline, say – stops and lists both; narrow with tags. A member that is itself a set stops too, since only a table can be a parent.

With x, table may name several tables and the result is always a list of parent records, even for one table, so it can be passed straight to the parents argument of datom_write(). Without x the result is one record, as it always was.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")

  # Resolve a parent declaration to pass to the parents argument of
  # datom_write.
  print(datom_parent(conn, "dm", datom_history(conn, "dm")$version[1]))

  unlink(tmp, recursive = TRUE)
}

List Projects Registered in the Governance Repo

Description

Returns a data frame with one row per project registered in the shared governance repo. Useful for managers and auditors who need to see the portfolio without having to clone every data repo.

Usage

datom_projects(x)

Arguments

x

A datom_conn or a datom_store with a governance component.

Details

Accepts either a datom_conn (typically the developer's existing connection – reads the local gov clone) or a datom_store (lets a caller enumerate the portfolio before connecting to any specific project).

Read path:

Corrupt registry entries (missing ref.json, unreadable JSON) emit a warning and are skipped – one bad project does not take down the listing.

Value

A data frame, sorted by name, with columns: name (character), data_backend (character), data_root (character), data_prefix (character; NA when absent), registered_at (character ISO8601 from clone mtime; NA on storage path).

Examples

# A governance store backed by a local directory. Projects are registered
# into it by the companion governance package (datomanager), so a freshly
# created governance store lists an empty portfolio.
tmp <- tempfile("datom-example-")
gov <- datom_store_local(file.path(tmp, "gov-storage"))

store <- datom_store(
  governance   = gov,
  data         = datom_store_local(file.path(tmp, "storage")),
  gov_repo_url = "https://github.com/example/acme-gov"
)

datom_projects(store)

unlink(tmp, recursive = TRUE)

Pull Latest Changes from Remote

Description

Fetches and merges the latest git changes from the remote repository. This is the recommended entry point at the start of each work session to ensure the local state is current before syncing or writing tables.

Usage

datom_pull(conn)

Arguments

conn

A datom_conn object from datom_get_conn().

Details

Git is the source of truth for all metadata (manifest, dispatch, table metadata). The manifest and other metadata files live in git and are pulled along with any other committed changes.

Requires developer role (readers have no git access).

Value

Invisibly, a list with:

commits_pulled

Integer count of new commits merged.

branch

Current branch name.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Nothing new on the remote yet, so this is a no-op.
  datom_pull(conn)

  unlink(tmp, recursive = TRUE)
}

Read a datom Table

Description

Returns a table as a data frame: the current version by default, or a past version when you pass version (copy it from datom_history()). Works with both developer and reader connections. To read a set, use datom_get_set().

Usage

datom_read(conn, name, version = NULL, context = NULL, ...)

Arguments

conn

A datom_conn object from datom_get_conn().

name

Table name.

version

Optional metadata_sha (datom version). If NULL, uses current.

context

Reserved; currently ignored.

...

Reserved; currently ignored.

Value

A data frame.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")

  # Current version
  dm <- datom_read(conn, "dm")
  print(head(dm))

  # A specific version, by its identifier -- byte-for-byte the same table
  v <- datom_history(conn, "dm")$version[1]
  print(identical(datom_read(conn, "dm", version = v), dm))

  unlink(tmp, recursive = TRUE)
}

Drop Members from a Set

Description

Removes the members you select and returns the set without them. Nothing is written: the object comes back edited, and the set is stored only when you pass the result to datom_write_set().

Usage

datom_remove_members(x, member = NULL, tags = NULL, version = NULL)

Arguments

x

A datom_set, from datom_get_set() or datom_assemble_set().

member

The member to drop, as its name, a member record, or a link. Optional only when tags or version selects on its own.

tags

Optional named list of labels selecting members, e.g. list(status = "draft"). A member matches when it carries every label listed.

version

Optional version, or a prefix of one, selecting the members pinned at it.

Details

It exists because the hand-rolled version is silently wrong. Filtering a member list by name drops every version of that name, so a set holding a live table beside a deliberately frozen baseline loses both; removing by position removes a different member the day somebody adds one.

Value

x without the selected members, its version and data_sha emptied because they described a payload it no longer holds, and what was dropped appended to its datom_edits attribute – which datom_write_set() turns into the commit message.

Why there is no connection argument

Not an oversight. Removing a member only has to find a pointer the set already holds, while adding or repointing one has to resolve it – read the artifact's metadata, confirm its kind, record the project that wrote it. So this verb does no IO at all and needs no credentials, which is also why it is the one edit verb that works on a set read through a storage-only connection with no clone.

Selecting what to drop

A selection is required. datom_remove_members(x) would mean removing every member, which the writer refuses anyway, so it aborts instead of building a payload the write then rejects. The safe default for a destructive verb is nothing – the opposite of datom_update_members(), where the safe default is everything because a refresh is idempotent.

Select by name, by a member record, or by a link, and narrow with tags or version; tags or version on their own select every member they match, so tags = list(status = "draft") drops the labelled ones.

Three things it refuses rather than doing quietly:

What you asked for Why it stops
a name matching more than one member that is the silently-wrong spelling this verb replaces -- it would drop a frozen baseline along with the live table
a selection matching nothing it is a typo, and Filter() reports success for it
every member a set with no members cannot be written, and the refusal belongs on the line that emptied it

An ambiguous name refuses here while datom_update_members() skips it, and the asymmetry is the consequence rather than a taste: skipping a repoint leaves a valid pinned version behind, while skipping a removal silently does nothing at all.

See Also

datom_update_members() to repoint members instead, datom_list_members() to see what a set holds, datom_write_set() to store the result.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  # A product repo declares itself as one and names the single set it owns.
  datom_init_repo(file.path(tmp, "repo"), "example_project", store,
                  mode = "product", set = "example_product")

  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  datom_write(conn, data = datom_example_data("lb"), name = "lb")
  datom_write_set(conn, list(
    datom_member(conn, "dm", datom_history(conn, "dm")$version[1],
                 tags = list(type = "input")),
    datom_member(conn, "lb", datom_history(conn, "lb")$version[1],
                 tags = list(status = "draft"))
  ))

  x <- datom_get_set(conn, "example_product")

  # By label, which is the selection that does not depend on position.
  x <- datom_remove_members(x, tags = list(status = "draft"))
  print(datom_list_members(x))

  datom_write_set(conn, x)

  unlink(tmp, recursive = TRUE)
}

Write the Data-Side Governance Attachment Record

Description

Writes governance.json – the data-side pointer recording which governance repository a project is attached to. This is the data-repo / data-storage half of attaching governance; the gov-repo registration (writing ref.json and dispatch.json, committing to the gov repo) is performed separately by the governance layer (datomanager::gov_attach()).

Usage

datom_repo_attach_governance(conn, gov_repo_url, gov_store, message = NULL)

Arguments

conn

A datom_conn object with role = "developer" and a local data clone (conn$path).

gov_repo_url

HTTPS clone URL of the governance git repository to record.

gov_store

A datom_store_s3 or datom_store_local component for the governance storage. Only its location fields are persisted; credentials are discarded.

message

Optional commit message. Defaults to "Attach governance: {project_name}".

Details

governance.json is the canonical data->gov pointer in the bidirectional governance link: the gov repo's ref.json points gov->data, and this file points data->gov, so either repo can find the other. It is written to two locations, mirroring the manifest pattern (git canonical, storage derived):

Routing this write through datom upholds the two-repos invariant: the governance layer never mutates the data repo directly.

Value

Invisibly, the SHA of the resulting data-repo commit.

See Also

datom_repo_delete(), datom_repo_set_data_store()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Data-side half of governance attachment. The gov-repo registration
  # is performed separately by the companion datomanager package.
  datom_repo_attach_governance(
    conn,
    gov_repo_url = "https://github.com/example/acme-gov",
    gov_store    = datom_store_local(file.path(tmp, "gov-storage"))
  )
  gov_json <- jsonlite::read_json(
    file.path(tmp, "repo", ".datom", "governance.json")
  )
  print(gov_json$gov_repo_url)

  unlink(tmp, recursive = TRUE)
}

Commit Content in the Data Repo

Description

Commits changes in the data repo clone and, by default, pushes them. This is datom's sanctioned git-mutation surface for downstream packages: a build or deployment package commits its own content – code, renv.lock, framework state – through this verb rather than importing git2r and writing to the data repo behind datom's back.

Usage

datom_repo_commit(conn, message, paths = NULL, push = TRUE)

Arguments

conn

A datom_conn object with role = "developer" and a local repo path (conn$path).

message

Commit message. Required, and used only when a commit is actually created.

paths

NULL (default) to stage everything ⁠git add .⁠ would stage, or a character vector of repo-relative paths to stage exactly those. Explicit paths must exist: to record a deletion, use paths = NULL.

push

Push after committing (default TRUE), through the same path datom's own writes use – so it inherits pull-before-push and upstream tracking. push = FALSE commits only; datom_repo_push() is the other half of that split.

Details

paths = NULL means what ⁠git add .⁠ means. It stages tracked modifications, deletions and untracked files, minus anything .gitignore excludes. That is the correct semantic for a human-invoked moment, and it is deliberately the opposite of what datom's own writes do: a commit created inside datom_write() stages an explicit file list, because it fires at a moment datom chose and must never sweep up work in progress.

One consequence of add-all worth knowing rather than discovering: if an earlier datom write failed after writing local metadata but before committing, those datom files are dirty and this verb will stage them. That is left intentional – silently excluding datom's own paths would make the argument lie about its contract, and the state is exactly what datom_validate() reports and datom_validate(fix = TRUE) repairs. It also moves git ahead of storage, which is the safe direction.

Commit is idempotent, push is convergent, and neither implies the other. A clean tree produces no commit and is not an error, so "commit everything" can be called twice. With push = TRUE the push still runs when the branch is ahead of the remote, even though no commit was created – otherwise one failed push would leave the remote behind forever, since every later call finds a clean tree and returns early.

Value

Invisibly, the commit SHA; invisible(NULL) when no commit was created (whether or not a push happened).

See Also

datom_repo_push(), datom_validate()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Content datom does not own, committed through datom.
  dir.create(file.path(tmp, "repo", "R"))
  writeLines("build <- function() NULL", file.path(tmp, "repo", "R", "build.R"))
  datom_repo_commit(conn, "Add build script")

  # Idempotent: a clean tree is a no-op, not an error.
  datom_repo_commit(conn, "Nothing to do")

  unlink(tmp, recursive = TRUE)
}

Delete the Data GitHub Repository and Local Clone

Description

Deletes the data-side GitHub repository via the GitHub REST API and removes the local clone directory. This is the data-side teardown step for a datom project.

Usage

datom_repo_delete(conn, confirm, force_gov_attached = FALSE)

Arguments

conn

A datom_conn object (developer role required).

confirm

Character string. Must equal conn$project_name exactly. No interactive prompts – this must be supplied explicitly.

force_gov_attached

Logical. FALSE (default) refuses to run when governance is attached (!is.null(conn$gov_root)). Pass TRUE only when called programmatically from datomanager::gov_decommission().

Details

Solo projects (no governance attached): call this together with datom_storage_delete_prefix() for a complete teardown.

Governed projects: use datomanager::gov_decommission() instead. That function calls datom_repo_delete() internally (with force_gov_attached = TRUE). Calling datom_repo_delete() directly on a governed project without that flag is refused to prevent accidentally orphaning the governance registration.

Steps:

  1. Delete the data GitHub repo via the GitHub REST API (requires conn$github_pat with delete_repo scope; skipped with a warning when conn$github_pat is NULL or when the remote is not GitHub). Aborts if conn$data_repo_url is not set.

  2. Remove the local clone directory (conn$path).

Each step is warn-and-continue on failure so the other still runs.

Value

Invisible TRUE on success.

See Also

datom_storage_delete_prefix(), datom_repo_set_data_store()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage. Because the remote is not GitHub,
# the API deletion step is skipped and only the local clone is removed.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Solo project teardown (no governance): storage, then repo.
  datom_storage_delete_prefix(conn)
  datom_repo_delete(conn, confirm = conn$project_name)

  unlink(tmp, recursive = TRUE)
}

Push the Data Repo to Its Remote

Description

Pushes the current branch of the data repo clone, through the same path datom's own writes use – so it inherits pull-before-push, upstream tracking, and the on-a-branch guard.

Usage

datom_repo_push(conn)

Arguments

conn

A datom_conn object with role = "developer" and a local repo path (conn$path).

Details

Convergent, not imperative. Nothing to push is an informational no-op rather than an error, so calling it twice is safe and "make sure the remote has everything" is a legal standalone operation.

This is the other half of datom_repo_commit()(push = FALSE). Without it, "push what I already committed" would only be expressible as another commit attempt – and since paths = NULL is add-all, a caller who merely wanted to push would risk committing whatever work in progress the tree happened to hold.

Value

Invisibly TRUE.

See Also

datom_repo_commit(), datom_pull()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Commit now, push later.
  writeLines("notes", file.path(tmp, "repo", "NOTES.md"))
  datom_repo_commit(conn, "Add notes", push = FALSE)
  datom_repo_push(conn)

  # Convergent: a second call has nothing to do and says so.
  datom_repo_push(conn)

  unlink(tmp, recursive = TRUE)
}

Rewrite the Data Store Pointer in project.yaml

Description

Updates storage.data in .datom/project.yaml to point at new_store, then commits and pushes the data repo. This is the data-side bookkeeping step of a store relocation.

Usage

datom_repo_set_data_store(conn, new_store, message = NULL)

Arguments

conn

A datom_conn object with role = "developer" and a local repo path (conn$path).

new_store

A datom_store_s3 or datom_store_local component (i.e. the data-side component of a datom_store() object, not the full composite).

message

Optional commit message. Defaults to "Update data store: {project_name}".

Details

Read-modify-write contract: the function reads the full existing project.yaml, modifies only storage.data, and writes back. It never reconstructs the file from conn fields. This preserves storage.governance on governed projects (it is permanent once written) and any other fields not owned by this function.

For governed projects the authoritative address is ref.json in the gov repo – this function updates only the local data clone so that datom_get_conn() stays consistent after migration. It is called by datomanager::gov_migrate_data() after the ref switch, never before.

Value

Invisibly, the SHA of the resulting commit.

See Also

datom_storage_copy(), datom_storage_verify(), datom_repo_delete()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Repoint project.yaml at a relocated data store.
  new_store <- datom_store_local(file.path(tmp, "storage-relocated"))
  datom_repo_set_data_store(conn, new_store)

  unlink(tmp, recursive = TRUE)
}

Check datom Repository Structure

Description

Returns detailed check results for each component.

Usage

datom_repository_check(path)

Arguments

path

Path to evaluate.

Value

List of TRUE/FALSE per check.


Show Repository Status

Description

Displays connection info, table count, and (for developers) uncommitted git changes and input file sync state.

Usage

datom_status(conn)

Arguments

conn

A datom_conn object from datom_get_conn().

Value

Invisibly, a list with connection, tables, and optionally git and input_files status details.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_status(conn)

  unlink(tmp, recursive = TRUE)
}

Copy All Objects Between Two datom Storage Namespaces

Description

Enumerates all objects under from_conn's datom namespace and streams each one to to_conn's datom namespace. All four backend combinations are supported:

Usage

datom_storage_copy(from_conn, to_conn)

Arguments

from_conn

A datom_conn object (source).

to_conn

A datom_conn object (destination).

Details

This is a policy-free primitive. It does not modify the source namespace, update project.yaml, or switch ref.json. For a complete managed migration (governed projects) use datomanager::gov_migrate_data(). For solo-project relocation combine this function with datom_repo_set_data_store().

Value

A data frame with columns key (character, relative key after ⁠{prefix}/datom/⁠) and bytes (numeric, byte count per object). Returns a zero-row data frame if the source namespace is empty.

See Also

datom_storage_verify(), datom_storage_list(), datom_storage_delete_prefix(), datom_repo_set_data_store()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and two
# local directories for the source and destination object stores.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  from_conn <- datom_get_conn(file.path(tmp, "repo"), store)
  datom_write(from_conn, data = datom_example_data("dm"), name = "dm")

  # The destination is addressed with a reader connection: no local repo,
  # just a store plus the project name.
  to_store <- datom_store(data = datom_store_local(file.path(tmp, "storage2")))
  to_conn <- datom_get_conn(store = to_store, project_name = "example_project")

  copied <- datom_storage_copy(from_conn, to_conn)
  print(nrow(copied))      # number of objects copied
  print(sum(copied$bytes)) # total bytes

  unlink(tmp, recursive = TRUE)
}

Delete All Objects Under a datom Storage Prefix

Description

Removes every file under {prefix}/datom/{prefix_key} from storage. Pass prefix_key = NULL (the default) to delete the entire datom namespace for this connection. A missing or empty prefix is a no-op.

Usage

datom_storage_delete_prefix(conn, prefix_key = NULL)

Arguments

conn

A datom_conn object.

prefix_key

Relative prefix to delete under (after ⁠{prefix}/datom/⁠). NULL (default) deletes the entire datom namespace root for this connection.

Details

Irreversible. Intended for package developers building tools on top of datom (e.g. datomanager for rollback or source deletion after migration). End users performing a full project teardown should use datom_repo_delete() instead.

Value

Invisibly, a backend-specific value. For S3: the count of deleted objects (0L if nothing found). For the local backend: 1L if the prefix directory existed and was removed, 0L otherwise.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)
  datom_write(conn, data = datom_example_data("dm"), name = "dm")

  # Delete a single table's objects
  datom_storage_delete_prefix(conn, prefix_key = "dm")

  # Delete the entire datom namespace (use with care)
  datom_storage_delete_prefix(conn)
  print(datom_storage_list(conn))

  unlink(tmp, recursive = TRUE)
}

List All Objects in a datom Storage Namespace

Description

Returns the full storage keys of every object under the datom namespace for this connection ({prefix}/datom/...). Intended for package developers building tools on top of datom (e.g. datomanager); end users typically do not need to inspect raw storage keys directly.

Usage

datom_storage_list(conn)

Arguments

conn

A datom_conn object.

Details

Keys are returned in their full storage-key form – for S3 that is "{prefix}/datom/..." relative to the bucket root; for local backends it is a path relative to conn$root. This mirrors the contract of the internal .datom_storage_list_objects() dispatch layer.

Value

A character vector of full storage keys. May be empty if the namespace contains no objects.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)
  datom_write(conn, data = datom_example_data("dm"), name = "dm")

  print(datom_storage_list(conn))

  unlink(tmp, recursive = TRUE)
}

Read a JSON Document from a datom Storage Namespace

Description

Reads and parses a JSON object from storage, dispatching on the connection's backend. Intended for package developers building tools on top of datom (e.g. datomanager) that need to inspect datom's own documents – manifests, metadata, version history – without reaching into internals via :::.

Usage

datom_storage_read_json(conn, key)

Arguments

conn

A datom_conn object.

key

Relative storage key (after ⁠{prefix}/datom/⁠).

Details

End users should prefer the purpose-built readers: datom_read() for table data, datom_list() and datom_summary() for the manifest, and datom_history() for version history. This export is a byte-level primitive and does not interpret what it reads.

key is a relative key – the portion after ⁠{prefix}/datom/⁠, e.g. "dm/.metadata/metadata.json". The backend prepends the namespace itself, so passing a full key (one containing a ⁠datom/⁠ segment) is refused rather than silently resolving under ⁠{prefix}/datom/{prefix}/datom/⁠ and finding nothing. Keys containing a .. segment or a leading / are refused too: reads are confined to this project's namespace.

There is no corresponding write export. Documents are written by purpose-built verbs (datom_write(), datom_repo_attach_governance()), not by a generic byte channel, so that every mutation of a datom namespace routes through a function that knows what it is writing.

Value

The parsed JSON document as an R list. Nested structures are kept as lists (simplifyVector = FALSE), so JSON arrays of objects do not collapse into data frames.

See Also

datom_storage_list() to discover keys; datom_read(), datom_list(), datom_history() for the interpreted equivalents.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)
  datom_write(conn, data = datom_example_data("dm"), name = "dm")

  # Relative key: the part after `{prefix}/datom/`
  meta <- datom_storage_read_json(conn, "dm/.metadata/metadata.json")
  print(meta$hash_algo)

  unlink(tmp, recursive = TRUE)
}

Verify a Copy Between Two datom Storage Namespaces

Description

Checks that objects in to_conn's datom namespace match their counterparts in from_conn. Two verification modes are available:

Usage

datom_storage_verify(
  from_conn,
  to_conn,
  keys = NULL,
  mode = c("structural", "content")
)

Arguments

from_conn

A datom_conn object (source / reference).

to_conn

A datom_conn object (destination to verify).

keys

Character vector of relative keys (after ⁠{prefix}/datom/⁠) to verify. NULL (default) verifies every key returned by datom_storage_list(from_conn). Pass a subset to verify a sample.

mode

"structural" (default) or "content". See above.

Details

Value

A data frame with columns:

See Also

datom_storage_copy(), datom_storage_list()

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and two
# local directories for the source and destination object stores.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  from_conn <- datom_get_conn(file.path(tmp, "repo"), store)
  datom_write(from_conn, data = datom_example_data("dm"), name = "dm")

  to_store <- datom_store(data = datom_store_local(file.path(tmp, "storage2")))
  to_conn <- datom_get_conn(store = to_store, project_name = "example_project")
  copied <- datom_storage_copy(from_conn, to_conn)

  # Verify all copied objects structurally (default, fast)
  results <- datom_storage_verify(from_conn, to_conn)
  print(all(results$ok))

  # Verify a subset with full content hash
  print(datom_storage_verify(from_conn, to_conn,
                             keys = copied$key[1],
                             mode = "content"))

  unlink(tmp, recursive = TRUE)
}

Create a datom Store

Description

A store tells datom where a project's data is kept – a local folder (datom_store_local()) or an S3 bucket (datom_store_s3()) – and, if you will be writing, your GitHub token. With a token you are a developer and can write; without one you are a reader and can only read. Pass the store to datom_init_repo() to start a project, or to datom_get_conn() to connect to one.

Usage

datom_store(
  governance = NULL,
  data,
  github_pat = NULL,
  data_repo_url = NULL,
  gov_repo_url = NULL,
  gov_local_path = NULL,
  github_org = NULL,
  github_api_url = NULL,
  validate = TRUE
)

Arguments

governance

A store component (e.g., datom_store_s3()) for governance files (dispatch, ref, migration history), or NULL for a no-governance store. A no-governance store represents a project that has not yet been promoted to governance (via the datomanager package); gov_repo_url and gov_local_path must also be NULL in that case.

data

A store component (e.g., datom_store_s3()) for data files (manifest, tables, metadata).

github_pat

GitHub personal access token. If provided, role is "developer". If NULL, role is "reader".

data_repo_url

GitHub remote URL for the data repository. Required when github_pat is provided and create_repo = FALSE in datom_init_repo().

gov_repo_url

GitHub remote URL for the shared governance repository. The governance repo is created once per org (via the datomanager package) and referenced here by every project that uses it.

gov_local_path

Local directory path for the governance clone. If NULL (default), the clone is placed as a sibling of the data repo, named after the basename of gov_repo_url (e.g., "acme-gov").

github_org

GitHub organization for repo creation. NULL for personal repos.

github_api_url

GitHub API base URL. NULL (default) uses "https://api.github.com", which is correct for github.com and GitHub Enterprise Cloud (GHEC). For GitHub Enterprise Server (GHES) pass the server's API root, e.g. "https://github.mycompany.com/api/v3". A trailing / is stripped for consistency.

validate

If TRUE (default), validate GitHub PAT via API. Set to FALSE for tests or offline use.

Value

A datom_store object.

Examples

tmp <- tempfile("datom_store_")
store <- datom_store(
  data = datom_store_local(path = tmp),
  data_repo_url = "https://github.com/example/my-project",
  validate = FALSE
)
store
is_datom_store(store)
unlink(tmp, recursive = TRUE)

Create a Local Filesystem Store Component

Description

Constructs a validated local filesystem storage component for use as either the governance or data component of a datom_store. Validates that the path exists (or is creatable) and is writable.

Usage

datom_store_local(path, prefix = NULL, validate = TRUE)

Arguments

path

Directory path for the store root.

prefix

Key prefix within the root (e.g., "project/"). NULL for no prefix.

validate

If TRUE (default), validate that path exists and is writable. Set to FALSE for tests or deferred creation.

Value

A datom_store_local object.

Examples

tmp <- tempfile("datom_store_")
store <- datom_store_local(path = tmp, validate = TRUE)
store
is_datom_store_local(store)
unlink(tmp, recursive = TRUE)

Create an S3 Store Component

Description

Constructs a validated S3 storage component for use as either the governance or data component of a datom_store. Validates credentials and bucket access at construction time (unless validate = FALSE).

Usage

datom_store_s3(
  bucket,
  prefix = NULL,
  region = "us-east-1",
  access_key,
  secret_key,
  session_token = NULL,
  validate = TRUE
)

Arguments

bucket

S3 bucket name.

prefix

S3 key prefix (e.g., "project/"). NULL for no prefix.

region

AWS region (default "us-east-1").

access_key

AWS access key ID.

secret_key

AWS secret access key.

session_token

Optional AWS session token (for temporary credentials).

validate

If TRUE (default), validate credentials and bucket access at construction time. Set to FALSE for tests or offline use.

Value

A datom_store_s3 object.

Examples

s3 <- datom_store_s3(
  bucket = "my-datom-bucket",
  prefix = "project/",
  region = "us-east-1",
  access_key = "AKIAIOSFODNN7EXAMPLE",
  secret_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
  validate = FALSE
)
s3
is_datom_store_s3(s3)

Create a Credentials-Only S3 Store Component

Description

Constructs an S3 store component that carries only AWS credentials – no bucket, prefix, or region. The data location is resolved at connection time from ref.json stored in the governance repo. This is the recommended construction style for readers when a governance store is in place.

Usage

datom_store_s3_creds(access_key, secret_key, session_token = NULL)

Arguments

access_key

AWS access key ID.

secret_key

AWS secret access key.

session_token

Optional AWS session token (for temporary credentials).

Details

A datom_store_s3_creds component must be paired with a governance component inside datom_store(). Attempting to create a composite store without governance will abort with a clear message.

Value

A datom_store_s3_creds object.

Examples

creds <- datom_store_s3_creds(
  access_key = "AKIAIOSFODNN7EXAMPLE",
  secret_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
)
creds
is_datom_store_s3_creds(creds)

Group a Set's Members into a Navigable View

Description

Groups members by the values of the label key(s) named in by and returns a nested list whose leaves are the members' links, so dp$output$adsl(conn) works and tab-completes.

Usage

datom_structure_members(x, by, missing = "untagged")

Arguments

x

A datom_set from datom_get_set().

by

Character vector of one or more label keys to group by, outermost first.

missing

Branch name for members carrying no value for an axis key.

Details

A pure function of x and the axis you ask for: nothing is stored, and datom takes no position on which hierarchy is the right one. Ask for by = c("domain", "type") and you get a different tree from the same set, which is that design working rather than being worked around.

Value

A nested list length(by) + 1 levels deep: one level per axis, then the member's own name holding its datom_link. An empty list for a set with no members.

One member, several branches

A member tagged domain = c("safety", "efficacy") appears under both safety and efficacy, so the total number of leaves can exceed the number of members. That is the point of labels over folders rather than a quirk of this verb: a folder holds an item in exactly one place, and a label does not.

What is refused, and why nothing is dropped

Two requests abort, because the alternative in both cases is a member the consumer cannot find and cannot see is absent:

A member that simply lacks the axis key is not refused and not dropped: it goes under missing, named, because a named bucket is visible and an omission is not.

See Also

datom_list_members() for the flat view, datom_fetch_member() to resolve one member by name.

Examples

# Built by hand to show the shape; in practice `x` comes from
# datom_get_set(). adsl carries two domains, so it appears under both.
x <- structure(
  list(
    name = "study001-adam", project = "study001", version = NULL,
    data_sha = NULL, tags = NULL,
    members = list(
      list(
        id = list(project = "study001", name = "adsl", kind = "table",
                  version = strrep("a", 64)),
        tags = list(domain = c("safety", "efficacy"))
      ),
      list(
        id = list(project = "study001", name = "dm", kind = "table",
                  version = strrep("b", 64))
      )
    )
  ),
  class = "datom_set"
)

dp <- datom_structure_members(x, by = "domain")
print(names(dp))
print(names(dp$safety))

# A leaf is a link: call it with a connection to resolve it.
print(dp$safety$adsl)

Summarize a datom Project

Description

Returns a compact, role-aware overview of a datom project: its name, backend, table/version totals, last write time, and (for developers) the git remote URL. Reads .metadata/manifest.json from the data store.

Usage

datom_summary(conn)

Arguments

conn

A datom_conn object from datom_get_conn().

Value

A datom_summary S3 object (a list with class "datom_summary") containing: project_name, role, backend, root, prefix, table_count, set_count, total_versions, last_updated, remote_url. table_count counts tables only and set_count counts sets; total_versions stays tables-only, so no counter changed meaning. remote_url is NULL for readers (no local data clone).

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  print(datom_summary(conn))

  unlink(tmp, recursive = TRUE)
}

Bring New and Changed Files Into a Project

Description

Takes the preview from datom_sync_manifest() and saves each new or changed file as a version of a table named after the file; unchanged files are skipped. This is the usual way to bring files into datom. To save a data frame you built in R, use datom_write().

Usage

datom_sync(
  conn,
  manifest,
  continue_on_error = TRUE,
  sources = NULL,
  tags = list(type = "input"),
  x = NULL
)

Arguments

conn

A datom_conn object from datom_get_conn().

manifest

Data frame from datom_sync_manifest(). On an ordinary repo, with columns name, file, format, original_file_sha, status. On a product repo, the set preview: project, name, kind, version_from, version_to, status. Any subset of its rows will do.

continue_on_error

If TRUE (default), continues processing remaining tables when one fails. If FALSE, stops on first error. Not accepted on a product repo, where one failure stops the call and nothing has been written.

sources

On a product repo only, and required there: one datom_conn, or a list of them, for the projects named by the rows being applied – the same connections the preview was built with. Refused on an ordinary repo.

tags

On a product repo only: the labels given to members added by new rows. Default list(type = "input"). A repointed member keeps its own labels. Refused on an ordinary repo.

x

On a product repo only: the datom_set to apply the preview to, from datom_get_set() or datom_assemble_set(). Omitted, the repo's stored set is read, or an empty one used when it has never been written. Refused on an ordinary repo.

Details

Reading files needs the rio package (install.packages("rio")).

Rows flagged "unsupported_format" by datom_sync_manifest() are reported as result = "error" with the recourse in the error column; the rest of the batch still processes.

On a product repo (mode: product) it applies a preview of the repo's set instead – see "On a product repo" below.

Value

On an ordinary repo, the manifest data frame augmented with result and error columns. result is "success", "skipped", or "error".

On a product repo, the updated datom_set, with what changed appended to its datom_edits attribute. Its version and data_sha are emptied when any row was applied.

On a product repo

A product repo owns one set, and this call applies a preview from datom_sync_manifest() to it: each new row adds a member at version_to, labelled with tags, and each changed row repoints the member it names from version_from to version_to, keeping that member's labels exactly. Rows of any other status do nothing. Filter the preview first to apply only part of it – subset(m, name != "lb") – or build the frame by hand with the same columns.

Nothing is written. The set comes back edited, and it is stored only when you pass it to datom_write_set(); the call ends by saying so. The write's default commit message then names what was added and repointed.

This is the one difference from syncing files, where each table is written as it syncs. A set is saved in one step, so the edit becomes one version, and you can look at the set, or add to it, before it does.

Every member added is read from its source first, which confirms the version exists and records the project that wrote it. The call stops, before changing anything, when:

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage. File import needs the optional
# rio package.
if (requireNamespace("git2r", quietly = TRUE) &&
    requireNamespace("rio", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  file.copy(
    system.file("extdata", "dm.csv", package = "datom"),
    file.path(tmp, "repo", "input_files", "dm.csv")
  )

  manifest <- datom_sync_manifest(conn)
  result <- datom_sync(conn, manifest)
  print(result[, c("name", "status", "result")])

  unlink(tmp, recursive = TRUE)
}

Preview What a Sync Will Change

Description

Looks at the files in the project's ⁠input_files/⁠ folder and returns one row per file, saying whether it is new, changed, unchanged since it was last synced, or in a format datom cannot read. Nothing is written. Review the result, drop any rows you do not want, then pass it to datom_sync().

Usage

datom_sync_manifest(conn, path = NULL, pattern = "*", sources = NULL)

Arguments

conn

A datom_conn object from datom_get_conn().

path

Optional path to input files directory. Defaults to ⁠input_files/⁠ inside the repo. Not accepted on a product repo, which reads no files.

pattern

Glob pattern for file matching. Default "*". On a product repo it filters source artifact names instead.

sources

On a product repo only, and required there: one datom_conn, or a list of them, for the projects the set's inputs come from. Each connection's project name is what members are matched on. Refused on an ordinary repo.

Details

A file counts as changed when its bytes differ from the file last synced under that name.

Files whose format is outside datom's ingestion allowlist (flat tabular formats only) are flagged "unsupported_format" up front, without blocking their allowlisted siblings.

On a product repo (mode: product) it maps the repo's set against source projects instead – see "On a product repo" below.

Value

On an ordinary repo, a data frame with columns: name, file, format, original_file_sha, status (one of "new", "changed", "unchanged", "unsupported_format").

On a product repo, a data frame with columns project, name, kind, version_from (NA for a new artifact), version_to (NA for a member that was not compared) and status (one of "new", "changed", "unchanged", "ambiguous", "not_checked", "excluded"). Versions are full 64-character strings.

On a product repo

A product repo owns one set (named in .datom/project.yaml), and this call compares that set, as stored, with what each source project holds now. It reads one manifest per source and the stored set; it writes nothing.

Tables and sets are treated alike: a set a source holds gets a row, and a member that is a set is compared exactly as a table member is, so a set built from other sets syncs the same way.

One row per artifact in the sources whose name matches pattern:

Plus one row for each member the call did not compare:

The preview never proposes removing a member. A member whose artifact is no longer listed in its source is named in the messages and left pinned. Members in the repo's own project are outputs and get no row: re-derive them, then move them with datom_update_members().

It stops when sources includes the repo's own project, and when a source connection's project name differs from the name that project's own manifest records.

Tables and sets are saved at different points. On an ordinary repo, datom_sync() writes each table as it syncs. On a product repo it hands the edited set back, and the set is saved only by datom_write_set() – one version for the whole edit, which you can look at or add to first.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # Drop a source file into the repo's input_files/ directory.
  file.copy(
    system.file("extdata", "dm.csv", package = "datom"),
    file.path(tmp, "repo", "input_files", "dm.csv")
  )

  manifest <- datom_sync_manifest(conn)
  print(manifest[, c("name", "format", "status")])

  unlink(tmp, recursive = TRUE)
}

Repoint a Set's Members at Newer Versions

Description

Moves members of a set forward to the versions that are current now, and returns the set with those pointers changed. Nothing is written: the report you see is the dry run, and the set is stored only when you pass the result to datom_write_set().

Usage

datom_update_members(
  x,
  conn,
  member = NULL,
  tags = NULL,
  version_from = NULL,
  version_to = NULL
)

Arguments

x

A datom_set, from datom_get_set() or datom_assemble_set().

conn

A datom_conn from datom_get_conn(), or a list of them – one per project the selected members belong to.

member

Optional: the member to repoint, as its name, a member record, or a link. NULL (the default) selects every member.

tags

Optional named list of labels narrowing the selection, e.g. list(type = "input"). A member matches when it carries every label listed.

version_from

Optional version, or a prefix of one, narrowing the selection to the member pinned at it.

version_to

Optional exact version to move to. NULL (the default) means whatever that artifact's project reports as current. Requires the selection to resolve to a single member.

Details

This is the operation a product needs when its inputs move on – a hundred members of which thirty upstream tables have advanced. Doing it by hand is list surgery on what datom_get_set() returned, and two of the obvious spellings are silently wrong: filtering members by name drops every version of that name, and rebuilding a pointer from its name plus a new version drops that member's labels.

Value

x with the matching members repointed, and what changed appended to its datom_edits attribute.

Which connection to pass

One per project the set spans, since access in datom is per project. A set whose members all live in the project that owns it needs only that one connection; a product drawing on three studies needs three, in a list.

A set's projects can be listed offline, with no connection at all:

unique(datom_list_members(x)$project)

Which members move

With no member, every member – refreshing everything is the common case and rerunning it changes nothing. Otherwise select one the way datom_fetch_member() does: by name, by a member record, or by a link. tags and version_from narrow either a name or the sweep, so tags = list(release = "live") repoints the labelled members and leaves the rest pinned.

A name that matches more than one member aborts, because the request cannot be honoured as typed; a sweep that meets the same pair skips it and says so, because refusing a whole refresh over one frozen baseline would make the first update on such a set an error.

Which version each member moves to

version_to omitted, each selected member moves to the version its own project reports as current, and every move is reported before anything is written. That is the one place datom infers "newest", and the reason it is allowed here is in the verb's name: a pointer constructor requires an explicit version (datom_member() refuses to guess), while a verb whose whole meaning is move this forward states the time-dependence up front and then says what it picked. What the set records is still an exact version, so a script that later reads that set is as reproducible as ever.

version_to supplied, the selected member moves to exactly that version – a rollback to a known-good, or a deliberate step to something that is not the newest. It requires the selection to resolve to one member, because one explicit version across several artifacts is not a meaning. It is also what makes this verb better than removing and re-adding a member: the labels come with it, where re-adding makes you retype them.

What it declines to do, and how

Situation Response
a member's project has no supplied connection the whole call is refused, naming that project
a member's artifact no longer appears in its project reported, and its pin is left alone
two members share a name and a project both skipped and reported

The first refuses because whether those members moved is unknowable, and reporting them as unchanged would state something nothing checked. The second does not, because the answer is known: the pinned version is immutable and still reads, so the set stays writable.

What comes back

The set it was handed, with matching members repointed and each moved member's labels byte-identical to what they were.

When something moved, a datom_set's version and data_sha are emptied: they described the payload it was read as, and that is no longer what the object holds. When nothing moved they are left alone, because the object still describes exactly that stored version.

The returned object also carries a log of what changed, which datom_write_set() uses for the commit message when you pass no message of your own – so ⁠git log⁠ names what moved instead of saying ⁠Update {name}⁠. datom_remove_members() adds to the same log, so editing both ways before you write produces one message describing both. Passing x$members rather than x to the write loses that and nothing else.

See Also

datom_remove_members() to drop members instead, datom_write_set() to store the result, datom_list_members() to see what a set holds, datom_member() to build a pointer from scratch.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  # A product repo declares itself as one and names the single set it owns.
  datom_init_repo(file.path(tmp, "repo"), "example_project", store,
                  mode = "product", set = "example_product")

  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  dm <- datom_example_data("dm")
  datom_write(conn, data = dm, name = "dm")
  datom_write_set(conn, list(
    datom_member(conn, "dm", datom_history(conn, "dm")$version[1],
                 tags = list(type = "input"))
  ))

  # The table moves on, so the set now cites an older version of it.
  datom_write(conn, data = dm[-1, , drop = FALSE], name = "dm")

  x <- datom_get_set(conn, "example_product")
  x <- datom_update_members(x, conn)

  # The label came with it, and nothing is stored until the write.
  print(datom_list_members(x))
  datom_write_set(conn, x)

  unlink(tmp, recursive = TRUE)
}

Validate Git-Storage Consistency

Description

Checks that git metadata matches S3 storage for all tables and repo-level files. Reports mismatches as a structured result.

Usage

datom_validate(conn, fix = FALSE)

Arguments

conn

A datom_conn object from datom_get_conn().

fix

If TRUE, attempts to fix inconsistencies by syncing data-side metadata (manifest + per-artifact metadata) to storage, and by restoring a set's payload when storage has lost it – git holds {name}/set.json, so those bytes are recoverable. A restore happens only when the stored object is absent and only when the clone's bytes hash to the document_sha already recorded; a stored payload is never overwritten and its recorded hash is never recomputed.

A missing table payload (data_missing_s3) cannot be repaired: the parquet bytes are never in the clone. Those tables are named in a warning and need datom_write() re-run with the source data.

Value

A list with:

valid

Logical — TRUE if everything is consistent.

repo_files

Data frame of repo-level file checks.

tables

Data frame of per-artifact checks, one row per artifact of either kind, with a kind column. Named tables for compatibility.

fixed

Logical — TRUE if fix = TRUE was applied.

What is checked per artifact

Both kinds of artifact are checked, and the payload check branches on kind: a table's payload is a parquet object, a set's is a JSON document at ⁠{name}/{data_sha}.json⁠. A set is checked further, because a payload whose members have gone is a citation that no longer resolves:

Statuses reported in the tables frame: metadata_missing_s3, history_missing_s3, data_missing_s3, members_unresolvable, document_sha_missing, and kind_unsupported for an artifact whose metadata declares a kind this version of datom does not know – reported rather than fatal, with that row's payload left unchecked.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  datom_validate(conn)

  unlink(tmp, recursive = TRUE)
}

Save a Data Frame as a datom Table

Description

Saves a data frame as a new version of a named table: the data goes to storage, and a record of the change is committed and pushed to the project's GitHub repository. If nothing has changed since the last version, nothing is saved. To bring in files rather than data frames, use datom_sync().

Usage

datom_write(
  conn,
  data = NULL,
  name = NULL,
  metadata = NULL,
  message = NULL,
  parents = NULL,
  .source_lineage = NULL,
  .table_type = "derived",
  .original_file_sha = NULL,
  .original_format = NULL
)

Arguments

conn

A datom_conn object from datom_get_conn().

data

Data frame to write. If NULL with name, does metadata-only sync.

name

Table name. If NULL with NULL data, mirrors the clone's storage-side documents for every artifact of either kind: the manifest, and each artifact's metadata, version history and versioned snapshots. On that route a set whose stored payload is missing also has it restored from the clone – see datom_validate(), which shares the mechanism, for the conditions on that.

metadata

Optional list of custom metadata.

message

Optional commit message.

parents

Optional list of parent records produced by datom_parent(), each carrying source, table, version, data_sha, and source_lineage. When supplied, the table's source_lineage is derived as the deduplicated union of the parents' source_lineage and each parent is recorded lean (source, table, version, data_sha). NULL if no lineage is recorded. There is no public source_lineage parameter; it is always derived from parents.

.source_lineage

Internal. Flat list of transitive non-derived source descriptors (each with project, table, version_sha) for the imported self-entry path, set by datom_sync(). Unused on the derived (parents) path.

.table_type

Internal. "derived" (default) or "imported" (set by datom_sync()).

.original_file_sha

Internal. SHA of source file (set by datom_sync()); NULL for derived.

.original_format

Internal. Original file format (set by datom_sync()); NULL for derived.

Value

List with deployment details.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # --- Basic write (no lineage) ---
  dm <- datom_example_data("dm")
  datom_write(conn, data = dm, name = "dm")

  # --- Write with a single parent ---
  # Each parent's data_sha and lineage are resolved by datom_parent.
  lb <- datom_example_data("lb")
  datom_write(conn, data = lb, name = "lb")
  lb_summary <- aggregate(
    list(n = lb$LBTESTCD), by = list(LBTESTCD = lb$LBTESTCD), FUN = length
  )
  datom_write(
    conn,
    data    = lb_summary,
    name    = "lb_summary",
    message = "Lab test counts",
    parents = list(
      datom_parent(conn, "lb", datom_history(conn, "lb")$version[1])
    )
  )

  # --- Write with multiple parents ---
  # The source lineage is derived as the union of the parents' lineages.
  dm_lb_merged <- merge(dm, lb, by = "USUBJID")
  datom_write(
    conn,
    data    = dm_lb_merged,
    name    = "dm_lb_merged",
    message = "Demographics joined with lab results",
    parents = list(
      datom_parent(conn, "dm", datom_history(conn, "dm")$version[1]),
      datom_parent(conn, "lb", datom_history(conn, "lb")$version[1])
    )
  )

  print(datom_list(conn))

  unlink(tmp, recursive = TRUE)
}

Save a Set as a New Version

Description

Saves a set: a named list of exact versions of tables (or other sets), plus labels, that you can cite with one version string. A set holds no data, so saving one copies nothing. Saving the same members and labels again creates no new version.

Usage

datom_write_set(
  conn,
  members,
  tags = NULL,
  name = NULL,
  message = NULL,
  include_paths = NULL
)

Arguments

conn

A datom_conn object from datom_get_conn(), scoped to the product repo (developer role).

members

A list of member records from datom_member(), each pinning one artifact version and optionally carrying its own tags – or a datom_set, from datom_get_set() or datom_assemble_set(), edited or not. Hand-assembled lists are refused.

tags

Optional named list of set-level text labels – facts about the collection itself, such as a description. Same grammar as a member's tags: a value is one string or several, text only.

name

The set's name. Defaults to the ⁠set:⁠ field in .datom/project.yaml; when supplied it must equal it.

message

Optional commit message. Omitted, it is ⁠Update {name}⁠ – except for a set edited with datom_add_member(), datom_update_members() or datom_remove_members(), where the default names the members added, moved or dropped, with their versions. So a set assembled in steps gets ⁠Update {name}: add N members⁠ on its first write. Pass x rather than x$members to get that, since the change list travels with the object.

include_paths

Optional character vector of repo-relative paths – your own code, renv.lock, build state – staged into the same commit as the set. Never mirrored to storage: the storage namespace holds datom artifacts and nothing else. See the section below.

Details

One repo holds one set. The repo declares which, in .datom/project.yaml:

mode: product
set: study001-adam

Both are checked before anything is hashed or written, so a repo that has not declared itself a product repo is refused with nothing left behind.

Value

Invisibly, a list with name, data_sha, metadata_sha (the version), member_count (the count after normalisation), action ("none" or "full") and commit_sha.

What a set carries, and what it does not

User metadata is tags, and there is no ⁠metadata =⁠ parameter: a description is a tag, and a second channel for the same thing would be two places to look. There is no view or navigation configuration either – a folder-like hierarchy is a projection a consumer computes over tags, and any number of them cost nothing precisely because none is stored.

A set records no parents and no source_lineage. Members are references, not derivation: lineage flows through tables, and the set is how you found a table rather than how data reached it.

Versions, and what moves one

The version covers the whole payload, members and tags alike. So editing a tag or a description mints a new version, which is intended: a set exists to be citable, and "same citation, different labels" would be a lie to whoever cited it. What does not mint a version is a purely syntactic edit – reordering tag values or members, repeating a label, or writing a single label as a one-element array. Those are normalised on the way in, so re-writing an identical payload is a no-op.

Your code does not move a version either, even though it travels in the same commit (see include_paths below). Refactor your build script, re-run it, get the same members and tags, and nothing is minted: the write is the usual no-op. datom_history() then shows the version it showed before, with a commit_sha pointing at the commit that first produced that payload – a commit that does not contain the code you just wrote. That is the recorded value doing its job rather than going stale; see datom_history() for why the commit is deliberately not part of the version.

Where the payload lives

Two copies, at two deliberately different addresses. Git holds {name}/set.json at one stable path, modified in place, so git carries the history and ⁠git diff⁠ between two versions shows which members changed. Storage holds the same bytes content-addressed at ⁠{name}/{data_sha}.json⁠, so a reader with no clone can fetch an exact version. Any past version is still reconstructible from the clone alone with ⁠git show <commit>:{name}/set.json⁠.

Carrying your code and environment into the same commit

include_paths stages paths you name into the one commit that carries the payload and its metadata. So checking out a set version's commit yields the data pointers, the logic that produced them and the environment they ran in – one clone, one checkout, the whole product. The joint version is structural: nothing records a link between the set and your files, because the commit is the link.

datom_write_set(conn, members,
                include_paths = c("R", "dp", "renv.lock"))

Four refusals, all of them before anything is hashed or written, so a refusal leaves nothing behind: a path that does not exist, a path outside the clone, a path datom owns (⁠.datom/⁠, the set, any artifact directory), and a path .gitignore excludes. The last one matters because git stages an ignored path silently and without complaint, which would leave the set version claiming a commit that omits exactly the file you named. Refusals win over the no-op below, since they are settled before change detection runs.

An unchanged set is still a no-op, however dirty those paths are. No commit, no version, and a message pointing at datom_repo_commit(), which is the verb for committing your own content at a moment you chose. A data write that quietly committed work in progress is the thing datom's explicit file lists exist to prevent, and idempotency must not become a side door into it.

Outputs must be built from the inputs the set pins

A table written with parents records which versions it was derived from. When a set lists such a table and also lists one of its parents, the write checks they agree: if the set pins the parent at a different version from the one the table was built from, and not at that version too, the write stops and names the member, the parent and both versions. Nothing is written. Re-derive the output with datom_parent(x = ), which takes the versions from the set, or move the input with datom_update_members().

Only tables in the set's own project are checked, and a parent the set does not list is not checked. A set carrying one table at two versions (a live copy beside a frozen baseline) passes as long as one of them is the version used. Each such member's recorded metadata is read from storage; if it cannot be read – a version that does not exist, or storage that cannot be reached – the write stops too.

Editing a set that already exists

Read it, change it, write it back. members accepts a datom_set from datom_get_set() directly, so the loop needs no unpacking:

x <- datom_get_set(conn, "study001-adam")
x$members <- c(x$members, list(datom_member(conn, "lb", v)))
datom_write_set(conn, x)

The set's own tags come along with it unless tags is supplied, so a read-append-write cannot silently drop the description. Passing x$members instead works too, and there tags is yours to carry.

A set built in steps with datom_assemble_set() and datom_add_member() is written the same way, and pipes into the write:

x |> datom_write_set(conn = conn)

A set is written into its own repo

A datom_set records its name and the project it belongs to. Both are checked before anything is hashed or written: a set named for another repo's declared set, or belonging to another project, stops the write. Writing it anyway would move it into this repo's project without saying so – the name check alone would miss that, since two product repos may declare the same set name. A name argument that disagrees with the set's own name stops it too. A plain list of member records carries neither, so neither is checked.

See Also

datom_member() to declare a member, datom_assemble_set() to build a set a member at a time, datom_write() for tables.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  # A product repo declares itself as one and names the single set it owns.
  datom_init_repo(file.path(tmp, "repo"), "example_project", store,
                  mode = "product", set = "example_product")

  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  # A set points at versions that already exist.
  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  datom_write(conn, data = datom_example_data("lb"), name = "lb")

  members <- list(
    datom_member(conn, "dm", datom_history(conn, "dm")$version[1],
                 tags = list(type = "input")),
    datom_member(conn, "lb", datom_history(conn, "lb")$version[1],
                 tags = list(type = "output", domain = c("safety", "labs")))
  )

  datom_write_set(
    conn, members,
    tags = list(description = "Example product for STUDY-001")
  )

  print(datom_list(conn))

  unlink(tmp, recursive = TRUE)
}

Check if Object is a datom Connection

Description

Check if Object is a datom Connection

Usage

is_datom_conn(x)

Arguments

x

Object to test.

Value

TRUE or FALSE.


Check if Object is a datom Store

Description

Check if Object is a datom Store

Usage

is_datom_store(x)

Arguments

x

Object to test.

Value

TRUE or FALSE.

Examples

tmp <- tempfile("datom_store_")
store <- datom_store(
  data = datom_store_local(path = tmp),
  data_repo_url = "https://github.com/example/my-project",
  validate = FALSE
)
is_datom_store(store)
is_datom_store("not a store")
unlink(tmp, recursive = TRUE)

Check if Object is a Local Store Component

Description

Check if Object is a Local Store Component

Usage

is_datom_store_local(x)

Arguments

x

Object to test.

Value

TRUE or FALSE.

Examples

tmp <- tempfile("datom_store_")
store <- datom_store_local(path = tmp, validate = TRUE)
is_datom_store_local(store)
is_datom_store_local("not a store")
unlink(tmp, recursive = TRUE)

Check if Object is an S3 Store Component

Description

Check if Object is an S3 Store Component

Usage

is_datom_store_s3(x)

Arguments

x

Object to test.

Value

TRUE or FALSE.

Examples

s3 <- datom_store_s3(
  bucket = "my-datom-bucket",
  access_key = "AKIAIOSFODNN7EXAMPLE",
  secret_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
  validate = FALSE
)
is_datom_store_s3(s3)
is_datom_store_s3("not a store")

Check if Object is a Credentials-Only S3 Store Component

Description

Check if Object is a Credentials-Only S3 Store Component

Usage

is_datom_store_s3_creds(x)

Arguments

x

Object to test.

Value

TRUE or FALSE.

Examples

creds <- datom_store_s3_creds(
  access_key = "AKIAIOSFODNN7EXAMPLE",
  secret_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
)
is_datom_store_s3_creds(creds)
is_datom_store_s3_creds("not a store")

Check if Path is a Valid datom Repository

Description

Validates datom repository structure. Used internally and by dpbuild.

Usage

is_valid_datom_repo(
  path,
  checks = c("all", "git", "datom", "renv"),
  verbose = FALSE
)

Arguments

path

Path to evaluate.

checks

Which checks to perform. Any combination of "all", "git", "datom", "renv".

verbose

If TRUE, prints which tests passed/failed.

Value

TRUE or FALSE.

Examples

# A plain directory is not a valid datom repository.
tmp <- tempfile("datom_valid_")
dir.create(tmp)
is_valid_datom_repo(tmp)
unlink(tmp, recursive = TRUE)

Create a datom Connection Object

Description

Internal constructor for the datom_conn S3 class. Two modes:

Usage

new_datom_conn(
  project_name,
  root,
  prefix = NULL,
  region = "us-east-1",
  client,
  path = NULL,
  role = c("reader", "developer"),
  endpoint = NULL,
  gov_root = NULL,
  gov_prefix = NULL,
  gov_region = NULL,
  gov_backend = NULL,
  gov_client = NULL,
  gov_local_path = NULL,
  backend = "s3",
  data_repo_url = NULL,
  github_pat = NULL,
  github_api_url = NULL,
  min_writer_version = NULL
)

Arguments

project_name

Project name string.

root

Storage root (S3 bucket name or local directory path).

prefix

Storage prefix (can be NULL).

region

AWS region string (data store). Ignored for local backend.

client

A storage client (paws S3 client or NULL for local).

path

Local repo path (NULL for readers).

role

One of "developer" or "reader".

endpoint

Optional S3 endpoint URL (e.g., for S3 access points). NULL for default.

gov_root

Governance storage root (can be NULL for legacy conns).

gov_prefix

Governance prefix (can be NULL).

gov_region

Governance region (can be NULL).

gov_backend

Governance storage backend ("s3" or "local"), set from the governance store component. NULL on solo (no-governance) conns. Independent of backend (the data backend): a project may keep data on one backend and governance on another.

gov_client

Governance storage client (can be NULL).

gov_local_path

Absolute path to the local gov clone (NULL for readers).

data_repo_url

HTTPS URL of the data GitHub repository. Populated at conn-construction time from the git remote or store. NULL for readers or when not yet known.

github_pat

GitHub personal access token held in memory only. Sourced from store$github_pat at conn-construction time. Never persisted to disk and never printed.

github_api_url

GitHub API base URL. Sourced from store$github_api_url at conn-construction time. Defaults to "https://api.github.com" when not set.

min_writer_version

The lowest version of datom this repo accepts writes from, read from project.yaml at conn-construction time. NULL means the repo declares no such limit, which is every repo written so far. Held on the connection because the file it comes from is already parsed there, so the write-entry check costs no extra read.

Details

The primary fields (root, prefix, region, client) refer to the data store. Governance store fields are prefixed with gov_.

Value

A datom_conn object.


Print a datom Connection

Description

Displays a clean summary without exposing credentials or the S3 client.

Usage

## S3 method for class 'datom_conn'
print(x, ...)

Arguments

x

A datom_conn object.

...

Ignored.

Value

Invisible x.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  print(conn)

  unlink(tmp, recursive = TRUE)
}

Description

Print a Member Link

Usage

## S3 method for class 'datom_link'
print(x, ...)

Arguments

x

A datom_link from a member of a set read with datom_get_set().

...

Ignored.

Value

Invisible x.

Examples

# See datom_get_set() for a runnable set example; a link is one of its
# members' `$fetch` elements.
print(names(formals(datom_get_set)))

Print a datom Set

Description

One line per member – name, kind, and its tags as compact key=value pairs, or - when it has none – plus the route to a member's content. Long member lists are truncated. A set not yet written shows version NA.

Usage

## S3 method for class 'datom_set'
print(x, ..., n = 20L)

Arguments

x

A datom_set, from datom_get_set() or datom_assemble_set().

...

Ignored.

n

Maximum number of members to list.

Details

Tags are open-keyed by design, so there is no fixed column layout to print them in.

Value

Invisible x.

Examples

# See datom_get_set() for a runnable example that prints a set.
print(names(formals(datom_get_set)))

Print a datom Store

Description

Displays store configuration with masked secrets.

Usage

## S3 method for class 'datom_store'
print(x, ...)

Arguments

x

A datom_store object.

...

Ignored.

Value

Invisible x.

Examples

tmp <- tempfile("datom_store_")
store <- datom_store(
  data = datom_store_local(path = tmp),
  data_repo_url = "https://github.com/example/my-project",
  validate = FALSE
)
print(store)
unlink(tmp, recursive = TRUE)

Print a Local Store Component

Description

Displays store configuration.

Usage

## S3 method for class 'datom_store_local'
print(x, ...)

Arguments

x

A datom_store_local object.

...

Ignored.

Value

Invisible x.

Examples

tmp <- tempfile("datom_store_")
store <- datom_store_local(path = tmp, validate = TRUE)
print(store)
unlink(tmp, recursive = TRUE)

Print an S3 Store Component

Description

Displays store configuration with masked secrets.

Usage

## S3 method for class 'datom_store_s3'
print(x, ...)

Arguments

x

A datom_store_s3 object.

...

Ignored.

Value

Invisible x.

Examples

s3 <- datom_store_s3(
  bucket = "my-datom-bucket",
  access_key = "AKIAIOSFODNN7EXAMPLE",
  secret_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY",
  validate = FALSE
)
print(s3)

Print a Credentials-Only S3 Store Component

Description

Displays masked credentials and a note that location is resolved from ref.json at connection time.

Usage

## S3 method for class 'datom_store_s3_creds'
print(x, ...)

Arguments

x

A datom_store_s3_creds object.

...

Ignored.

Value

Invisible x.

Examples

creds <- datom_store_s3_creds(
  access_key = "AKIAIOSFODNN7EXAMPLE",
  secret_key = "wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY"
)
print(creds)

Print a datom_summary

Description

Print a datom_summary

Usage

## S3 method for class 'datom_summary'
print(x, ...)

Arguments

x

A datom_summary object.

...

Ignored.

Value

Invisible x.

Examples

# Offline, self-contained: a bare git repo stands in for GitHub and a
# local directory for object storage.
if (requireNamespace("git2r", quietly = TRUE)) {
  tmp <- tempfile("datom-example-")
  remote <- file.path(tmp, "remote.git")
  dir.create(remote, recursive = TRUE)
  git2r::init(remote, bare = TRUE)

  store <- datom_store(
    data = datom_store_local(file.path(tmp, "storage")),
    github_pat = "example-token", # role selector; a local remote needs none
    data_repo_url = remote,
    validate = FALSE
  )
  datom_init_repo(file.path(tmp, "repo"), "example_project", store)
  conn <- datom_get_conn(file.path(tmp, "repo"), store)

  datom_write(conn, data = datom_example_data("dm"), name = "dm")
  print(datom_summary(conn))

  unlink(tmp, recursive = TRUE)
}

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.