The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
Goal: Stand up a versioned datom project whose data
lives in Amazon S3, and onboard a study’s files with
datom_sync(). The steps are the ones from Getting Started – one file, an update, a
no-op, a batch – and the only difference is the store you build.
Want to try datom locally first? Start with Getting Started. It uses a folder on your machine instead of S3 and needs no AWS account. The functions are the same.
You look after the data for study001, a clinical trial, and your team already works in S3. You want the first extract to land in the shared bucket, versioned from day one. datom keeps the data in S3 and the record of every version in a git repository, so the history can be read and reproduced from any machine.
Two locations, two roles:
datom_sync() takes a
file in, S3 holds a versioned copy. The input file is no longer
needed.repo. datom creates the metadata repository
through the GitHub API with it; the gh CLI is not
needed.git2r, rio and
keyring packages. datom uses git2r
for the metadata repository and rio to read the files you
sync; keyring holds your secrets in the OS keychain.Store the three secrets in your keychain once:
This is the only place the keychain is read. Everything after it
reads the environment with Sys.getenv(), so the secrets
never appear in your code.
Sys.setenv(
GITHUB_PAT = keyring::key_get(service = "GITHUB_PAT"),
AWS_ACCESS_KEY_ID = keyring::key_get(service = "AWS_ACCESS_KEY_ID"),
AWS_SECRET_ACCESS_KEY = keyring::key_get(service = "AWS_SECRET_ACCESS_KEY")
)On CI or in a container, set these three environment variables through your platform’s secret store and skip this chunk.
Every value used more than once is set here, so changing one means changing it in one place.
library(datom)
# --- Settings you control ----------------------------------------------------
bucket <- "study001" # one bucket per study
region <- "us-east-1"
project_imported <- "study001-imported" # recorded in the project's metadata
prefix_imported <- "imported/" # this project's folder in the bucket
repo_imported <- "study001-imported" # GitHub repo name
# Local working folder for the metadata repository. The data never lands here;
# it goes straight to S3.
workdir_imported <- fs::path(tempdir(), "study001-imported")The project, repo and folder names match by convention only. They are separate settings because they do not have to match.
s3://study001/
imported/datom/ onboarded tables (this article)
imported is datom’s word for a table that came in from a
file. datom manages the datom/ folder under each prefix: it
reads and writes the data, metadata and version records there, and
leaves anything else in the bucket alone.
A store says where the data lives and how to reach it. Giving it a GitHub token makes it a writer store: it can create the metadata repository and record new versions.
store_write_imported <- datom_store(
data = datom_store_s3(
bucket = bucket,
prefix = prefix_imported,
region = region,
access_key = Sys.getenv("AWS_ACCESS_KEY_ID"),
secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY")
),
github_pat = Sys.getenv("GITHUB_PAT")
)datom checks that it can reach the bucket as the store is built, so a wrong credential shows up here rather than at the first sync.
datom_init_repo(
path = workdir_imported,
project_name = project_imported,
store = store_write_imported,
create_repo = TRUE,
repo_name = repo_imported
)
#> v Created GitHub repo ".../study001-imported".
#> v Initialized datom repository "study001-imported" at '.../study001-imported'This creates the GitHub repository, clones it into
workdir_imported, and commits a project.yaml
recording where the project’s data lives. No data goes to GitHub, only
metadata.
There is no mode argument here. Leaving it out gives an
ordinary repository, one that takes files in with
datom_sync().
The clone also gets an input_files/ folder. git ignores
it, so nothing placed there is ever committed. It is where
datom_sync() looks for new files.
conn_write_imported <- datom_get_conn(
path = workdir_imported,
store = store_write_imported
)
print(conn_write_imported)
#>
#> -- datom connection
#> * Project: "study001-imported"
#> * Backend: "s3"
#> * Role: "developer"
#> * Data root: "study001"
#> * Data prefix: "imported/"
#> * Data region: "us-east-1"
#> * Governance: not attached
#> * Path: '.../study001-imported'
#> * Data repo: <https://github.com/.../study001-imported.git>The month-1 extract of the demographics table has arrived. Put it in the input folder:
inputs_imported <- fs::path(workdir_imported, "input_files")
write.csv(
x = datom_example_data(domain = "dm", cutoff_date = "2026-01-28"),
file = fs::path(inputs_imported, "dm.csv"),
row.names = FALSE
)Scan the folder, then sync:
manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 1 new, 0 changed, 0 unchanged.
synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 1 table...
#> v Wrote "dm" (full): "153bff41"
#> v "dm" synced (new).
#> i Sync complete: 1 succeeded, 0 failed, 0 skipped.The file was converted to parquet and uploaded to S3. The version
record was committed to the metadata repository and pushed to GitHub.
datom_sync() returns the manifest with a
result for each file, kept here as synced.
datom_list(conn = conn_write_imported)
#> name kind current_version current_data_sha last_updated
#> 1 dm table 153bff41 decbafd2 2026-09-27T05:58:03Z
dm_history <- datom_history(conn = conn_write_imported, name = "dm",
short_hash = TRUE)
dm_history[, c("version", "timestamp", "commit_message")]
#> version timestamp commit_message
#> 1 153bff41 2026-09-27T05:58:03Z Sync dm (new)The input file is no longer needed. The table reads from S3:
The month-2 extract arrives with new subjects:
write.csv(
x = datom_example_data(domain = "dm", cutoff_date = "2026-02-28"),
file = fs::path(inputs_imported, "dm.csv"),
row.names = FALSE
)
manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 0 new, 1 changed, 0 unchanged.
synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 1 table...
#> v Wrote "dm" (full): "0fac26cd"
#> v "dm" synced (changed).
#> i Sync complete: 1 succeeded, 0 failed, 0 skipped.Both versions stay readable. The newest is the default; an older one is read by its version. The first 8 characters of a version are enough, as with a git commit:
dm_history <- datom_history(conn = conn_write_imported, name = "dm",
short_hash = TRUE)
dm_history[, c("version", "timestamp", "commit_message")]
#> version timestamp commit_message
#> 1 0fac26cd 2026-09-27T05:58:10Z Sync dm (changed)
#> 2 153bff41 2026-09-27T05:58:03Z Sync dm (new)
nrow(datom_read(conn = conn_write_imported, name = "dm"))
#> [1] 16
dm_version <- dm_history$version[nrow(dm_history)] # oldest row: month 1
nrow(datom_read(conn = conn_write_imported, name = "dm", version = dm_version))
#> [1] 4manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 0 new, 0 changed, 1 unchanged.
synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i No new or changed files. Nothing to sync.A file datom already holds is not uploaded again, so the sync is safe to run on a schedule.
The month-3 extract brings four tables at once: demographics
(dm), dosing (ex), labs (lb) and
adverse events (ae).
for (domain in c("dm", "ex", "lb", "ae")) {
write.csv(
x = datom_example_data(domain = domain, cutoff_date = "2026-03-28"),
file = fs::path(inputs_imported, paste0(domain, ".csv")),
row.names = FALSE
)
}
manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 4 files: 3 new, 1 changed, 0 unchanged.
synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 4 tables...
#> v Wrote "ae" (full): "075773e9"
#> v "ae" synced (new).
#> v Wrote "dm" (full): "773e6862"
#> v "dm" synced (changed).
#> v Wrote "ex" (full): "8dbcc9a7"
#> v "ex" synced (new).
#> v Wrote "lb" (full): "435bccb0"
#> v "lb" synced (new).
#> i Sync complete: 4 succeeded, 0 failed, 0 skipped.All four tables are now versioned in S3:
datom_list(conn = conn_write_imported)
#> name kind current_version current_data_sha last_updated
#> 1 dm table 773e6862 e547f03d 2026-09-27T05:58:22Z
#> 2 ae table 075773e9 d5f8dd5a 2026-09-27T05:58:17Z
#> 3 ex table 8dbcc9a7 ab96afc3 2026-09-27T05:58:26Z
#> 4 lb table 435bccb0 5d419c60 2026-09-27T05:58:31ZA colleague who only reads needs bucket credentials and nothing else: no GitHub token and no clone. A store without a token is a reader store.
store_read_imported <- datom_store(
data = datom_store_s3(
bucket = bucket,
prefix = prefix_imported,
region = region,
access_key = Sys.getenv("AWS_ACCESS_KEY_ID"),
secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY")
)
)
conn_read_imported <- datom_get_conn(
store = store_read_imported,
project_name = project_imported
)
nrow(datom_read(conn = conn_read_imported, name = "lb"))
#> [1] 205The read goes straight to S3, not through GitHub.
s3://study001/imported/datom/.Two things are deliberately left out of this article:
datomanager, and starting on S3 does not
commit you to it.datomanager workflow. You started on S3, so you do not need
it now.Delete the project’s storage first, then its repository:
datom_storage_delete_prefix(conn = conn_write_imported)
datom_repo_delete(conn = conn_write_imported, confirm = project_imported)datom_storage_delete_prefix() deletes everything under
imported/datom/, and it does not ask first. The rest of the
bucket is left alone, and so is the bucket itself.
datom_repo_delete() deletes the GitHub repository and
the local clone. It needs the project name as confirm, and
it does not touch storage.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.