Starting on S3

Goal: Stand up a versioned datom project whose data lives in Amazon S3, and onboard a study’s files with datom_sync(). The steps are the ones from Getting Started – one file, an update, a no-op, a batch – and the only difference is the store you build.

Want to try datom locally first? Start with Getting Started. It uses a folder on your machine instead of S3 and needs no AWS account. The functions are the same.

You look after the data for study001, a clinical trial, and your team already works in S3. You want the first extract to land in the shared bucket, versioned from day one. datom keeps the data in S3 and the record of every version in a git repository, so the history can be read and reproduced from any machine.

Two locations, two roles:

Requirements

Store the three secrets in your keychain once:

keyring::key_set(service = "GITHUB_PAT")
keyring::key_set(service = "AWS_ACCESS_KEY_ID")
keyring::key_set(service = "AWS_SECRET_ACCESS_KEY")

Load your secrets

This is the only place the keychain is read. Everything after it reads the environment with Sys.getenv(), so the secrets never appear in your code.

Sys.setenv(
  GITHUB_PAT            = keyring::key_get(service = "GITHUB_PAT"),
  AWS_ACCESS_KEY_ID     = keyring::key_get(service = "AWS_ACCESS_KEY_ID"),
  AWS_SECRET_ACCESS_KEY = keyring::key_get(service = "AWS_SECRET_ACCESS_KEY")
)

On CI or in a container, set these three environment variables through your platform’s secret store and skip this chunk.

Settings

Every value used more than once is set here, so changing one means changing it in one place.

library(datom)

# --- Settings you control ----------------------------------------------------
bucket           <- "study001"            # one bucket per study
region           <- "us-east-1"

project_imported <- "study001-imported"   # recorded in the project's metadata
prefix_imported  <- "imported/"           # this project's folder in the bucket
repo_imported    <- "study001-imported"   # GitHub repo name

# Local working folder for the metadata repository. The data never lands here;
# it goes straight to S3.
workdir_imported <- fs::path(tempdir(), "study001-imported")

The project, repo and folder names match by convention only. They are separate settings because they do not have to match.

Where the data goes

s3://study001/
    imported/datom/        onboarded tables        (this article)

imported is datom’s word for a table that came in from a file. datom manages the datom/ folder under each prefix: it reads and writes the data, metadata and version records there, and leaves anything else in the bucket alone.

Build the store

A store says where the data lives and how to reach it. Giving it a GitHub token makes it a writer store: it can create the metadata repository and record new versions.

store_write_imported <- datom_store(
  data = datom_store_s3(
    bucket     = bucket,
    prefix     = prefix_imported,
    region     = region,
    access_key = Sys.getenv("AWS_ACCESS_KEY_ID"),
    secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY")
  ),
  github_pat = Sys.getenv("GITHUB_PAT")
)

datom checks that it can reach the bucket as the store is built, so a wrong credential shows up here rather than at the first sync.

Create the repository

datom_init_repo(
  path         = workdir_imported,
  project_name = project_imported,
  store        = store_write_imported,
  create_repo  = TRUE,
  repo_name    = repo_imported
)
#> v Created GitHub repo ".../study001-imported".
#> v Initialized datom repository "study001-imported" at '.../study001-imported'

This creates the GitHub repository, clones it into workdir_imported, and commits a project.yaml recording where the project’s data lives. No data goes to GitHub, only metadata.

There is no mode argument here. Leaving it out gives an ordinary repository, one that takes files in with datom_sync().

The clone also gets an input_files/ folder. git ignores it, so nothing placed there is ever committed. It is where datom_sync() looks for new files.

Connect

conn_write_imported <- datom_get_conn(
  path  = workdir_imported,
  store = store_write_imported
)
print(conn_write_imported)
#> 
#> -- datom connection
#> * Project: "study001-imported"
#> * Backend: "s3"
#> * Role: "developer"
#> * Data root: "study001"
#> * Data prefix: "imported/"
#> * Data region: "us-east-1"
#> * Governance: not attached
#> * Path: '.../study001-imported'
#> * Data repo: <https://github.com/.../study001-imported.git>

Step 1: Sync one file

The month-1 extract of the demographics table has arrived. Put it in the input folder:

inputs_imported <- fs::path(workdir_imported, "input_files")

write.csv(
  x         = datom_example_data(domain = "dm", cutoff_date = "2026-01-28"),
  file      = fs::path(inputs_imported, "dm.csv"),
  row.names = FALSE
)

Scan the folder, then sync:

manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 1 new, 0 changed, 0 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 1 table...
#> v Wrote "dm" (full): "153bff41"
#> v "dm" synced (new).
#> i Sync complete: 1 succeeded, 0 failed, 0 skipped.

The file was converted to parquet and uploaded to S3. The version record was committed to the metadata repository and pushed to GitHub. datom_sync() returns the manifest with a result for each file, kept here as synced.

datom_list(conn = conn_write_imported)
#>   name  kind current_version current_data_sha         last_updated
#> 1   dm table        153bff41         decbafd2 2026-09-27T05:58:03Z

dm_history <- datom_history(conn = conn_write_imported, name = "dm",
                            short_hash = TRUE)
dm_history[, c("version", "timestamp", "commit_message")]
#>    version            timestamp commit_message
#> 1 153bff41 2026-09-27T05:58:03Z  Sync dm (new)

The input file is no longer needed. The table reads from S3:

fs::file_delete(fs::path(inputs_imported, "dm.csv"))
nrow(datom_read(conn = conn_write_imported, name = "dm"))
#> [1] 4

Step 2: Update one file

The month-2 extract arrives with new subjects:

write.csv(
  x         = datom_example_data(domain = "dm", cutoff_date = "2026-02-28"),
  file      = fs::path(inputs_imported, "dm.csv"),
  row.names = FALSE
)

manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 0 new, 1 changed, 0 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 1 table...
#> v Wrote "dm" (full): "0fac26cd"
#> v "dm" synced (changed).
#> i Sync complete: 1 succeeded, 0 failed, 0 skipped.

Both versions stay readable. The newest is the default; an older one is read by its version. The first 8 characters of a version are enough, as with a git commit:

dm_history <- datom_history(conn = conn_write_imported, name = "dm",
                            short_hash = TRUE)
dm_history[, c("version", "timestamp", "commit_message")]
#>    version            timestamp    commit_message
#> 1 0fac26cd 2026-09-27T05:58:10Z Sync dm (changed)
#> 2 153bff41 2026-09-27T05:58:03Z     Sync dm (new)

nrow(datom_read(conn = conn_write_imported, name = "dm"))
#> [1] 16

dm_version <- dm_history$version[nrow(dm_history)]   # oldest row: month 1
nrow(datom_read(conn = conn_write_imported, name = "dm", version = dm_version))
#> [1] 4

Step 3: Sync again with nothing new

manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 1 file: 0 new, 0 changed, 1 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i No new or changed files. Nothing to sync.

A file datom already holds is not uploaded again, so the sync is safe to run on a schedule.

Step 4: A batch of files

The month-3 extract brings four tables at once: demographics (dm), dosing (ex), labs (lb) and adverse events (ae).

for (domain in c("dm", "ex", "lb", "ae")) {
  write.csv(
    x         = datom_example_data(domain = domain, cutoff_date = "2026-03-28"),
    file      = fs::path(inputs_imported, paste0(domain, ".csv")),
    row.names = FALSE
  )
}

manifest <- datom_sync_manifest(conn = conn_write_imported)
#> i Scanned 4 files: 3 new, 1 changed, 0 unchanged.

synced <- datom_sync(conn = conn_write_imported, manifest = manifest)
#> i Syncing 4 tables...
#> v Wrote "ae" (full): "075773e9"
#> v "ae" synced (new).
#> v Wrote "dm" (full): "773e6862"
#> v "dm" synced (changed).
#> v Wrote "ex" (full): "8dbcc9a7"
#> v "ex" synced (new).
#> v Wrote "lb" (full): "435bccb0"
#> v "lb" synced (new).
#> i Sync complete: 4 succeeded, 0 failed, 0 skipped.

All four tables are now versioned in S3:

datom_list(conn = conn_write_imported)
#>   name  kind current_version current_data_sha         last_updated
#> 1   dm table        773e6862         e547f03d 2026-09-27T05:58:22Z
#> 2   ae table        075773e9         d5f8dd5a 2026-09-27T05:58:17Z
#> 3   ex table        8dbcc9a7         ab96afc3 2026-09-27T05:58:26Z
#> 4   lb table        435bccb0         5d419c60 2026-09-27T05:58:31Z

Reading as a reader

A colleague who only reads needs bucket credentials and nothing else: no GitHub token and no clone. A store without a token is a reader store.

store_read_imported <- datom_store(
  data = datom_store_s3(
    bucket     = bucket,
    prefix     = prefix_imported,
    region     = region,
    access_key = Sys.getenv("AWS_ACCESS_KEY_ID"),
    secret_key = Sys.getenv("AWS_SECRET_ACCESS_KEY")
  )
)

conn_read_imported <- datom_get_conn(
  store        = store_read_imported,
  project_name = project_imported
)

nrow(datom_read(conn = conn_read_imported, name = "lb"))
#> [1] 205

The read goes straight to S3, not through GitHub.

Where you are

What next

Governance and migration come later

Two things are deliberately left out of this article:

Teardown

Delete the project’s storage first, then its repository:

datom_storage_delete_prefix(conn = conn_write_imported)
datom_repo_delete(conn = conn_write_imported, confirm = project_imported)

datom_storage_delete_prefix() deletes everything under imported/datom/, and it does not ask first. The rest of the bucket is left alone, and so is the bucket itself.

datom_repo_delete() deletes the GitHub repository and the local clone. It needs the project name as confirm, and it does not touch storage.