The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

Getting started

delta.sharing connects R to tables exposed through Delta Sharing. A typical workflow is to create a client, select a table, describe a read, and choose the form of the result.

Try the public example data

The Delta Sharing project hosts an open server that can be used without registering or creating a private credential:

library(delta.sharing)

client <- sharing_client(demo_profile())

housing <- client$table("delta_sharing.default.boston-housing")
housing$snapshot(
  columns = c("chas", "medv"),
  limit = 5
)$to_tibble()

demo_profile() fetches the public profile maintained by the Delta Sharing project. The same client and table methods work with a private share.

Connect to your share

Pass the path to a .share profile:

client <- sharing_client("~/config.share")

sharing_client() also accepts a parsed profile list. See ?sharing_client for the supported profile fields and authentication methods.

Find a table

Use the client to discover the data available through its profile:

client$list_shares()
client$list_schemas("sales")
client$list_tables("sales", "default")

These methods follow pagination automatically and return printable lists of records. Create a reusable table handle from a listed table:

orders <- client$table("sales.default.orders")

orders
orders$version()
orders$schema()

Creating a table handle does not read table rows. Its protocol() and metadata() methods expose additional table details.

Read a snapshot

Configure the read with snapshot(), then materialize it. to_tibble() is the usual choice for R analysis:

snapshot <- orders$snapshot(
  columns = c("order_id", "status", "amount"),
  limit = 1000
)

orders_tbl <- snapshot$to_tibble()

Snapshots can also target a table version or point in time:

orders$snapshot(version = 42)$to_tibble()
orders$snapshot(timestamp = "2024-01-01T00:00:00Z")$to_tibble()

Choose the materializer that matches the next consumer:

Result Method
Tibble for ordinary R analysis to_tibble()
Base data frame to_data_frame()
In-memory Arrow table to_arrow()
Lazy Arrow reader, including for DuckDB to_arrow_reader()
Low-level Arrow C Stream to_arrow_stream()

Arrow is a required dependency. The tibble and data-frame methods automatically convert BIGINT columns to bit64::integer64, including empty and nested results. A valid -9223372036854775808 raises a conversion error because bit64 reserves it for NA; Arrow materializers retain that value. Each materializer call performs a new read, so reuse the result when the same data is needed more than once.

columns selects the returned columns and limit caps the number of returned rows. predicate accepts nested R lists that are sent to the sharing server as best-effort hints; predicates are not exact row filters. See ?SharingTable for the complete snapshot options and predicate structure.

Read changes

For a table with change data feed enabled, changes() reads an inclusive version or timestamp range:

changes_tbl <- orders$changes(
  starting_version = 120,
  ending_version = 125,
  columns = c(
    "order_id",
    "status",
    "_change_type",
    "_commit_version",
    "_commit_timestamp"
  )
)$to_tibble()

_change_type distinguishes inserts, deletes, and the before and after rows of updates. The commit columns identify when each change was recorded. Change reads expose the same materializers as snapshots.

Downloads and caching

Selected files are downloaded concurrently and cached for the R session. New handles for the same table reuse files already present in the cache, so overlapping reads may avoid downloading them again.

The cache is stored under R’s session temporary directory and normally needs no user management.

See vignette("performance-caching") for cache identity and lifetime, cold and repeated reads, download concurrency, materializer costs, and Arrow batching.

Next steps

See ?SharingTable, ?SharingSnapshot, and ?SharingChanges for the complete read options. The README shows how to query a shared table with DuckDB without first creating an R data frame.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.