The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

Getting Started with staRburst

Introduction

staRburst makes it trivial to scale your parallel R code from your laptop to 100+ AWS workers. This vignette walks through setup and common usage patterns.

Which API should I use?

staRburst offers a few entry points. They share the same engine — pick the one that fits how your code is written:

Need Use
Map a function over inputs (simplest) starburst_map()
Existing future / furrr code plan(starburst) then future_map()
Long-running job that survives disconnects starburst_session()
Explicit, reusable worker cluster starburst_cluster()
Configure account/backend defaults starburst_config()

If you’re new, start with starburst_map() — the next section walks through a first job with it. If you already have future/furrr code, skip to Migrating existing future / furrr code; the two styles are equivalent and shown side by side there.

Which backend? EC2 (default) vs Fargate

Every API above runs on one of two compute backends, chosen with launch_type:

  • EC2 — the default and recommended choice. Faster (no cold start), cheaper with Spot (the default), and more flexible (any instance type, warm pools). starburst_setup() provisions its capacity provider for you, so it works out of the box.
  • Fargate — an optional serverless alternative (launch_type = "FARGATE"). Fully managed and needs no starburst_setup_ec2 / capacity provider, but has cold-start latency and is bounded by your account’s Fargate vCPU quota. Reach for it if you specifically want task-based serverless execution.

You don’t have to choose up front — leave the default (EC2) and add launch_type = "FARGATE" later if you want to compare.

How it works (the mental model)

        Local R session
              |
              |  (1) serialize your function + inputs + detected globals (qs2)
              |      AWS credentials are read HERE, from your local environment
              v
   S3 bucket  +  staRburst control plane (ECS/ECR)
              |
              |  (2) launch workers: EC2 (default, Spot) or Fargate
              |      each worker's R packages come from your renv.lock, baked
              |      into a Docker image cached in ECR
              v
   Remote R workers  (one task per input element)
              |
              |  (3) each worker pulls a task from S3, runs it, writes the
              |      result + logs back to S3
              v
   Local R session  <-- results collected in order (starburst_map / cluster)
        or
   Detached-session store in S3  <-- collected later via starburst_session_attach()

What to know from this picture:

Installation

# Install from GitHub
remotes::install_github("scttfrdmn/starburst")

One-Time Setup

Before using staRburst, you need to configure AWS resources. This only needs to be done once.

library(starburst)

# Interactive setup wizard
starburst_setup()

This will: - Validate your AWS credentials - Create an S3 bucket for data transfer - Create an ECR repository for Docker images - Set up an ECS cluster and VPC resources - Provision the default EC2 capacity (launch template + Auto Scaling Group + capacity provider for c7g.xlarge) so the default EC2 backend works immediately — created at zero instances, so no compute cost until you run a job (skip with setup_ec2 = FALSE if you only use Fargate) - Check Fargate quotas and offer to request increases - Build the initial worker image

Provisioning the AWS resources takes about 2 minutes. On the first run, staRburst also builds the worker Docker image, which adds 5–10 minutes (it is cached and reused afterwards, and you can skip it with starburst_setup(build_image = FALSE) and build lazily on first job launch).

Your first job

The simplest way to use staRburst is starburst_map() — hand it your inputs and a function, and it runs the function over each input on AWS workers:

library(starburst)

# Define your work
expensive_simulation <- function(i) {
  # Some computation that takes a few minutes
  results <- replicate(1000, {
    x <- rnorm(10000)
    mean(x^2)
  })
  mean(results)
}

# Run across 50 cloud workers (EC2 by default)
results <- starburst_map(1:100, expensive_simulation, workers = 50)
#> [Starting] Starting starburst cluster with 50 workers
#> [Status] Processing 100 items with 50 workers
#> [Starting] Submitting 100 tasks...
#> [Wait] Progress: 100/100 (128.0s)
#> [OK] Completed in 128.0 seconds
#> [Cost] Estimated cost: $0.85

That is the whole workflow: starburst_setup() once, then starburst_map() per job. Everything below builds on this.

Migrating existing future / furrr code

If you already use future or furrr, you don’t need starburst_map() — set the plan to starburst and your existing future_map() / future_lapply() calls run on AWS unchanged. The two styles are equivalent; use whichever matches your code:

library(furrr)
library(starburst)

# Local baseline
plan(sequential)
results_local <- future_map(1:100, expensive_simulation)

# Same code, now on 50 cloud workers — just change the plan
plan(starburst, workers = 50)
results_cloud <- future_map(1:100, expensive_simulation)

# Results are identical to the local run
identical(results_local, results_cloud)
#> [1] TRUE

The two calls below do the same thing — pick the one that fits your codebase:

# Direct API
results <- starburst_map(1:100, expensive_simulation, workers = 50)

# future / furrr
plan(starburst, workers = 50)
results <- future_map(1:100, expensive_simulation)

Real-world examples

Rather than repeat toy snippets here, the article gallery carries complete, measured end-to-end examples — each with real runtimes, costs, and the batching choices that make (or break) a workload:

Before you size a job, read the two guides that report real numbers on when the cloud wins, when it loses, and how to pick task/worker counts: Workload Shapes and Performance. The short version: each task should be real work (seconds or more), and thousands of tiny tasks should be batched into dozens–hundreds — see Batch Small Tasks below.

Working with Data

Base R’s read.csv()/readRDS() cannot read s3:// URLs — you need an S3 client on the worker. A small helper keeps the examples below concrete and copy-pasteable (the worker image already includes paws.storage):

# Download an s3://bucket/key object to a temp file and read it on the worker.
read_s3 <- function(s3_uri, reader = readRDS) {
  m <- regmatches(s3_uri, regexec("^s3://([^/]+)/(.+)$", s3_uri))[[1]]
  bucket <- m[2]; key <- m[3]
  tmp <- tempfile()
  obj <- paws.storage::s3()$get_object(Bucket = bucket, Key = key)
  writeBin(obj$Body, tmp)
  reader(tmp)
}

Data Already in S3

If your data is already in S3, workers can read it directly:

plan(starburst, workers = 50)

results <- future_map(file_list, function(file) {
  # Each worker pulls its file from S3 (read_s3 defined above)
  data <- read_s3(sprintf("s3://my-bucket/%s", file), reader = read.csv)
  process(data)
})

Uploading Local Data

For smaller datasets, you can pass data as arguments:

# Load data locally
data <- read.csv("local_file.csv")

# staRburst automatically uploads to S3 and distributes
plan(starburst, workers = 50)

# `replicates` is your vector of inputs (e.g. bootstrap replicate ids)
results <- future_map(replicates, function(i) {
  # Each worker gets a copy of 'data'
  bootstrap_analysis(data, i)
})

Large Data Optimization

For very large objects, upload once to S3 yourself and have each worker read it from there, rather than serializing the object into every task:

# Upload once from your machine
paws.storage::s3()$put_object(
  Bucket = "my-bucket", Key = "large_data.rds",
  Body = "huge_file.rds"
)
s3_path <- "s3://my-bucket/large_data.rds"

# Workers read from S3 inside the task (read_s3 helper defined above)
plan(starburst, workers = 100)

# `tasks` is your vector of work items
results <- future_map(tasks, function(i) {
  data <- read_s3(s3_path)   # readRDS by default
  process(data, i)
})

Cost Management

Estimate Costs

# Check cost before running
plan(starburst, workers = 100, cpu = 4, memory = "8GB")
#> Estimated cost: ~$3.50/hour

Set Cost Limits

# Set an hourly cost ceiling (USD/hour) — jobs above this rate won't start
starburst_config(
  max_hourly_cost = 10,       # Don't start jobs estimated over $10/hour
  cost_alert_threshold = 5     # Warn at $5/hour
)

# Now jobs exceeding limit will error before starting
plan(starburst, workers = 1000)  # Would cost ~$35/hour
#> Error: Estimated cost ($35/hr) exceeds limit ($10/hr)

Cost estimates

After a run, staRburst reports an estimated cost — the measured worker runtime multiplied by the current AWS price for your instance type (On-Demand or Spot), looked up live from the AWS Pricing API (and cached; it falls back to built-in rates offline). It is a close estimate, not a figure from AWS billing / Cost Explorer.

plan(starburst, workers = 50)

results <- future_map(data, process)

#> Cluster runtime: ~23 minutes
#> Estimated cost: ~$1.34

Quota Management (Fargate backend)

This section applies to the Fargate backend. On the default EC2 backend, worker count is bounded by your normal EC2 On-Demand / Spot instance limits, not the Fargate vCPU quota below. If you stay on EC2 you can usually skip this.

Check Your Quota

starburst_quota_status()
#> Fargate vCPU Quota: 100 / 100 used
#> Allows: ~25 workers with 4 vCPUs each
#>
#> Recommended: Request increase to 500 vCPUs

Request Quota Increase

starburst_request_quota_increase(vcpus = 500)
#> Requesting Fargate vCPU quota increase:
#>   Current: 100 vCPUs
#>   Requested: 500 vCPUs
#>
#> [OK] Quota increase requested (Case ID: 12345678)
#> [OK] AWS typically approves within 1-24 hours

Wave-Based Execution

If you request more workers than your quota allows, staRburst automatically uses wave-based execution:

# Quota allows 25 workers, but you request 100
plan(starburst, workers = 100, cpu = 4)

#> [!] Requested: 100 workers (400 vCPUs)
#> [!] Current quota: 100 vCPUs (allows 25 workers max)
#>
#> [Plan] Execution plan:
#>   - Running in 4 waves of 25 workers each
#>
#> [TIP] Request quota increase to 500 vCPUs? [y/n]: y
#>
#> [OK] Quota increase requested
#> [Starting] Starting wave 1 (25 workers)...

results <- future_map(inputs, expensive_function)

#> [Wave] Wave 1: 100% complete (250 tasks)
#> [Wave] Wave 2: 100% complete (500 tasks)
#> [Wave] Wave 3: 100% complete (750 tasks)
#> [Wave] Wave 4: 100% complete (1000 tasks)

Troubleshooting

View Worker Logs

# View logs from most recent cluster
starburst_logs()

# View logs from specific task
starburst_logs(task_id = "abc-123")

# View last 100 log lines
starburst_logs(last_n = 100)

Check Cluster Status

starburst_status()
#> Active Clusters:
#>   • starburst-xyz123: 50 workers running
#>   • starburst-abc456: 25 workers running

Common Issues

Environment mismatch: Packages not found on workers

# Rebuild environment
starburst_rebuild_environment()

Task failures: Some tasks failing

# Check logs
starburst_logs(task_id = "failed-task-id")

# Often due to memory limits - increase worker memory
plan(starburst, workers = 50, memory = "16GB")  # Default is 8GB

Slow data transfer: Large objects taking too long

# Use Arrow for data frames
library(arrow)
write_parquet(my_data, "s3://bucket/data.parquet")

# Workers read Arrow
results <- future_map(1:100, function(i) {
  data <- read_parquet("s3://bucket/data.parquet")
  process(data, i)
})

Best Practices

1. Use for Right-Sized Workloads

Good: Each task takes >5 minutes

# 100 tasks, each takes 10 minutes
# Local: 1000 minutes, Cloud: ~10 minutes

Bad: Each task takes <1 minute

# 10000 tasks, each takes 30 seconds
# Startup overhead (45s) dominates

2. Batch Small Tasks

Instead of:

# 10,000 tiny tasks
results <- future_map(1:10000, small_function)

Do:

# 100 batches of 100 tasks each
batches <- split(1:10000, ceiling(seq_along(1:10000) / 100))

results <- future_map(batches, function(batch) {
  lapply(batch, small_function)
})

# Flatten results
results <- unlist(results, recursive = FALSE)

3. Use S3 for Large Data

Don’t:

big_data <- read.csv("10GB_file.csv")  # Upload for every task
results <- future_map(1:1000, function(i) process(big_data, i))

Do:

# Upload once to S3
tmp <- tempfile(fileext = ".csv"); write.csv(big_data, tmp, row.names = FALSE)
paws.storage::s3()$put_object(Bucket = "bucket", Key = "big_data.csv", Body = tmp)

# Workers read from S3 (read_s3 helper from "Working with Data" above)
results <- future_map(tasks, function(i) {
  data <- read_s3("s3://bucket/big_data.csv", reader = read.csv)
  process(data, i)
})

4. Set Reasonable Limits

starburst_config(
  max_hourly_cost = 50,            # Cap the hourly rate (USD/hour)
  cost_alert_threshold = 25        # Get warned early
)

5. Clean Up

# staRburst auto-cleans, but you can force it
plan(sequential)  # Switch back to local
# Old cluster resources are cleaned up automatically

Advanced: Custom Configuration

CPU and Memory

# High CPU, low memory (CPU-bound work)
plan(starburst, workers = 50, cpu = 8, memory = "16GB")

# Low CPU, high memory (memory-bound work)
plan(starburst, workers = 25, cpu = 4, memory = "32GB")

Timeout

# Increase timeout for long-running tasks (default 1 hour)
plan(starburst, workers = 10, timeout = 7200)  # 2 hours

Region

# Use specific region (default from config)
plan(starburst, workers = 50, region = "us-west-2")

Next Steps

Getting Help

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.