---
title: "Updating code written for version 0.4"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 2
    math_method: mathml
vignette: >
  %\VignetteIndexEntry{Updating code written for version 0.4}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r}
#| label: knitr-options
#| include: false

knitr::opts_chunk$set(
  collapse = FALSE,
  comment = "#>",
  message = FALSE,
  fig.width = 7,
  fig.height = 5
)
```

catchmentACS 0.5.0 corrects the estimates of medians and per-person values and
checks some arguments more strictly than 0.4 did. For code written for 0.3,
`vignette("porting-v03-to-v04", package = "catchmentACS")` describes the
changes made in 0.4.

```{r}
#| label: setup
#| eval: true

library(catchmentACS)
library(dplyr)
library(sf)  # needed to subset the bundled sf objects with [
```

```{r}
#| label: setup-cache
#| include: false

# Compute every result in this article instead of reading saved ones; the
# option is restored at the end of the article.
old_options <- options(catchmentACS.cache_enabled = FALSE)
```

## Medians and per-person values

`cacs_run()` averages the medians and per-person values of three tables of the
American Community Survey (ACS) over the census tracts that a drive-time area
overlaps: median household income (`B19013_001`), median home value
(`B25077_001`), and per capita income (`B19301_001`). A median or per-person
value from any other table is added up like a count. The weights of the
average, called area shares, are proportional to the area of each tract's
overlap with the drive-time area and sum to one for each site, drive time, and
variable. Before 0.5.0, they summed to one over all the variables of a site
and drive time taken together. In a call with more than one variable, the
estimates computed with these weights were therefore too small, and so were
their margins of error, the half-widths of their confidence intervals (90
percent by default). When every tract has a row for each variable, they were
smaller by a factor equal to the number of variables in the call. Counts and
rates are computed with other weights and did not change.

Medians and per-person values computed with 0.4 need to be computed again. In
a result, they are the rows whose `weight_basis` is `"area_mean"`. The example
below finds them in a run on made-up data installed with the package, in which
the drive-time areas are circles and the ACS estimates are random numbers:

```{r}
#| label: area-share-rows
#| eval: true

iso <- readRDS(system.file(
  "extdata", "legacy_2025_isochrones.rds", package = "catchmentACS"
))
acs <- readRDS(system.file(
  "extdata", "sample_alabama_subset.rds", package = "catchmentACS"
))
site_07 <- data.frame(site_id = "AL_SITE_07", lon = -85.365, lat = 31.655)

result <- cacs_run(
  site_07, state = "AL",
  precomputed_isochrones = iso[iso$site_id == "AL_SITE_07", ],
  acs = acs, verbose = FALSE
)
tibble::as_tibble(result) |>
  filter(weight_basis == "area_mean") |>
  select(drive_time_min, variable, estimate, moe)
```

## The `year` argument

`cacs_acs_prefetch()` accepts as `year` only a whole number from 2009 to 2024,
and so does `cacs_run()` when it downloads the ACS estimates. In 0.4, they
accepted any year from 2009 to the year before the current one and truncated a
fractional `year`, such as `2023.5`. In 0.5.0, a fraction and a year after
2024 give an error. When the estimates are supplied through `acs`,
`cacs_run()` only records `year`.

## Saved ACS estimates and the Census API key

`cacs_acs_prefetch()` reads estimates saved in the cache without
a Census API key, as it did in 0.4. In 0.5.0, it looks for them before
sending any request, where 0.4 first asked tidycensus for the list of ACS
variables and went on if that failed. It needs the key only to download: when
no saved result matches the call, or with `force_refresh = TRUE`.

A saved result matches only while the installed versions of R, catchmentACS,
tidycensus, tigris, and sf stay the same (`?cacs_acs_prefetch`), so copies
saved with 0.4 are not read. The first call with 0.5.0 downloads the
estimates again and needs the key. Results that `cacs_intersect_weight()`
saved in the cache with 0.4 are not reused either; they are computed again.
From version 0.6.0 on, saved results are kept only until the R session ends
unless a cache folder that lasts between sessions is set, and the folder used
by versions 0.5.1 and earlier is not read (see `?cacs_cache_dir`).

## Sites as a plain data frame

`cacs_isochrone()` now also accepts the sites as a base R data frame with the
columns `site_id`, `lon`, and `lat`, like `site_07` above. In 0.4, it required
an sf object of points or a tibble, and both are still accepted.

## The `isochrone` column when `cacs_run()` builds the areas

With `output = "list_column"` or `output = "both"`, the list-column form of
`cacs_run()` has a column `isochrone`, which holds the drive-time area of each
row as a one-row sf object. In 0.4, it held the areas only when they were
supplied through `precomputed_isochrones`, and `NULL` when `cacs_run()` built
them. In 0.5.0 it holds them in both cases. A message about the column is
shown on every such call, even with `verbose = FALSE`.

## Choices that are not implemented

`weight_method = "population"` is not implemented yet; area weighting
(`"area"`, the default) is the only method. In 0.5.0, `cacs_run()` gives the
error before any step runs, with the class `catchmentACS_error_credential`. In
0.4, the error came at the intersection step, after the areas and the ACS
estimates were ready, when `bg_pop_sf` was supplied; without `bg_pop_sf`, it
came at the start, with the class `catchmentACS_error_schema`.

Choosing `provider = "mapbox"` or `provider = "r5r"` to build the areas, or a
`rates` list other than `cacs_acs_default_rates`, the list of the five built-in
rates, gives an error, as it did in 0.4.

```{r}
#| label: restore-options
#| include: false

options(old_options)
rm(old_options)
```
