---
title: "Updating code written for versions 0.1 and 0.2"
output:
  rmarkdown::html_vignette:
    toc: true
    toc_depth: 2
    math_method: mathml
vignette: >
  %\VignetteIndexEntry{Updating code written for versions 0.1 and 0.2}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r}
#| label: knitr-options
#| include: false

knitr::opts_chunk$set(
  collapse = FALSE,
  comment = "#>",
  message = FALSE,
  fig.width = 7,
  fig.height = 5
)
```

Scripts written for catchmentACS 0.1 or 0.2 can stop with an error under
version 0.6.0, or run and give different results.
`vignette("porting-v03-to-v04", package = "catchmentACS")` and
`vignette("porting-v04-to-v05", package = "catchmentACS")` describe the changes
of versions 0.4 and 0.5 in more detail.

Several errors about drive-time areas (isochrones), such as a missing column or
the wrong coordinate reference system, refer to this article. The section on
saved drive-time areas explains them.

```{r}
#| label: setup
#| eval: true

library(catchmentACS)
library(sf)  # needed to subset the bundled sf objects with [
```

```{r}
#| label: setup-cache
#| include: false

# Compute every result in this article instead of reading saved ones; the
# option is restored at the end of the article.
old_options <- options(catchmentACS.cache_enabled = FALSE)
```

Two files that ship with the package stand in for data of your own.
`fixture_032_3site_fresh.rds` holds the inputs and results of a run with
version 0.2 for three points at public places in Birmingham, Mobile, and
Huntsville, Alabama: 10-minute areas from the public OSRM (Open Source Routing
Machine) demo server and American Community Survey (ACS) estimates for
2019–2023. The second file, `legacy_2025_isochrones.rds`, holds circles around
20 made-up sites, used here as drive-time areas that pass the current checks.

```{r}
#| label: read-data
#| eval: true

# Inputs and results of a run with version 0.2
old_run <- readRDS(system.file(
  "testdata", "fixture_032_3site_fresh.rds",
  package = "catchmentACS", mustWork = TRUE
))

# Drive-time areas saved with version 0.2
old_areas <- old_run$iso_sf_v2

# Drive-time areas that pass the current checks
example_areas <- readRDS(system.file(
  "extdata", "legacy_2025_isochrones.rds",
  package = "catchmentACS", mustWork = TRUE
))
```

## Summary of changes

::: {style="overflow-x: auto;"}

| Version | Change | Effect on an older script | What to do |
|:--|:--|:--|:--|
| 0.3 | Drive-time areas need a `ring_topology` column equal to `"cumulative"` | Saved areas give an error | Add the column, or join bands with `cacs_rings_to_cumulative()` |
| 0.3 | `cacs_isochrone()` joins the bands that the osrm package returns for several drive times into cumulative areas | Saved areas built with OSRM for several drive times may be bands, and the estimates for their longer drive times then cover only a band | Check them as below; build them again, or join them with `cacs_rings_to_cumulative()` after adding `isomin` and `isomax` |
| 0.3, 0.4 | Without `res`, OSRM areas use `30L` with `osrm_mode = "demo"` (the default) and `70L` with `osrm_mode = "docker"` (50 in version 0.2) | New areas differ from the old ones | Set `res = 50L` for the grid of version 0.2 (`iso_args = list(res = 50L)` in `cacs_run()`) |
| 0.3 | `ssi_rate` was `B19056_001 / B11001_001` and is now `B19056_002 / B19056_001` | Old values cannot be compared, and ACS data downloaded with the default variables of version 0.2 give `NA` | Download the ACS data again |
| 0.5 | The weights that average medians and per-person values sum to one for each variable | Old values from a call with more than one variable are too small: when every tract has all of them, an old value is the correct one divided by the number of variables | Compute them again |
| 0.4 | The rows of `rates_breakdown` in `summary()` are no longer shifted | Old means, standard deviations, and counts of missing values belong to other rates | Compute them again |
| 0.4 | `tibble::as_tibble()`, and functions that call it, put the rows of the five rates first within each site and drive time | In versions 0.4 to 0.5.1, rows picked by position after `tibble::as_tibble()` differed, and `dplyr::semi_join()`, `dplyr::anti_join()`, and summaries after `dplyr::rowwise()` could give wrong results without a warning | Nothing from version 0.6.0 on, which keeps the order again. `rate_first = TRUE`, or `options(catchmentACS.rate_first_default = TRUE)` for calls from your own code, puts the rate rows first |
| 0.3 | A result saved in the cache needs a checksum file | Results saved by versions 0.1 and 0.2 are not used | Nothing; the old cache folder can be deleted by hand (see `?cacs_cache_dir`) |
| 0.2 | `cacs_acs_prefetch()` leaves out tracts numbered 9900 or higher (water and special-purpose tracts) and tracts with no area | Fewer tracts, with a message | `drop_water_tracts = FALSE` keeps them in a new download; data saved in the cache without them need `force_refresh = TRUE` |
| 0.2, 0.3 | Results gain the columns `n_tracts_num`, `n_tracts_den`, and `ring_topology` | Columns move | Pick columns by name |

:::

## Drive-time areas saved with versions 0.1 and 0.2

`cacs_intersect_weight()` checks the drive-time areas it is given, and so does
`cacs_run()` with `precomputed_isochrones`, after it downloads the ACS
estimates if `acs` is not supplied. The areas must be an `sf` object in
EPSG:4326 (longitude and latitude) with the 16 columns of a `cacs_isochrone()`
result, including `ring_topology`, which versions 0.1 and 0.2 did not have.
`cacs_validate_iso()` makes the same checks of the columns, their types, and
the coordinate reference system, and returns one row for each problem, with a
line of R code for a fix in its `example` column. Checking saved areas with it
first avoids a download that ends in an error:

```{r}
#| label: s01-saved-areas
#| eval: true

issues <- cacs_validate_iso(old_areas)
issues[, c("check", "col", "actual")]
writeLines(issues$example)
```

This fix is right only if each row holds the whole area reachable within its
`drive_time_min`, so that each area contains the shorter ones of the same site.
`cacs_validate_iso()` flags rows whose `isomin` is above 0, but it does not
compare the areas. These areas have one drive time for each site, so there is
nothing to compare.

Drive-time bands, also called rings, cover only the minutes between two drive
times, so a band shares no area with the band inside it. When several drive
times are requested in one call, the osrm package returns bands, and since
version 0.3 `cacs_isochrone()` joins them into cumulative areas. Saved areas
that versions 0.1 and 0.2 built with OSRM for several drive times may be bands.
The share of a shorter area that lies inside the longer one tells bands from
cumulative areas, as in this example of one site with two bands:

```{r}
#| label: s02-bands-or-cumulative
#| eval: true

make_square <- function(xmin, ymin, xmax, ymax) {
  sf::st_polygon(list(rbind(
    c(xmin, ymin), c(xmax, ymin), c(xmax, ymax),
    c(xmin, ymax), c(xmin, ymin)
  )))
}
core <- make_square(-87.10, 33.00, -87.00, 33.10)

# Two bands: 0 to 5 minutes, and 5 to 10 minutes (a larger square
# with the first one cut out)
bands <- sf::st_sf(
  site_id = "S01",
  isomin = c(0L, 5L),
  isomax = c(5L, 10L),
  geometry = sf::st_sfc(
    core,
    sf::st_difference(make_square(-87.20, 32.90, -86.90, 33.20), core),
    crs = 4326
  )
)

# Share of the shorter area that lies inside the longer one
inside_share <- function(shorter, longer) {
  overlap <- sf::st_intersection(sf::st_geometry(longer),
                                 sf::st_geometry(shorter))
  sum(as.numeric(sf::st_area(overlap))) /
    sum(as.numeric(sf::st_area(shorter)))
}
inside_share(shorter = bands[1, ], longer = bands[2, ])
```

`cacs_rings_to_cumulative()` joins each band with the bands inside it, using the
columns `isomin` and `isomax` (the start and end of each band, in minutes):

```{r}
#| label: bands-to-cumulative
#| eval: true

cumulative <- cacs_rings_to_cumulative(bands)
sf::st_drop_geometry(cumulative)
inside_share(shorter = cumulative[1, ], longer = cumulative[2, ])
```

The result still lacks the other columns of a `cacs_isochrone()` result. For
areas made by another tool, these columns can hold the values that
`cacs_isochrone()` gives an area built without a problem, with the meanings
given in the Value section of `?cacs_isochrone`. `provider` and
`provider_requested` must each be one of the four routing service names,
`"osrm"`, `"ors"`, `"mapbox"`, or `"r5r"`, and `provider` is copied into the
result, so for such areas it is only a label:

```{r}
#| label: s03-fill-columns
#| eval: true

areas <- cumulative
areas$provider <- "osrm"
areas$provider_requested <- "osrm"
areas$provider_downgrade <- FALSE
areas$profile <- "car"
areas$routing_engine_version <- "osrm-pkg/unknown"
areas$polygon_simplification_tolerance <- NA_real_
areas$osm_snapshot_date <- "unknown"
areas$osm_snapshot_status <- "unknown_best_effort"
areas$generated_at <- Sys.time()
areas$isochrone_empty <- FALSE
areas$failure_reason <- NA_character_
areas$retry_count <- 1L

nrow(cacs_validate_iso(areas))
```

For example, `cacs_run()` uses an area only when its `isochrone_empty` is
`FALSE` and its `failure_reason` is `NA`. An `isochrone_empty` of `NA` or a
`failure_reason` of `""` passes the checks, but the area then counts as a
routing failure.

`cacs_intersect_weight()` takes areas only in EPSG:4326, and
`cacs_validate_iso()` reports areas in another system:

```{r}
#| label: s04-crs-check
#| eval: true

bad_crs <- sf::st_transform(example_areas[1:2, ], 4269)
cacs_validate_iso(bad_crs) |>
  dplyr::filter(check == "iso_crs") |>
  dplyr::select(check, actual, expected, example)
```

`isochrone_empty` and `provider_downgrade` must be logical, and
`drive_time_min` and `retry_count` integers:

```{r}
#| label: s05-provider-downgrade-type
#| eval: true

bad_type <- example_areas[1, ]
bad_type$provider_downgrade <- NA_character_
issues <- cacs_validate_iso(bad_type)
issues[, c("col", "actual", "expected")]
writeLines(issues$example)
```

## Running an older analysis again

The run saved in `fixture_032_3site_fresh.rds` shows what changes when an
analysis from version 0.2 is run again on its saved inputs. With the areas and
the ACS estimates supplied, `cacs_run()` builds no areas and downloads nothing:

```{r}
#| label: s06-rerun-saved-inputs
#| eval: true

old_sites <- sf::st_drop_geometry(old_run$sites_sf)
old_sites <- old_sites[, c("site_id", "lon", "lat")]
old_run_areas <- old_run$iso_sf_v2
old_run_areas$ring_topology <- "cumulative"

rerun <- cacs_run(
  old_sites,
  state = "AL",
  drive_times = 10,
  precomputed_isochrones = old_run_areas,
  acs = old_run$acs_sf,
  weight_args = list(keep_tract_audit = TRUE),
  verbose = FALSE
)
```

The warning gives one line for each reason found among the rate rows, and here
there is one: the rows whose `failure_origin` is `"carrier"`, rates that are
`NA` because the estimate for the numerator or the denominator, or its margin
of error (the half-width of its 90 percent confidence interval), is missing.
Here they are the `ssi_rate` rows, one for each site. The new results compare
with the saved ones as follows:

```{r}
#| label: s07-compare-with-saved
#| eval: true

compare <- dplyr::inner_join(
  tibble::as_tibble(old_run$result_run_v2) |>
    dplyr::select(site_id, variable, saved = estimate),
  tibble::as_tibble(rerun) |>
    dplyr::select(site_id, variable, new = estimate),
  by = c("site_id", "variable")
)
dplyr::filter(
  compare,
  site_id == "AL_BHM_01",
  variable %in% c("B01003_001", "B19013_001", "B19301_001",
                  "poverty_rate", "ssi_rate")
)

# All other counts and rates at the three sites
same <- dplyr::filter(
  compare,
  !variable %in% c("B19013_001", "B19301_001", "ssi_rate")
)
all.equal(same$saved, same$new)
```

The counts and four of the rates, `r nrow(same)` rows in all, did not change.
Median household income (`B19013_001`) and per capita income (`B19301_001`)
grew by the same factor wherever they have a value:

```{r}
#| label: median-ratio
#| eval: true

proxies <- dplyr::filter(compare, variable %in% c("B19013_001", "B19301_001"))
dplyr::mutate(proxies, ratio = new / saved)
```

The factor, `r length(unique(old_run$acs_sf$variable))`, is the number of ACS
variables in the run, in which every tract has a row for each variable. Before
version 0.5.0, the weights that average medians and per-person values summed to
one over all the variables of a site and drive time;
`vignette("porting-v04-to-v05", package = "catchmentACS")` describes the
correction.

In the saved results, `ssi_rate` is
`r unique(compare$saved[compare$variable == "ssi_rate"])` at every site, because
version 0.2 divided `B19056_001` by `B11001_001`, two counts of all households.
It now divides the households with Supplemental Security Income in the
past 12 months, `B19056_002`, by all households, `B19056_001`. The saved ACS
estimates lack `B19056_002`:

```{r}
#| label: ssi-codes
#| eval: true

cacs_acs_default_rates$ssi_rate
setdiff(unlist(cacs_acs_default_rates), old_run$acs_sf$variable)
```

A new download with the default variables of `cacs_acs_prefetch()` or
`cacs_run()` includes it, since it is in `cacs_acs_default_vars`.

Tables made with `summary()` in version 0.2 are also affected. Until version
0.4, each row of its `rates_breakdown` element showed the mean, standard
deviation, and count of missing values of the next rate in alphabetical order,
and the last row those of the first rate. Version 0.4 corrected this and added
per-site tables of the rates, which `cacs_summary_as_markdown()` formats as
Markdown.

Since version 0.3, `keep_tract_audit = TRUE`, passed above to
`cacs_intersect_weight()` through `weight_args`, adds the attribute
`cacs_tract_audit`. It has one row for each tract in each drive-time area, with
the tract's coverage weight, the share of its area inside the drive-time area.
It can replace code that intersects the areas with tract boundaries to list the
tracts:

```{r}
#| label: s08-tract-table
#| eval: true

tracts <- attr(rerun, "cacs_tract_audit")
names(tracts)
dplyr::count(tracts, site_id, drive_time_min, name = "tracts")
```

## Building drive-time areas again

Areas built again can differ from the saved ones. One reason is the OSRM grid
resolution `res`: the osrm package draws each area from travel times to a grid
of `res` by `res` points around the site. Without `res`, version 0.2 used 50;
version 0.6.0 uses `30L` with `osrm_mode = "demo"`, the default, and `70L` with
`osrm_mode = "docker"`. Version 0.2 dropped `res` from `iso_args` with a
warning, so a script that gave it there used 50 as well. To build the areas on
the grid of version 0.2, set `res = 50L` in `cacs_isochrone()`, or
`iso_args = list(res = 50L)` in `cacs_run()`. `?cacs_isochrone` describes the
grid, an option that changes the values, and the waiting time on the demo
server. The road data of the server and the coordinates of the sites can
change too.

## Saved results in the cache

Since version 0.3, a result saved in the cache is used only when the checksum
file next to it (`.fingerprint`) is present and its checksum matches. Results
saved by versions 0.1 and 0.2 have no checksum file, so the first run after the
upgrade builds the areas and downloads the ACS estimates again.
Versions 0.5.1 and earlier kept saved results in the user cache folder of the
operating system, which version 0.6.0 neither reads nor deletes;
`?cacs_cache_dir` gives the path of that folder on each system, and the folder
can be deleted by hand.

`cacs_set_cache(FALSE)` turns the cache off for the rest of the R session, so
that a script uses no saved results. The option `catchmentACS.cache_dir` or the
environment variable `CACS_CACHE_DIR` sets another cache folder (see
`cacs_cache_dir()`). Results saved in the default folder, inside the temporary
folder of the R session, are deleted when the session ends.

```{r}
#| label: restore-options
#| include: false

options(old_options)
rm(old_options)
```
