The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
snap_rate with
the ratio formulalabor_force_participation
with the proportion formulaFor each drive-time area, cacs_derive_rates() computes
five rates from American Community Survey (ACS) estimates for census
tracts, such as the share of people below the poverty level. Each rate
comes with a margin of error (MOE), the half-width of its confidence
interval, at the 90 percent level unless another level was set with the
level argument of cacs_propagate_moe(), whose
result cacs_derive_rates() takes. A level
given to cacs_derive_rates() itself gives an error. The
margin of error is computed with one of two formulas, chosen for each
rate according to the argument formula_dispatch, and the
last section of this article checks two rates against hand
calculations.
Each rate divides one ACS count by another, such as people below the
poverty level by people for whom poverty status is determined. We use
the notation of
vignette("methodology", package = "catchmentACS"). For
tract
,
and
are the estimates of the numerator and the denominator, and
is the coverage weight, the share of the tract’s area inside the
drive-time area of site
.
The rate for the site is
The numerator and the denominator are counts and are combined like
any other count, which assumes that whatever each of the two variables
counts is spread evenly over the tract’s area (Comber and Zeng 2019, 8). The coverage weights
are explained in
vignette("theory-spatial-aggregation", package = "catchmentACS").
The package adds up the numerator and the denominator over the tracts first and divides once; it does not compute a rate for each tract. When the numerator and the denominator cover the same tracts and every is positive, the result is an average of the tract rates with weights , the part of each tract’s denominator counted in the drive-time area:
An average of the tract rates weighted by coverage weight or by overlap area alone would ignore how large each tract’s denominator is. Because both counts of a tract get the same coverage weight, an area that overlaps only one tract has that tract’s rate.
The rates are computed from the weighted counts alone, without the geometry.
The numerator and denominator codes of the five rates are listed in
cacs_acs_default_rates:
#> $poverty_rate
#> num den
#> "B17001_002" "B17001_001"
#>
#> $snap_rate
#> num den
#> "B22003_002" "B22003_001"
#>
#> $ssi_rate
#> num den
#> "B19056_002" "B19056_001"
#>
#> $unemp_rate
#> num den
#> "B23025_005" "B23025_003"
#>
#> $labor_force_participation
#> num den
#> "B23025_002" "B23025_001"
| Rate | Numerator | Denominator |
|---|---|---|
poverty_rate |
people whose income in the past 12 months was below the poverty
level (B17001_002) |
people for whom poverty status is determined
(B17001_001) |
snap_rate |
households that received Food Stamps or the Supplemental Nutrition
Assistance Program (SNAP) in the past 12 months
(B22003_002) |
all households (B22003_001) |
ssi_rate |
households with Supplemental Security Income (SSI) in the past 12
months (B19056_002) |
all households (B19056_001) |
unemp_rate |
unemployed people in the civilian labor force
(B23025_005) |
the civilian labor force, among people 16 years and over
(B23025_003) |
labor_force_participation |
people in the labor force, including the armed forces
(B23025_002) |
people 16 years and over (B23025_001) |
In each rate the numerator is part of the denominator: the people or
households it counts are among those that the denominator counts.
unemp_rate and labor_force_participation come
from the same ACS table, B23025, but divide by different
counts: the civilian labor force for the unemployment rate, and everyone
16 years and over for labor force participation.
The two counts of a rate are summed over different tracts when a
tract in the area has a row for one of the two codes and not for the
other. The package keeps rows whose estimate is missing, and a rate that
needs such a tract is NA rather than wrong. The two counts
cover different tracts when those rows are dropped before
cacs_intersect_weight() runs, for example by a filter of
your own. The numerator is then no longer part of the denominator, and
the rate can be much larger or smaller than the share for the area, even
above 1. No warning is given. The columns n_tracts_num and
n_tracts_den of the rate row, described in the help page of
cacs_derive_rates(), give the two numbers of tracts.
These five are the only rates that cacs_derive_rates()
computes. Its rates argument accepts only
cacs_acs_default_rates; any other list, including a subset
or a reordering of it, gives an error. Other rates, such as a poverty
rate for children built from the age and sex groups of table
B17001, can be computed from the results of
cacs_propagate_moe() or cacs_run(). These have
a row with the weighted count and its margin of error for each ACS count
requested. A numerator that adds up several of these counts gets its
margin of error from the formula that the Census Bureau’s handbook gives
for a sum, which again treats the counts as independent (U.S. Census Bureau 2020, chap. 8). The formulas
below then give the margin of error of the rate.
The Census Bureau’s handbook has a formula for a proportion (its formula 6) and one for a ratio (formula 7). A proportion is a ratio whose numerator is part of its denominator; the ratio formula is for a ratio whose numerator is not. Both compute the margin of error of from the margins of error and of the two weighted counts (U.S. Census Bureau 2020, chap. 8):
The package’s warnings and the help page of
cacs_propagate_moe() call them C1 and C2, and the columns
moe_formula_requested and
moe_formula_effective record them as
"proportion_subset" and
"general_ratio_conservative". The ratio formula divides by
the absolute value of the denominator, which the hand calculations below
use as well; for the five rates, whose denominators are counts, that is
the denominator itself.
The two values under the square root differ by
,
so the ratio formula never gives the narrower margin of error. How much
wider it is varies from rate to rate and from one area to another, and
vignette("theory-moe-propagation", package = "catchmentACS")
gives the factor. Both formulas start from margins of error of the
weighted counts that treat the tract estimates as independent, an
assumption discussed in the same article.
By default, cacs_derive_rates() and
cacs_run() use the ratio formula for all five rates,
although the numerator of each is part of its denominator. For these
proportions the handbook gives the proportion formula, so the default is
a choice of the package. The argument formula_dispatch of
either function takes one of three values:
| Rate | "general_ratio_conservative" (default) |
"auto" |
"proportion_subset" |
|---|---|---|---|
poverty_rate |
ratio | proportion | proportion |
labor_force_participation |
ratio | proportion | proportion |
snap_rate |
ratio | ratio | ratio |
ssi_rate |
ratio | ratio | ratio |
unemp_rate |
ratio | ratio | ratio |
"proportion_subset" gives the same formulas as
"auto" and, in addition, a warning naming the three rates
that keep the ratio formula. unemp_rate keeps the ratio
formula with every value of formula_dispatch, although its
numerator is part of its denominator, to reproduce the 2025 analysis the
package was first written for. The column
moe_formula_requested of each rate row records the formula
chosen for the rate. For snap_rate, ssi_rate,
and unemp_rate, the proportion formula can be applied by
hand to the estimate and moe of the rows of
the two ACS codes, unless the value under its square root is
negative.
If the proportion formula is chosen for a rate and the value under
its square root is negative, the package uses the ratio formula for that
row instead, as the handbook advises.
vignette("theory-moe-propagation", package = "catchmentACS")
describes the columns that record the substitution.
The checks below compare hand calculations with the output of
cacs_derive_rates() for two rates: snap_rate
with the ratio formula and labor_force_participation with
the proportion formula. For labor_force_participation, the
margin of error from the ratio formula is also computed by hand, for
comparison.
The example is the 15-minute area around site AL_SITE_03
in the data shipped with the package. Its drive-time area is a circle
and its tract estimates are made-up values, so the rates show the
arithmetic and are not estimates for any real area. The code runs only
the three steps of cacs_run() that come after the download
and the routing, and it uses formula_dispatch = "auto" so
that two of the rates use the proportion formula:
# Drive-time areas and made-up ACS data bundled with the package
iso <- readRDS(system.file(
"extdata", "legacy_2025_isochrones.rds", package = "catchmentACS"
))
acs <- readRDS(system.file(
"extdata", "sample_alabama_subset.rds", package = "catchmentACS"
))
site_id <- "AL_SITE_03"
drive_time <- 15L
iso_one <- iso[
iso$site_id == site_id & iso$drive_time_min == drive_time, ,
drop = FALSE
]
weighted <- cacs_intersect_weight(
iso_sf = iso_one,
acs_sf = acs,
weight_method = "area",
verbose = FALSE
)
propagated <- cacs_propagate_moe(weighted, verbose = FALSE)
# "auto": the proportion formula for poverty_rate and
# labor_force_participation, the ratio formula for the other three rates
rates <- suppressWarnings(
cacs_derive_rates(propagated, formula_dispatch = "auto", verbose = FALSE)
)suppressWarnings() hides the warning that
cacs_derive_rates() gives at the end, which counts the rows
where the ratio formula replaced the proportion formula.
The checks read
,
,
and their variances, the numbers that cacs_derive_rates()
itself uses, from the attribute cacs_aggregation_carriers
of the result of cacs_intersect_weight().
cacs_derive_rates() removes that attribute, so the checks
below need the result of the earlier step, not a rate table or a result
of cacs_run(), where it is NULL. The variances
(var_total_raw) are squared standard errors (SE), so a
margin of error is
,
with
at the 90 percent level:
z <- 1.645
totals <- attr(weighted, "cacs_aggregation_carriers")
# The weighted count (est_total) or its variance (var_total_raw) for one ACS code
total_of <- function(code, col) totals[[col]][totals$variable == code]
totals |>
filter(variable %in% c(
"B22003_002", "B22003_001", # snap_rate num / den
"B23025_002", "B23025_001", # labor_force_participation num / den
"B17001_002", "B17001_001" # poverty_rate num / den
)) |>
select(variable, est_total, var_total_raw, weight_sum, n_tracts)#> # A tibble: 6 × 5
#> variable est_total var_total_raw weight_sum n_tracts
#> <chr> <dbl> <dbl> <dbl> <int>
#> 1 B17001_001 8474 320617. 3 3
#> 2 B17001_002 2165 11820. 3 3
#> 3 B22003_001 3556 48137. 3 3
#> 4 B22003_002 891 2533. 3 3
#> 5 B23025_001 8136 208358. 3 3
#> 6 B23025_002 4359 73889. 3 3
In this example, weight_sum equals
n_tracts: the 3 tracts lie entirely inside the drive-time
area, so each weighted count is the plain sum of the tract
estimates.
snap_rate with the ratio formulasnap_rate uses the ratio formula with every value of
formula_dispatch. Written with the variances, the formula
is
:
A_snap <- total_of("B22003_002", "est_total")
B_snap <- total_of("B22003_001", "est_total")
VarA_snap <- total_of("B22003_002", "var_total_raw")
VarB_snap <- total_of("B22003_001", "var_total_raw")
R_snap <- A_snap / B_snap # the rate A / B
hand_moe_snap <- z * sqrt(VarA_snap + R_snap^2 * VarB_snap) / abs(B_snap)
pkg_snap <- rates |>
filter(variable == "snap_rate") |>
select(estimate, moe, moe_formula_effective, moe_fallback)
list(
A = A_snap, B = B_snap,
hand_estimate = R_snap,
hand_moe_C2 = hand_moe_snap,
pkg = pkg_snap
)#> $A
#> [1] 891
#>
#> $B
#> [1] 3556
#>
#> $hand_estimate
#> [1] 0.2505624
#>
#> $hand_moe_C2
#> [1] 0.0344782
#>
#> $pkg
#> # A tibble: 1 × 4
#> estimate moe moe_formula_effective moe_fallback
#> <dbl> <dbl> <chr> <lgl>
#> 1 0.251 0.0345 general_ratio_conservative FALSE
The package’s row records the ratio formula with no substitution, and its estimate, 0.2506, and margin of error, 0.0345, are the values computed by hand:
all.equal(
c(estimate = R_snap, moe = hand_moe_snap),
c(estimate = pkg_snap$estimate, moe = pkg_snap$moe)
)#> [1] TRUE
labor_force_participation with the proportion
formulaWith formula_dispatch = "auto",
labor_force_participation uses the proportion formula,
,
unless the value under the square root is negative. The code computes
that value, the margin of error from the proportion formula, and, for
comparison, the margin of error from the ratio formula:
A_lfp <- total_of("B23025_002", "est_total")
B_lfp <- total_of("B23025_001", "est_total")
VarA_lfp <- total_of("B23025_002", "var_total_raw")
VarB_lfp <- total_of("B23025_001", "var_total_raw")
R_lfp <- A_lfp / B_lfp
under_root <- VarA_lfp - R_lfp^2 * VarB_lfp # under the proportion formula's root
hand_moe_C1 <- z * sqrt(under_root) / B_lfp
hand_moe_C2 <- z * sqrt(VarA_lfp + R_lfp^2 * VarB_lfp) / abs(B_lfp) # for comparison
list(
under_root = under_root, # positive, so the proportion formula can be used
hand_moe_C1 = hand_moe_C1,
hand_moe_C2 = hand_moe_C2 # the ratio formula
)#> $under_root
#> [1] 14080.91
#>
#> $hand_moe_C1
#> [1] 0.02399222
#>
#> $hand_moe_C2
#> [1] 0.07392929
The value under the square root is positive, so the proportion formula applies. In this area, the proportion formula gives a margin of error of 0.0240, and the ratio formula gives 0.0739, 3.1 times as wide. The package’s row records the proportion formula with no substitution, and its margin of error is the value computed by hand:
pkg_lfp <- rates |>
filter(variable == "labor_force_participation") |>
select(estimate, moe, moe_formula_effective, moe_fallback, moe_fallback_reason)
pkg_lfp#> # A tibble: 1 × 5
#> estimate moe moe_formula_effective moe_fallback moe_fallback_reason
#> <dbl> <dbl> <chr> <lgl> <chr>
#> 1 0.536 0.0240 proportion_subset FALSE n/a
#> [1] TRUE
The rate rows of the example record, for each rate, the formula chosen and the formula used, and the tract counts:
rates |>
filter(estimand_family == "derived_rate") |>
select(variable, estimate, moe,
moe_formula_requested, moe_formula_effective,
moe_fallback, moe_fallback_reason,
n_tracts, n_tracts_num, n_tracts_den) |>
print(width = Inf)#> # A tibble: 5 × 10
#> variable estimate moe moe_formula_requested
#> <chr> <dbl> <dbl> <chr>
#> 1 poverty_rate 0.255 0.0351 proportion_subset
#> 2 snap_rate 0.251 0.0345 general_ratio_conservative
#> 3 ssi_rate 0.0852 0.0113 general_ratio_conservative
#> 4 unemp_rate 0.0664 0.00838 general_ratio_conservative
#> 5 labor_force_participation 0.536 0.0240 proportion_subset
#> moe_formula_effective moe_fallback moe_fallback_reason n_tracts
#> <chr> <lgl> <chr> <int>
#> 1 general_ratio_conservative TRUE negative_variance NA
#> 2 general_ratio_conservative FALSE n/a NA
#> 3 general_ratio_conservative FALSE n/a NA
#> 4 general_ratio_conservative FALSE n/a NA
#> 5 proportion_subset FALSE n/a NA
#> n_tracts_num n_tracts_den
#> <int> <int>
#> 1 3 3
#> 2 3 3
#> 3 3 3
#> 4 3 3
#> 5 3 3
moe_formula_requested follows the "auto"
column of the table in the formula section:
"proportion_subset" for poverty_rate and
labor_force_participation, and
"general_ratio_conservative" for the other three rates.
moe_formula_effective differs from it only for
poverty_rate. n_tracts is NA on
every rate row, and n_tracts_num equals
n_tracts_den because each tract of the area has a row for
every ACS code. poverty_rate records a substitution: in
this area the value under the square root of its proportion formula is
negative, so the ratio formula was used.
vignette("theory-moe-propagation", package = "catchmentACS")
works through such a case by hand.
A rate and its margin of error are NA when the numerator
or the denominator, or the margin of error of either, is missing for the
site and drive time. This happens when one of the two ACS codes has no
rows in the ACS data or when a tract in the area has a missing estimate
or margin of error for one of them. It also happens for every rate of a
pair with no tract left. The code below removes every row of the SSI
numerator, B19056_002, from the example data and runs the
three steps again, this time with the default value of
formula_dispatch:
acs_drop <- acs[sf::st_drop_geometry(acs)$variable != "B19056_002", ]
weighted_d <- cacs_intersect_weight(
iso_sf = iso_one, acs_sf = acs_drop,
weight_method = "area", verbose = FALSE
)
propagated_d <- cacs_propagate_moe(weighted_d, verbose = FALSE)
rates_d <- suppressWarnings(cacs_derive_rates(propagated_d, verbose = FALSE))
rates_d |>
filter(estimand_family == "derived_rate") |>
select(variable, estimate, moe,
moe_fallback_reason, failure_origin, n_tracts_num, n_tracts_den) |>
print(width = Inf)#> # A tibble: 5 × 7
#> variable estimate moe moe_fallback_reason failure_origin
#> <chr> <dbl> <dbl> <chr> <chr>
#> 1 poverty_rate 0.255 0.0351 n/a none
#> 2 snap_rate 0.251 0.0345 n/a none
#> 3 ssi_rate NA NA n/a carrier
#> 4 unemp_rate 0.0664 0.00838 n/a none
#> 5 labor_force_participation 0.536 0.0739 n/a none
#> n_tracts_num n_tracts_den
#> <int> <int>
#> 1 3 3
#> 2 3 3
#> 3 NA NA
#> 4 3 3
#> 5 3 3
Here suppressWarnings() hides the warning at the end,
which counts the rate rows that are NA for this reason.
ssi_rate has NA in estimate and
moe, and "carrier" in
failure_origin, the column that records the step at which a
row failed; "carrier" means that a count or margin of error
that the rate needs is missing. n_tracts_num and
n_tracts_den are NA, and
moe_fallback_reason is "n/a" because no
margin-of-error formula was applied. The other four rates have the same
estimates as before. With the default formula_dispatch, all
four use the ratio formula, so labor_force_participation
now has the ratio-formula margin of error computed for comparison in the
second check, and poverty_rate records no substitution.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.