Reference distributions and reproducibility

The reference law of the increment-hull gauge is \(T_q = \gamma_{K(\mathbb B_q)}\{B_q(1)\}\), the gauge of the increment hull of a standard \(q\)-dimensional Brownian bridge evaluated at its independent Brownian endpoint. It depends on the dimension \(q\) but not on the long-run covariance matrix: the same matrix multiplies the estimation error and the normalizing set and cancels from the gauge. Serial dependence therefore enters only through the asymptotic approximation, not through the reference law.

Closed-form scalar law

For \(q = 1\), \(T_1 = |Z|/R\) with \(Z\) standard normal and \(R\) the range of an independent Brownian bridge. Its distribution function is \(\Pr(|M| \le c) = c - 2c^3\sum_{k\ge1}(c^2 + 4k^2)^{-3/2}\); the package evaluates it through an exponentially convergent Bessel-function representation and cross-checks the two forms in its tests.

library(aersn)
aersn_scalar_quantile(c(0.90, 0.95, 0.99))
#> [1] 1.397390 1.705776 2.367378
curve(aersn_scalar_density(x), -4, 4, ylab = "density of M = Z / R")

Matched-grid references

For a sample of size n, the manuscript uses quantiles of the statistic computed from a Brownian bridge on the same grid of n intervals. The finite-grid gauge decreases to the continuous-path gauge as the grid is refined, so matched-grid quantiles are larger:

sapply(c(50, 200, 1000), function(n)
  aersn_critical_value(aersn_reference(1, n = n, draws = 20000, seed = 1)))
#> [1] 2.042077 1.848527 1.752016
aersn_scalar_quantile(0.95)
#> [1] 1.705776

For \(q \ge 2\) the law is simulated with the same linear program used for the sample statistic. Every setting that changes the law is part of the object: dimension, grid (uniform n or profile nodes), number of draws, seed, chunk size and solver settings. Draws are generated in fixed-size chunks with chunk-specific seeds, so the result is reproducible for a given seed and does not depend on the number of cores.

r1 <- aersn_reference(2, n = 100, draws = 2000, seed = 5)
r2 <- aersn_reference(2, n = 100, draws = 2000, seed = 5)
identical(r1$draws, r2$draws)
#> [1] TRUE
r1
#> Reference distribution (aersn_reference): increment-hull gauge
#>   matched-grid Monte Carlo law for increment-hull gauge: q = 2, uniform grid with n = 100 intervals, 2000 draws, seed 5
#>   quantiles (type 8):
#>      90.0%: 2.2387  (Monte Carlo s.e. 0.040)
#>      95.0%: 2.6604  (Monte Carlo s.e. 0.056)
#>      99.0%: 3.5324  (Monte Carlo s.e. 0.10)

Monte Carlo standard errors of quantiles are estimated from interleaved batches; two-sided p-values report a standard error and a resolution of one over the number of draws. Scalar one-sided tails use symmetry and have half that resolution. aersn_pvalue() also reports a 95 percent binomial interval for the reference tail probability, so zero exceedances need not be interpreted as a zero true probability. Increase draws when a decision is borderline. The reported rejection decision compares the statistic to the interpolated quantile; empirical p-values can differ at the boundary by a Monte Carlo resolution unit.

A profile grid needs at least q + 1 positive increments. Flat segments are allowed when this condition holds. Numerical failures stop the simulation; failed draws are not removed to estimate quantiles. Grid matching approximates the Brownian reference law and does not make dependent-data inference finite-sample exact.

Matching statistics and references

aersn_test(), confint() and aersn_region() refuse a reference object that was simulated for a different dimension, grid size or profile:

set.seed(1)
fit <- aersn_mean(matrix(rnorm(300), 150, 2))
aersn_test(fit, reference = r1)
#> Error:
#> ! The reference distribution was computed on a grid with n = 100 intervals but the sample has n = 150.

The verified registry

aersn_registry contains matched-grid quantiles from the manuscript’s replication package (30,000 draws for \(q \ge 2\), 2,000,000 for \(q = 1\)), with batch Monte Carlo standard errors. The package tests compare their own simulations with these values. They can also be used directly as a tabulated reference for the available (q, n) pairs:

subset(aersn_registry, prob == 0.95)
#>     q    n prob       cv        mcse   draws batches quantile_type     seed
#> 2   1  200 0.95 1.849057 0.001331279 2000000     100             8 20260715
#> 6   1  500 0.95 1.794344 0.001285959 2000000     100             8 20260715
#> 10  1  738 0.95 1.777630 0.001297601 2000000     100             8 20260715
#> 14  1 1000 0.95 1.771534 0.001373603 2000000     100             8 20260715
#> 18  2  200 0.95 2.566424 0.012871566   30000     100             8 20260717
#> 22  2  500 0.95 2.495273 0.013014530   30000     100             8 20260717
#> 26  2 1000 0.95 2.439701 0.012533809   30000     100             8 20260717
#> 30  3  200 0.95 3.240205 0.015404391   30000     100             8 20260717
#> 34  3  500 0.95 3.122605 0.013821078   30000     100             8 20260717
#> 38  3 1000 0.95 3.043275 0.014095537   30000     100             8 20260717
#> 42  5  200 0.95 4.505366 0.018430026   30000     100             8 20260717
#> 46  5  500 0.95 4.283925 0.017070230   30000     100             8 20260717
#> 50  5  738 0.95 4.252784 0.016716769   30000     100             8 20260717
#> 54  5 1000 0.95 4.192039 0.015634613   30000     100             8 20260717
#> 58 10  500 0.95 7.210666 0.019740122   30000     100             8 20260717
#> 62 10 1000 0.95 6.993052 0.019777966   30000     100             8 20260717
#>                                  source
#> 2    matched_scalar_critical_values.csv
#> 6    matched_scalar_critical_values.csv
#> 10   matched_scalar_critical_values.csv
#> 14   matched_scalar_critical_values.csv
#> 18 multivariate_matched_cv_registry.csv
#> 22 multivariate_matched_cv_registry.csv
#> 26 multivariate_matched_cv_registry.csv
#> 30 multivariate_matched_cv_registry.csv
#> 34 multivariate_matched_cv_registry.csv
#> 38 multivariate_matched_cv_registry.csv
#> 42 multivariate_matched_cv_registry.csv
#> 46 multivariate_matched_cv_registry.csv
#> 50 multivariate_matched_cv_registry.csv
#> 54 multivariate_matched_cv_registry.csv
#> 58 multivariate_matched_cv_registry.csv
#> 62 multivariate_matched_cv_registry.csv
fit200 <- aersn_mean(matrix(rnorm(400), 200, 2))
aersn_test(fit200, reference = "registry")
#> 
#>  Adjusted-range increment hull test, q = 2
#> 
#> data:  fit200
#> T = 1.395, q = 2, n = 200
#> alternative hypothesis: true parameter is not equal to the null value
#> null value: theta1 = 0, theta2 = 0 
#> estimate:   theta1 = 0.006234, theta2 = -0.1220 
#> critical value at level 0.95: 2.566  (do not reject)
#> p-value bracket from tabulated quantiles: 0.100 < p <= 1.00
#> reference: tabulated quantiles for increment-hull gauge: q = 2, n = 200 (aersn_registry (multivariate_matched_cv_registry.csv))

Caching

References are cached within the session by their full key, including batches; aersn_clear_cache() empties the cache. Objects can be saved with saveRDS() and reused across sessions; they record the package version, seed and settings used to create them.