The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

Inference from supplied influence contributions

The core of the package accepts an estimate and the observation-level influence contributions of its asymptotically linear representation. This vignette explains the interface, the conventions, and the conditions that must hold for the inference to be valid. All data are synthetic.

The representation

Let \(\hat\theta_n \in \mathbb R^q\) estimate \(\theta_0\) from \(n\) observations in time order, with \[\sqrt n(\hat\theta_n - \theta_0) = n^{-1/2}\sum_{t=1}^n \psi_t + o_p(1),\] where \(\psi_t \in \mathbb R^q\) is the (mean-zero) influence function. The user supplies an estimate \(\hat\psi_{n,t}\) of each contribution. Column \(j\) of the contribution matrix corresponds to coordinate \(j\) of estimate, and nrow(psi) must equal the number of observations in the sum.

library(aersn)
set.seed(3)
n <- 200
## A ratio of two means, with the contributions derived by the delta method.
Y <- cbind(rnorm(n, 2), rnorm(n, 1))
m <- colMeans(Y)
ratio <- m[2] / m[1]
J <- matrix(c(-m[2] / m[1]^2, 1 / m[1]), 1, 2)   # derivative of h(m) = m2 / m1
psi <- sweep(Y, 2, m) %*% t(J)                     # n x 1 contributions
fit <- aersn(ratio, psi, n = n, names = "ratio")
fit
#> Affine-equivariant adjusted-range self-normalization (increment hull)
#>   model: user-supplied influence contributions 
#>   n = 200 observations; q = 1 parameter
#>   centering: calendar time, tau(r) = r 
#>   estimate:
#>  ratio 
#> 0.5181 
#>   hull diagnostics: numerical rank 1 of 1  ; condition 1.000 ; min projected range 0.6113 ; spread 1.000 
#> Use aersn_test(), confint(), aersn_contrast(), aersn_region(), plot().
aersn_test(fit, null = 0.5, draws = 20000, seed = 1)
#> 
#>  Adjusted-range increment hull test, q = 1
#> 
#> data:  fit
#> T = 0.4195, q = 1, n = 200
#> alternative hypothesis: true parameter is not equal to the null value
#> null value: ratio = 0.5000 
#> estimate:   ratio = 0.5181 
#> critical value at level 0.95: 1.849  (do not reject)
#> p-value = 0.623 (Monte Carlo s.e. 0.0034, resolution 0.000050)
#> reference: matched-grid Monte Carlo law for increment-hull gauge: q = 1, uniform grid with n = 200 intervals, 20000 draws, seed 1

The same result is obtained by fitting the mean vector and transforming the target with aersn_target(), which applies the Jacobian to the contributions:

fit2 <- aersn_target(aersn_mean(Y), function(b) b[2] / b[1], jacobian = J,
                     names = "ratio")
all.equal(fit$psi, fit2$psi, check.attributes = FALSE)
#> [1] TRUE

Conventions made explicit

Conditions required for validity

The package computes the statistic and its reference law; it cannot verify that the conditions of the manuscript hold for a user-supplied estimator. Those conditions are (manuscript Assumption 1):

  1. Functional limit. The partial-sum process of the true contributions, \(n^{-1/2}\sum_{t \le \lfloor nr \rfloor}\psi_t\), converges weakly to \(A B_q(r)\) with \(A\) nonsingular, so the long-run covariance matrix \(\Omega = AA^\top\) is positive definite.
  2. Asymptotic linearity. The estimation error satisfies the representation above.
  3. Feasible path. The increment hull of the path built from the estimated contributions converges in Hausdorff distance to the hull of the path built from the true contributions. Uniform convergence of the estimated path is sufficient.

Condition 3 is where estimated nuisance parameters matter. For smooth GMM and for conditional likelihood the manuscript verifies it when the contributions incorporate the nuisance-parameter effect through the Jacobian and the inverse Hessian (see aersn_gmm() and aersn_mle()). Discarding nuisance coordinates from raw moment contributions or scores does not give the correct path.

Nuisance parameters: a concrete illustration

For a regression slope with an estimated intercept, the correct contribution is the second element of \(\hat Q^{-1} z_t \hat u_t\), not \(x_t \hat u_t\) alone. The two differ whenever the regressor is not centered:

x <- rnorm(n) + 1
y <- 1 + 0.5 * x + rnorm(n)
Z <- cbind(1, x); b <- solve(crossprod(Z), crossprod(Z, y)); u <- as.numeric(y - Z %*% b)
correct <- ((Z * u) %*% solve(crossprod(Z) / n))[, 2]
naive <- x * u / mean(x^2)
c(correct = aersn_gauge(aersn(b[2], correct), 0.5),
  naive = aersn_gauge(aersn(b[2], naive), 0.5))
#>   correct     naive 
#> 0.4071279 0.6705966

aersn_lm() and aersn_gmm() form the correct contributions automatically.

Effective sample size

The statistic is not invariant to n: multiplying the number of rows by a constant while keeping the same estimate changes both the path scaling and the estimation-error scaling. Supply exactly the contributions of the representation, one per observation.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.