| Type: | Package |
| Title: | Implementation of Logistic Box-Cox Regression |
| Version: | 2.0.1 |
| Date: | 2026-09-05 |
| Maintainer: | Li Xing <jnjfayaa@gmail.com> |
| Description: | Implements a logistic Box-Cox model that adds a shape parameter to a routine logistic regression model to flexibly estimate the shape and strength of the relationship between a binary outcome and a continuous predictor, adjusting for covariates and survey weights. This model is fully described in Xing, L. et al. (2021) <doi:10.1002/cjs.11587>. This version extends the original 'lboxcox' package (1.1) with numerically stabilized likelihood/gradient calculations, vectorized data preprocessing, and a bootstrap-ensemble estimator. |
| License: | GPL-3 |
| Encoding: | UTF-8 |
| LazyData: | true |
| RoxygenNote: | 7.3.3 |
| Suggests: | knitr, rmarkdown, testthat (≥ 3.0.0) |
| VignetteBuilder: | knitr |
| Config/testthat/edition: | 3 |
| Depends: | R (≥ 3.5.0), survey |
| Imports: | maxLik, caret, doParallel, foreach, MASS, stats |
| NeedsCompilation: | no |
| Packaged: | 2026-10-02 16:47:39 UTC; xushiyu |
| Author: | Li Xing [cre, aut], Shiyu Xu [aut], Jing Wang [aut], Kohlton Booth [aut], Xuekui Zhang [aut], Igor Burstyn [aut], Paul Gustafson [aut] |
| Repository: | CRAN |
| Date/Publication: | 2026-10-04 23:00:22 UTC |
lboxcox: Logistic Box-Cox Regression
Description
Implements a logistic Box-Cox model that adds a shape parameter to a routine logistic regression model, allowing for a flexible relationship between the log-odds of a binary outcome and a continuous predictor. See Xing, L. et al. (2021) for the full model description.
Fitting a model
lbc_maxlik, lbc_train_ms,
lbc_train_bagging, and lbc_train_all fit the
logistic Box-Cox model with, respectively, a single starting value,
multiple starting values, a bootstrap ensemble around
lbc_maxlik, and a bootstrap ensemble around
lbc_train_ms. lboxcox_cv.fit instead selects
the Box-Cox shape parameter by cross-validated deviance residual.
Predicting and interpreting a fitted model
lboxcox_maxLik.predict and lboxcox_maxLik_el.predict
produce predicted probabilities from a fitted model, lboxcox_cv.predict
does the same for a lboxcox_cv.fit model, median_effect
summarizes the fitted exposure-response slope, and devr computes
the deviance-residual prediction error used throughout the package.
Author(s)
Maintainer: Li Xing jnjfayaa@gmail.com
Authors:
Shiyu Xu xushiyu98@outlook.com
Jing Wang
Kohlton Booth booth.kohl@gmail.com
Xuekui Zhang
Igor Burstyn
Paul Gustafson
References
Xing, L., Zhang, X., Burstyn, I., & Gustafson, P. (2021). On logistic Box-Cox regression for flexibly estimating the shape and strength of exposure-disease relationships. Canadian Journal of Statistics, 49(3), 808-825.
Depression dataset
Description
The depress data frame contains 8,893 adults aged 20 years or older from the 2005-2006 and 2007-2008 National Health and Nutrition Examination Survey (NHANES) cycles.
Usage
depress
Format
Sample survey data with columns:
- depression
binary response variable equal to 1 for a Patient Health Questionnaire-9 (PHQ-9) score of at least 10 and 0 otherwise
- mercury
total blood mercury concentration in micrograms per litre
- age
age of the participant in years
- gender
0 if the participant is female and 1 if the participant is male
- weight
four-year Day 1 dietary sampling weight, formed as
WTDRD1 / 2for the two combined NHANES cycles
Source
National Center for Health Statistics, NHANES 2005-2006 and 2007-2008. The analytic data were adapted from the data used by Xing, L., Zhang, X., Burstyn, I., & Gustafson, P. (2021). On logistic Box-Cox regression for flexibly estimating the shape and strength of exposure-disease relationships. Canadian Journal of Statistics, 49(3), 808-825.
Deviance residual prediction error
Description
Computes the sum of absolute signed deviance residuals between binary
outcomes and predicted probabilities. Used as the cross-validation
criterion in lboxcox_cv.fit and, more generally, as a
prediction-error summary for comparing fitted models.
Usage
devr(y.bin, y.pred)
Arguments
y.bin |
binary (0/1) outcome vector. |
y.pred |
predicted probabilities, same length as |
Value
a single numeric value: the sum of absolute deviance residuals.
Fit a logistic Box-Cox model by direct likelihood maximization
Description
Trains the given formula using a logistic Box-Cox model whose parameters
(intercept, slope, Box-Cox shape, and covariate coefficients) are found by
maximizing the log-likelihood directly with maxLik::maxLik. If the
resulting lambda estimate falls outside [0, 2], falls back to the
best of a grid of svyglm fits at different lambda values (see
lbc_train_ms for a version that always searches the grid).
Usage
lbc_maxlik(
formula,
weight_column_name,
data,
init = NULL,
svy_lambda_vector = seq(0, 2, length = 4),
init_lambda_vector = seq(0, 2, length = 100),
num_cores = 1,
seed = NULL
)
Arguments
formula |
a formula of the form |
weight_column_name |
the name of the column in |
data |
dataframe containing the dataset to train on. |
init |
initial estimates for the coefficients. If |
svy_lambda_vector |
values of lambda used in training the |
init_lambda_vector |
values of lambda used, as a fallback, to find the
best-log-likelihood starting point if the direct maximization does not
converge to a lambda in |
num_cores |
the number of cores used when searching |
seed |
deprecated compatibility argument; ignored because the BFGS
optimizer is deterministic. Control bootstrap reproducibility by calling
|
Value
object of class maxLik from the maxLik package (or,
on fallback, of class svyglm). Contains the coefficient estimates
in $estimate (named Beta_0, Beta_1, Lambda,
then covariates) and a convergence flag in $error.
Note
This is reliant on the following work:
Henningsen, A., Toomet, O. (2011). maxLik: A package for maximum likelihood estimation in R. Computational Statistics, 26(3), 443-458.
Bootstrap-ensemble logistic Box-Cox model around lbc_train_ms
Description
Fits lbc_train_ms on each of 100 bootstrap resamples of
data in parallel, takes the median of the per-resample lambda
estimates, and refits the final model at that lambda on the full data.
More expensive than lbc_train_bagging (each resample itself
searches a lambda grid via lbc_train_ms) but more robust to
poor starting values within each resample.
Bootstrap samples use the caller's current random-number state. Call
set.seed() before this function when reproducible resamples are
required.
Usage
lbc_train_all(
formula,
weight_column_name,
data,
init = NULL,
svy_lambda_vector = seq(0, 2, length = 100),
num_cores = 1,
cores = 5
)
Arguments
formula |
a formula of the form |
weight_column_name |
the name of the column in |
data |
dataframe containing the dataset to train on. |
init |
initial estimates for the coefficients. If |
svy_lambda_vector |
values of lambda used to build the grid of
starting vectors that |
num_cores |
the number of cores used when searching |
cores |
number of cores to parallelize the bootstrap resamples over. |
Value
a list of length 3: the final refit model (as returned by
lbc_train_ms), the list of per-resample fits (NA
for any resample where lbc_train_ms errored), and the
count of resamples that errored.
Note
This is reliant on the following work:
Microsoft Corporation, Weston, S. (2020). foreach: Provides Foreach Looping Construct. R package version 1.5.1.
Microsoft Corporation, Weston, S. (2020). doParallel: Foreach Parallel Adaptor for the 'parallel' Package. R package version 1.0.16.
Bootstrap-ensemble logistic Box-Cox model around lbc_maxlik
Description
Fits lbc_maxlik on each of 100 bootstrap resamples of
data in parallel, takes the median of the per-resample lambda
estimates, and refits the final model at that lambda on the full data.
Bootstrap samples use the caller's current random-number state. Call
set.seed() before this function when reproducible resamples are
required.
Usage
lbc_train_bagging(
formula,
weight_column_name,
data,
init = NULL,
svy_lambda_vector = seq(0, 2, length = 4),
init_lambda_vector = seq(0, 2, length = 100),
num_cores = 1,
cores = 10
)
Arguments
formula |
a formula of the form |
weight_column_name |
the name of the column in |
data |
dataframe containing the dataset to train on. |
init |
initial estimates for the coefficients. If |
svy_lambda_vector |
values of lambda used in training the |
init_lambda_vector |
values of lambda used, as a fallback, to find the
best-log-likelihood starting point if the direct maximization does not
converge to a lambda in |
num_cores |
the number of cores used when searching |
cores |
number of cores to parallelize the bootstrap resamples over. |
Value
a list of length 3: the final refit model (as returned by
lbc_maxlik/its fallback), the list of per-resample fits
(NA for any resample where lbc_maxlik errored), and
the count of resamples that errored.
Note
This is reliant on the following work:
Microsoft Corporation, Weston, S. (2020). foreach: Provides Foreach Looping Construct. R package version 1.5.1.
Microsoft Corporation, Weston, S. (2020). doParallel: Foreach Parallel Adaptor for the 'parallel' Package. R package version 1.0.16.
Fit a logistic Box-Cox model from multiple starting points
Description
Trains the given formula using a logistic Box-Cox model by running
maxLik::maxLik from a grid of svyglm-based starting vectors
(one per entry of svy_lambda_vector) and keeping the fit with the
highest achieved log-likelihood. More expensive than
lbc_maxlik but more robust to a poor single starting point.
Usage
lbc_train_ms(
formula,
weight_column_name,
data,
init = NULL,
svy_lambda_vector = seq(0, 2, length = 100),
num_cores = 1
)
Arguments
formula |
a formula of the form |
weight_column_name |
the name of the column in |
data |
dataframe containing the dataset to train on. |
init |
initial estimates for the coefficients. If |
svy_lambda_vector |
values of lambda used to build the grid of
starting vectors that |
num_cores |
the number of cores used when searching |
Value
the refit svyglm model at the selected lambda, with
$estimate (named Beta_0, Beta_1, Lambda, then
covariates) and a convergence flag in $error; error = 1
means no continuous BFGS run converged and the exact profile grid was
used as a fallback.
Note
This is reliant on the following work:
Henningsen, A., Toomet, O. (2011). maxLik: A package for maximum likelihood estimation in R. Computational Statistics, 26(3), 443-458.
Select the Box-Cox lambda by cross-validated deviance residual
Description
Trains the given formula using a logistic Box-Cox model whose lambda is
chosen, instead of by likelihood maximization, as the value in
lambda_vector with lowest mean k-fold cross-validated
devr (deviance residual prediction error). The model is then
refit on the full data at the selected lambda.
Usage
lboxcox_cv.fit(
mydata,
ixx,
iyy,
formula,
weight_column_name = NULL,
lambda_vector = seq(0, 2, length = 100),
k
)
Arguments
mydata |
data frame containing the dataset to train on. |
ixx |
the (untransformed) primary predictor vector, e.g. |
iyy |
the binary outcome vector, e.g. |
formula |
a formula of the form |
weight_column_name |
the name of the column in |
lambda_vector |
grid of Box-Cox lambda values to select from. |
k |
number of cross-validation folds. |
Value
the refit svyglm model at the selected lambda, with
$estimate (named Beta_0, Beta_1, Lambda, then
covariates) attached.
Predict from a lboxcox_cv.fit model
Description
Predict from a lboxcox_cv.fit model
Usage
lboxcox_cv.predict(myCVfit, newdata, formula)
Arguments
myCVfit |
a model fit by |
newdata |
data frame of new observations to predict on. |
formula |
the same formula used to fit |
Value
a numeric vector of predicted probabilities.
Predict from a fitted logistic Box-Cox model
Description
Produces predicted probabilities for newdata from a model fit by
lbc_maxlik, lbc_train_ms, or the first
($estimate-bearing) element returned by lbc_train_bagging
or lbc_train_all.
Usage
lboxcox_maxLik.predict(myMaxLikfit, newdata, formula)
Arguments
myMaxLikfit |
a fitted model with a named |
newdata |
data frame of new observations to predict on. |
formula |
the same formula used to fit |
Value
a numeric vector of predicted probabilities.
Predict from a bootstrap-ensemble logistic Box-Cox model
Description
Averages predicted probabilities across the (up to 100) per-resample fits
returned as the second element of lbc_train_bagging or
lbc_train_all, skipping resamples that failed to converge.
Usage
lboxcox_maxLik_el.predict(myMaxLik_elfit, newdata, formula)
Arguments
myMaxLik_elfit |
the length-3 list returned by
|
newdata |
data frame of new observations to predict on. |
formula |
the same formula used to fit |
Value
a numeric vector of predicted probabilities, averaged over the converged bootstrap resamples.
Median exposure-response slope of a fitted logistic Box-Cox model
Description
Calculates a number that represents the overall gradient measurement
between the (log-scale) predictor and the log-odds of the outcome,
evaluated at the survey-weighted mean of log(x), along with a
Wald-type 95% confidence interval derived from the model's Hessian.
Usage
median_effect(formula, weight_column_name, data, trained_model)
Arguments
formula |
the formula used to train the logistic Box-Cox model. |
weight_column_name |
the name of the column in |
data |
dataframe containing the dataset the model was trained on. |
trained_model |
the already-trained model, e.g. the output of
|
Value
a named numeric vector with elements median effect,
lower 95% ci, and upper 95% ci.