The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

coldstart

Calibrate AI-generated items before you have the pretest seats to do it the old way.

Automatic item generation and LLM-assisted writing produce items faster than pretesting can calibrate them. coldstart predicts each new item’s difficulty from its features, uses the prediction as a robust prior, finds the families where prediction fails, and tells you how many responses each item needs.

library(coldstart)

sim  <- cs_simulate(seed = 1)                     # legacy + new items with features
it   <- sim$items; tr <- it$set == "train"

pr   <- cs_predictor(it$b_legacy[tr], sim$features[tr, ], it$family[tr])
pred <- predict(pr, sim$features[!tr, ], it$family[!tr])   # mean + honest SD

cs_plan(pred, target_sd = 0.3)                    # responses needed per item

resp <- cs_responses(sim, n_per_item = 25)        # or your pretest data
cal  <- cs_calibrate(resp, pred)                  # t-prior Bayes vs baseline
chk  <- cs_check(cal, setNames(it$family[!tr], it$item[!tr]))
cal  <- cs_calibrate(resp, cs_distrust(pred, chk))  # drop priors that failed

Installation

From CRAN (once released):

install.packages("coldstart")

Development version from GitHub:

install.packages("pak")
pak::pak("edidatasolutions/coldstart")

Features can be anything numeric: embeddings from any text model, cognitive attribute codes, content metadata. The package does not call a model itself.

Design choices

Validation (known truth, 5 replications, inst/validation/known_truth.R)

Predictive SDs are honest: stated 0.54 vs actual RMSE 0.53 (seen families), 0.70 vs 0.64 (unseen family); 90% intervals cover 90.8% and 90.4%.

RMSE of difficulty, seen families:

responses per item baseline predicted prior
15 0.72 0.40
25 0.53 0.35
50 0.36 0.30
100 0.27 0.23

With the prior, 25 responses do what 50 do without it.

Drifted (“rogue”) template family (+1.2 logits vs its history): the prior alone hurts (0.66 vs 0.61 at n = 25). cs_check flags the family in 80% of replications at n = 25 and 100% at n ≥ 50, with 0.2 false family flags per replication. After cs_distrust(), RMSE is back at baseline (0.57 at n = 25).

Planner: a target posterior SD of 0.30 needs a median of 40 responses per item with the prior vs 57 without; achieved SD 0.302, RMSE 0.300.

Status

Done: cs_simulate, cs_responses, cs_predictor (+predict), cs_calibrate, cs_plan, cs_check, cs_distrust. Next: 2PL (discrimination priors), sequential updating as responses arrive, ability uncertainty for pretest examinees (currently treated as known from operational scoring), and non-linear predictors.

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.