The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.
Calibrate AI-generated items before you have the pretest seats to do it the old way.
Automatic item generation and LLM-assisted writing produce items faster than pretesting can calibrate them. coldstart predicts each new item’s difficulty from its features, uses the prediction as a robust prior, finds the families where prediction fails, and tells you how many responses each item needs.
library(coldstart)
sim <- cs_simulate(seed = 1) # legacy + new items with features
it <- sim$items; tr <- it$set == "train"
pr <- cs_predictor(it$b_legacy[tr], sim$features[tr, ], it$family[tr])
pred <- predict(pr, sim$features[!tr, ], it$family[!tr]) # mean + honest SD
cs_plan(pred, target_sd = 0.3) # responses needed per item
resp <- cs_responses(sim, n_per_item = 25) # or your pretest data
cal <- cs_calibrate(resp, pred) # t-prior Bayes vs baseline
chk <- cs_check(cal, setNames(it$family[!tr], it$item[!tr]))
cal <- cs_calibrate(resp, cs_distrust(pred, chk)) # drop priors that failedFrom CRAN (once released):
install.packages("coldstart")Development version from GitHub:
install.packages("pak")
pak::pak("edidatasolutions/coldstart")Features can be anything numeric: embeddings from any text model, cognitive attribute codes, content metadata. The package does not call a model itself.
cs_distrust() withdraws
priors from failing families.inst/validation/known_truth.R)Predictive SDs are honest: stated 0.54 vs actual RMSE 0.53 (seen families), 0.70 vs 0.64 (unseen family); 90% intervals cover 90.8% and 90.4%.
RMSE of difficulty, seen families:
| responses per item | baseline | predicted prior |
|---|---|---|
| 15 | 0.72 | 0.40 |
| 25 | 0.53 | 0.35 |
| 50 | 0.36 | 0.30 |
| 100 | 0.27 | 0.23 |
With the prior, 25 responses do what 50 do without it.
Drifted (“rogue”) template family (+1.2 logits vs
its history): the prior alone hurts (0.66 vs 0.61 at n = 25).
cs_check flags the family in 80% of replications at n = 25
and 100% at n ≥ 50, with 0.2 false family flags per replication. After
cs_distrust(), RMSE is back at baseline (0.57 at n =
25).
Planner: a target posterior SD of 0.30 needs a median of 40 responses per item with the prior vs 57 without; achieved SD 0.302, RMSE 0.300.
Done: cs_simulate, cs_responses,
cs_predictor (+predict),
cs_calibrate, cs_plan, cs_check,
cs_distrust. Next: 2PL (discrimination priors), sequential
updating as responses arrive, ability uncertainty for pretest examinees
(currently treated as known from operational scoring), and non-linear
predictors.
These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.