The hardware and bandwidth for this mirror is donated by METANET, the Webhosting and Full Service-Cloud Provider.
If you wish to report a bug, or if you are interested in having us mirror your free-software or open-source project, please feel free to contact us at mirror[@]metanet.ch.

evaluatellm 0.1.0

Pre-submission audit fixes, ahead of the first CRAN release.

First release.

Statistical inference for language model evaluations, following Miller (2024) doi:10.48550/arXiv.2411.00640 for the core standard error and experiment design results, and the prediction-powered inference literature for the model judge functions.

Scoring

Comparison

Planning

Model judges

Leaderboards

Reporting

These binaries (installable software) and packages are in development.
They may not be fully stable and should be used with caution. We make no claims about them.