Statistical Inference for Language Model Evaluations


[Up] [Top]

Documentation for package ‘evaluatellm’ version 0.1.0

Help Pages

as_eval Build an Evaluation Object
ev_bootstrap Cluster Bootstrap for an Arbitrary Statistic
ev_cluster Cluster-Robust Standard Error for an Evaluation
ev_elo Bradley-Terry Ratings from Pairwise Preferences
ev_icc Intra-Cluster Correlation
ev_judge_agreement Agreement Between a Model Judge and a Human Gold Standard
ev_judge_debias Debias a Model Judge with a Small Human Sample
ev_judge_power How Many Human Labels a Debiased Evaluation Needs
ev_mde Smallest Difference an Evaluation Can Detect
ev_multi Multiplicity Adjustment Across a Benchmark Suite
ev_paired Paired Comparison of Two Models
ev_plot Plot Results with Error Bars
ev_power Number of Questions Needed to Detect a Difference
ev_rank Bootstrap Rank Intervals for a Leaderboard
ev_resample Variance Decomposition for Repeated Sampling
ev_score Evaluation Score with a Standard Error
ev_table Collect Results into a Table
ev_unpaired Unpaired Comparison of Two Models
ev_variance_reduction Variance Reduction with a Reference Model