| as_eval | Build an Evaluation Object |
| ev_bootstrap | Cluster Bootstrap for an Arbitrary Statistic |
| ev_cluster | Cluster-Robust Standard Error for an Evaluation |
| ev_elo | Bradley-Terry Ratings from Pairwise Preferences |
| ev_icc | Intra-Cluster Correlation |
| ev_judge_agreement | Agreement Between a Model Judge and a Human Gold Standard |
| ev_judge_debias | Debias a Model Judge with a Small Human Sample |
| ev_judge_power | How Many Human Labels a Debiased Evaluation Needs |
| ev_mde | Smallest Difference an Evaluation Can Detect |
| ev_multi | Multiplicity Adjustment Across a Benchmark Suite |
| ev_paired | Paired Comparison of Two Models |
| ev_plot | Plot Results with Error Bars |
| ev_power | Number of Questions Needed to Detect a Difference |
| ev_rank | Bootstrap Rank Intervals for a Leaderboard |
| ev_resample | Variance Decomposition for Repeated Sampling |
| ev_score | Evaluation Score with a Standard Error |
| ev_table | Collect Results into a Table |
| ev_unpaired | Unpaired Comparison of Two Models |
| ev_variance_reduction | Variance Reduction with a Reference Model |