Toward Estimating the Rank Correlation between the Test Collection Results and the True System Performance.
Urbano J, Marrero M. Toward Estimating the Rank Correlation between the Test Collection Results and the True System Performance. International ACM SIGIR Conference on Research and Development in Information Retrieval, 2016
The Kendall and AP rank correlation coefficients have become mainstream in Information Retrieval research for comparing the rankings of systems produced by two different evaluation conditions, such as di erent e ectiveness measures or pool depths. However, in this paper we focus on the expected rank correlation between the mean scores observed with a test collection and the true unobservable means under the same conditions. In particular, we propose statistical estimators of and AP correlations following both parametric and non-parametric approaches, and with special emphasis on small topic sets. Through large scale simulation with TREC data, we study the error and bias of the estimators. In general, such estimates of expected correlation with the true ranking may accompany the results reported from an evaluation experiment, as an easy to understand gure of reliability. All the results in this paper are fully reproducible with data and code available online.
Keywords: Evaluation; Test Collection; Correlation; Kendall; Average Precision; Estimation