On the discriminative power of credit scoring systems trained on independent samples

The aim of this work is to assess the importance of independence assumption in behavioral scorings created using logistic regression. We develop four sampling methods that control which observations associated to each client are to be included in the training set, avoiding a functional dependence between observations of the same client. We then calibrate logistic regressions with variable selection on the samples created by each method, plus one using all the data in the training set (biased base method), and validate the models on an independent data set. We find that the regression built using all the observations shows the highest area under the ROC curve and Kolmogorv–Smirnov statistics, while the regression that uses the least amount of observations shows the lowest performance and highest variance of these indicators. Nevertheless, the fourth selection algorithm presented shows almost the same performance as the base method using just 14 % of the dataset, and 14 less variables. We conclude that violating the independence assumption does not impact strongly on results and, furthermore, trying to control it by using less data can harm the performance of calibrated models, although a better sampling method does lead to equivalent results with a far smaller dataset needed

10.1007/978-3-319-01595-8_27

247-254

Springer

Biron, Miguel

6ad4f5b8-d615-42b1-bd98-a2ae254869d0

Bravo, Cristian

b22c4145-644e-40ee-85d8-431c59c3c71b

10 October 2013

Biron, Miguel

6ad4f5b8-d615-42b1-bd98-a2ae254869d0

Bravo, Cristian

b22c4145-644e-40ee-85d8-431c59c3c71b

Biron, Miguel and Bravo, Cristian (2013) On the discriminative power of credit scoring systems trained on independent samples. In, Data Analysis, Machine Learning and Knowledge Discovery. (Data Analysis, Machine Learning and Knowledge Discovery) New York, US. Springer, pp. 247-254. (doi:10.1007/978-3-319-01595-8_27).