Clustering-based validation splits for model selection under domain shift

Napoli, Andrea and White, Paul (2025) Clustering-based validation splits for model selection under domain shift. Transactions on Machine Learning Research.

Record type: Article

Abstract

This paper considers the problem of model selection under domain shift. Motivated by principles from distributionally robust optimisation and domain adaptation theory, it is proposed that the training-validation split should maximise the distribution mismatch between the two sets. By adopting the maximum mean discrepancy (MMD) as the measure of mismatch, it is shown that the partitioning problem reduces to kernel k-means clustering. A constrained clustering algorithm, which leverages linear programming to control the size, label, and (optionally) group distributions of the splits, is presented. The algorithm does not require additional metadata, and comes with convergence guarantees. In experiments, the technique consistently outperforms alternative splitting strategies across a range of datasets and training algorithms, for both domain generalisation and unsupervised domain adaptation tasks. Analysis also shows the MMD between the training and validation sets to be well-correlated with test domain accuracy, further substantiating the validity of this approach.

Text

Clustering-Based Validation Splits for Model Selection under Domain Shift - Version of Record

Available under License Creative Commons Attribution.

Download (567kB)