Understanding the effect of collocating heterogeneous ML models on GPUs for scaling inference services
The goal of this thesis is to understand how the spatial co-deployment of different ML models affects their performance on GPUs and to derive criteria for placing and allocating model instances (as a responsibility of inference serving systems) to preserve performance at a cost-efficient level.