Community guidelines
Be specific and constructive. No vendor spam — promoting your own product belongs in a listing. Anyone can read; posting needs a free account.
We added a few NVIDIA GPU nodes for training and inference. Now one team reserves a whole GPU for a notebook while another has jobs waiting. MIG helps on some cards, but model sizes vary and the normal scheduler doesn't understand business priority. What are people using?
yeah, we stopped treating training jobs like normal deployments. They go through a queue with quotas and priority. Inference stays as services. That alone made utilization easier to reason about
MIG is useful when profiles match the workload, but it's not free flexibility. Changing profiles can mean draining the node. Measure memory and duty cycle first. Some small inference services did better sharing through the serving layer than getting their own slice.
edit: the notebook use is the political part. Researchers want instant access and production wants predictable latency. Separate pools may be the only sane answer.
yeah, we use a small interactive pool with time limits and a larger queued pool. Idle notebooks are stopped automatically. Production has dedicated capacity and can burst only if the batch queue is empty. Nobody loves the policy, which is usually a sign it's reasonably fair