IT EventsBook

Discussions

Community guidelines

Be specific and constructive. No vendor spam — promoting your own product belongs in a listing. Anyone can read; posting needs a free account.

Kubernetes GPU sche...
 
Notifications
Clear all
Kubernetes GPU scheduling... quotas and queues, what actually works?
5 Posts
3 Users
0 Reactions
3 Views
argo_al
(@argo_al)
Active Member
Joined: 3 weeks ago
Posts: 9
Topic starter   [#57]

We added a few NVIDIA GPU nodes for training and inference. Now one team reserves a whole GPU for a notebook while another has jobs waiting. MIG helps on some cards, but model sizes vary and the normal scheduler doesn't understand business priority. What are people using?



   
Quote
finops_k8s
(@finops_k8s)
Active Member
Joined: 1 month ago
Posts: 8
 

yeah, we stopped treating training jobs like normal deployments. They go through a queue with quotas and priority. Inference stays as services. That alone made utilization easier to reason about



   
ReplyQuote
idp_builder
(@idp_builder)
Active Member
Joined: 1 month ago
Posts: 6
 

MIG is useful when profiles match the workload, but it's not free flexibility. Changing profiles can mean draining the node. Measure memory and duty cycle first. Some small inference services did better sharing through the serving layer than getting their own slice.



   
ReplyQuote
argo_al
(@argo_al)
Active Member
Joined: 3 weeks ago
Posts: 9
Topic starter  

edit: the notebook use is the political part. Researchers want instant access and production wants predictable latency. Separate pools may be the only sane answer.



   
ReplyQuote
finops_k8s
(@finops_k8s)
Active Member
Joined: 1 month ago
Posts: 8
 

yeah, we use a small interactive pool with time limits and a larger queued pool. Idle notebooks are stopped automatically. Production has dedicated capacity and can burst only if the batch queue is empty. Nobody loves the policy, which is usually a sign it's reasonably fair



   
ReplyQuote
Share:
Scroll to Top