Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud
Nvidia
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
100 open kubernetes roles across 41 companies are on ApplySarthi right now, most of them in Bengaluru (11), Hyderabad (4), Pune (3).
- Senior Application Support & DevOps Engineer (Kubernetes)Fis
- Principal Kubernetes Platform EngineerMastercard
- Sr Engineer - Managed KubernetesTarget · bengaluru
- Lead Kubernetes Platform Engineer3M
- Kubernetes Platform ArchitectBroadcom
What kubernetes roles keep asking for: Kubernetes (73%), AWS (48%), CI/CD (44%), Observability (41%), Python (36%), Terraform (35%), Azure (30%), GCP (27%) — counted across their open postings here.
AWS jobs · Azure jobs · CI/CD jobs · GCP jobs
Nvidia has 2,295 open roles listed here.
- Senior Compiler Optimization Engineer – LLVMbengaluru
- Engineering Manager - OpenBMC Platform
- Principal Firmware Engineer - Data Center Server Management
- Senior Systems Software Engineer- EDA Infrastructure
- Senior Data Backend Engineer
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for kubernetes roles keep coming back to Kubernetes, AWS, CI/CD, Observability. Practise those questions before you sit with Nvidia.
Questions you are likely to be asked
- Why do you want to join Nvidia?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- How do you decide what to monitor, and what should wake someone up at night?
- How would you cut the cloud bill of a system without hurting it?
- How do you keep secrets and access safe in your infrastructure?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Systems Software Engineer, Kubernetes Scale - DGX Cloud at Nvidia interview free →The DGX Cloud organization at NVIDIA brings together cutting-edge hardware and software innovation to deliver industry-leading accelerated computing for the world's most adventurous AI workloads. We're a team of innovative engineers dedicated to solving some of the world's biggest challenges, constantly driving advancements, and impacting millions of lives worldwide! We are looking for an outstanding Senior Systems Software Engineer with deep experience in distributed systems, open-source technologies such as Kubernetes and containers, and a strong background in systems performance and scalability. The ideal candidate brings broad, end-to-end experience across the stack - from GPU operator and device plugins to distributed inference serving and cloud platforms - along with the technical depth to investigate and address exciting, real-world problems at scale. In this pivotal role, you will take on the challenge of scaling AI infrastructure while optimizing total cost of ownership, driving down cost per token to unlock the next generation of AI innovation and AI factories! What you'll be doing: Drive end-to-end performance and scale characterization for the NVIDIA DGX Cloud software stack, from Kubernetes control and data planes through NVIDIA components such as GPU Operator, Network Operator, DCGM, NIM, and distributed inference serving, following issues from orchestration down to the metal. Collaborate with AI researchers, developers and customers to develop innovative, automated tests that simulate real user workloads using custom-built and leading open-source tools and frameworks. Deep dive into performance and scale issues in complex distributed systems, including interactions between Kubernetes and the NVIDIA software stack, to identify and resolve root causes. Design and develop monitoring, reporting and analysis tools for performance and scale testing across software, GPU and CPU resources. Triage, debug and root cause issues related to operating Kubernetes clusters at ultra-large scale, ensuring reliability and efficiency. Build and maintain a high-velocity framework that enables continuous, always-on performance and scale testing via a modern CI/CD pipeline. Document research, methodologies and results clearly and concisely, and present findings at internal and external venues, including community conferences such as KubeCon and GTC. Engage efficiently with upstream communities — including Kubernetes, CNCF and NVIDIA open-source projects — to validate performance and scalability of AI workloads early and help shape design and development decisions. What we need to see: 8+ years of experience Computer Architecture, Networking, Storage systems, Accelerators and Bachelors/Masters in Engineering (preferably, Electrical Engineering, Computer Engineering, or Computer Science) or equivalent experience Expertise in Kubernetes and familiarity with related CNCF projects Background in working with large scale parallel and distributed accelerator-based systems Expertise optimizing performance and AI workloads on large scale systems Experience with performance modeling and benchmarking at scale Proficiency in Golang/Python Background with the NVIDIA software ecosystem in both training and inference domains Expertise with at least one of public CSP infrastructure (GCP, AWS, Azure, OCI for example) Ways to stand out from the crowd: Strong operational experience with any one of the Kubernetes distributions Prior experience scaling Kubernetes clusters to ultra-large node and object counts Demonstrated history of working in the open-source community Excellent communication and interpersonal abilities PhD in relevant areas NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Head of Startups - India and South AsiaNvidia · bengaluru
- Manager, AI/HPC Infrastructure Technical Delivery — IndiaNvidia · pune
- Architect – AI-Powered Performance Verification AutomationNvidia · bengaluru
- Data Center Infrastructure SpecialistNvidia · bengaluru
- Senior Developer Relations ManagerNvidia · bengaluru
- Server Performance Architect - HardwareNvidia · bengaluru
- Senior Solutions Architect, Infiniband and Networking Ethernet - NVISNvidia · bengaluru
- Senior Software Engineer, Fabric Networking - GPUNvidia · bengaluru
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on workday · posted 2026-10-02. ApplySarthi collects openings and links to application pages; the role is advertised by Nvidia, not by us.