ApplySarthi Match jobs to your CV

AI Infrastructure System Engineer

Jobgether

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

2,034 open infrastructure roles across 290 companies are on ApplySarthi right now, most of them in Bengaluru (68), Hyderabad (28), Delhi NCR (12).

What infrastructure roles keep asking for: AWS (30%), System design (22%), Kubernetes (21%), Python (21%), Observability (19%), Terraform (15%), CI/CD (13%), Linux (12%) — counted across their open postings here.

Ansible jobs · Go jobs · Kubernetes jobs · Linux jobs

Jobgether has 3,942 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for infrastructure roles keep coming back to AWS, System design, Kubernetes, Python. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with System design? Tell me one thing you learned the hard way.
  3. What would you check first if a model's accuracy dropped after going live?
  4. When would you not use machine learning for a problem?
  5. Walk me through a model you built, from the data to how it was used.

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the AI Infrastructure System Engineer at Jobgether interview free →

Accountabilities:: Design and build fleet automation systems capable of provisioning, validating, deploying, upgrading, repairing, and retiring GPU clusters with minimal human intervention. Develop AI infrastructure agents that automate deployment workflows, investigate root causes, triage incidents, and support autonomous remediation. Build fleet intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance. Develop predictive capabilities that identify potential infrastructure failures before they affect customers or workloads. Build software and automation systems that maximize GPU availability, utilization, performance, and reliability across large accelerator fleets. Create automated validation frameworks for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage systems, and distributed AI workloads. Develop internal infrastructure platforms and developer tools that enable infrastructure to be managed programmatically rather than through manual operations. Continuously improve deployment velocity, system reliability, and operational efficiency through automation and software engineering. Collaborate with hardware, networking, platform, and AI teams to identify infrastructure challenges and develop scalable solutions. Apply strong systems thinking to problems spanning hardware and software components across large-scale AI infrastructure. Requirements: 3+ years of experience building distributed systems, infrastructure platforms, or large-scale backend software. Strong software engineering skills in Python, Go, or Rust. Proven experience developing platforms, automation systems, developer infrastructure, or similar software-driven infrastructure solutions. Experience working with Linux and modern infrastructure technologies such as Kubernetes, Terraform, Ansible, or comparable tools. Strong understanding of distributed systems and the ability to reason across hardware and software layers. Passion for solving complex infrastructure problems through software and automation. Strong automation-first mindset, with an instinct to build systems that eliminate repetitive manual tasks. Ability to work effectively on complex technical challenges involving reliability, performance, scalability, and operational efficiency. Strong collaboration skills and the ability to partner effectively with engineering teams across infrastructure, hardware, networking, and AI. Experience with GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch, or related technologies is a plus. Experience with InfiniBand or RoCE networking is advantageous. Familiarity with bare-metal provisioning and infrastructure lifecycle management is desirable. Experience supporting large-scale AI training or inference clusters is a plus. Knowledge of hardware health monitoring and predictive failure detection is beneficial. Experience with distributed storage systems is advantageous. Familiarity with AI agents or autonomous infrastructure operations is a plus. Benefits: Opportunity to work on large-scale AI infrastructure supporting advanced training and inference workloads. Exposure to complex systems spanning GPUs, networking, storage, distributed computing, and AI workloads. Opportunity to build highly automated infrastructure systems and developer platforms. Collaboration with multidisciplinary engineering teams working across hardware, networking, platform, and AI. Environment focused on software-driven infrastructure, automation, scalability, reliability, and performance. Opportunity to contribute to systems operating at significant GPU scale. Remote-friendly job listing based in Bangalore, India.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-25. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.