ApplySarthi

GPU Cluster Architect

Jobgether

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

3,422 open architect roles across 326 companies are on ApplySarthi right now, most of them in Bengaluru (221), Hyderabad (100), Delhi NCR (53).

What architect roles keep asking for: AWS (32%), Python (21%), Azure (19%), GCP (13%), Customer success (12%) — counted across their open postings here.

LLMs jobs · Python jobs

Jobgether has 3,737 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for architect roles keep coming back to AWS, Python, Azure, GCP. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with LLMs? Tell me one thing you learned the hard way.
  3. What is a weakness you are working on, and how?
  4. Tell me about yourself, and why this role is the right next step.
  5. Tell me about a problem you solved at work that you are proud of.

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the GPU Cluster Architect at Jobgether interview free →

Accountabilities: Architect scalable GPU cluster topologies encompassing compute nodes, high-performance interconnects, storage systems, and control planes. Define and evaluate infrastructure architectures capable of supporting large-scale AI and machine-learning workloads across multiple data center sites. Model workload requirements for applications such as large language model training and inference, using latency, bandwidth, GPU density, and other performance factors to guide architectural decisions. Design and validate high-throughput, low-latency networking architectures at both POD and data-center scale, including InfiniBand and Ethernet-based environments. Work with network architecture teams to evaluate and validate technologies such as InfiniBand HDR/NDR and RoCEv2. Partner with storage engineering teams to optimize infrastructure for training datasets, checkpointing, and other demanding AI workloads. Analyze monitoring and telemetry signals to identify design issues, reliability risks, and opportunities for architectural improvement. Collaborate closely with site reliability, networking, storage, and data center engineering teams to operationalize, deploy, and scale infrastructure architectures. Contribute to automation and telemetry initiatives that improve the visibility, performance, and reliability of large-scale GPU environments. Make end-to-end architectural decisions that balance scalability, performance, reliability, operational complexity, and infrastructure efficiency. Requirements 5+ years of experience designing and architecting large-scale computing or GPU clusters. Deep understanding of modern GPU architectures, including NVIDIA, AMD, or comparable platforms. Strong experience with high-performance computing interconnects, particularly InfiniBand and RoCE. Solid background in systems architecture, networking, hardware infrastructure, and hardware reliability. Understanding of GPU cluster design principles, including compute topology, network architecture, storage integration, and control-plane considerations. Experience evaluating infrastructure performance and making architecture decisions based on workload characteristics such as latency, bandwidth, and compute density. Experience with scripting or software development for automation, telemetry, monitoring, or infrastructure tooling using languages such as Python or Go. Ability to analyze technical signals and operational data to identify infrastructure issues and inform design improvements. Strong cross-functional collaboration skills, with the ability to work effectively with networking, storage, site reliability, and data center engineering teams. Strong analytical and problem-solving abilities, with a practical approach to complex infrastructure challenges. Ability to work independently in a fast-moving environment while taking ownership of significant architectural decisions. Excellent communication skills and the ability to explain complex technical architectures and tradeoffs to technical stakeholders. Benefits Competitive compensation ranging from $184,000 to $318,000 OTE , including base salary and performance bonus. Equity in the form of RSUs may be available at certain salary grades. 100% company-paid medical, dental, and vision insurance for employees and their families. 401(k) plan with up to a 4% company match and immediate vesting. Paid parental leave: 20 weeks for primary caregivers and 12 weeks for secondary caregivers. Remote work reimbursement of up to $85 per month for mobile and internet expenses. Company-paid short-term disability, long-term disability, and life insurance. Remote work flexibility for U.S.-based employees. Career growth and ongoing learning opportunities. Opportunity to work on large-scale AI and machine-learning infrastructure projects. Collaborative, international environment with experienced engineering and AI professionals. High level of ownership and opportunities to contribute to technically ambitious infrastructure initiatives. Work environment that encourages innovation, initiative, and continuous improvement.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-30. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.