Staff Site Reliability Operations
Nvidia
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
946 open reliability roles across 175 companies are on ApplySarthi right now, most of them in Bengaluru (39), Delhi NCR (12), Pune (9).
- Observability DevOps Engineer - RDT Digital Operations and ReliabilityRoche
- Sr. Site Reliability Engineer (US Federal)Workday
- Intern - NAND Wafer ReliabilityMicron
- Reliability, Maintainability & System Health Systems Engineer (Experienced or Lead)Boeing
- Senior Site Reliability EngineerMastercard
What reliability roles keep asking for: Python (37%), Kubernetes (34%), Observability (34%), AWS (28%), Terraform (23%), Linux (21%), System design (19%), CI/CD (17%) — counted across their open postings here.
Excel jobs · Linux jobs · Python jobs · ServiceNow jobs
Nvidia has 2,295 open roles listed here.
- Senior Compiler Optimization Engineer – LLVMbengaluru
- Engineering Manager - OpenBMC Platform
- Principal Firmware Engineer - Data Center Server Management
- Senior Systems Software Engineer- EDA Infrastructure
- Senior Data Backend Engineer
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for reliability roles keep coming back to Python, Kubernetes, Observability, AWS. Practise those questions before you sit with Nvidia.
Questions you are likely to be asked
- Why do you want to join Nvidia?
- What is your experience with Linux? Tell me one thing you learned the hard way.
- Tell me about an outage you handled. What did you learn from it?
- How do you decide what to monitor, and what should wake someone up at night?
- How would you cut the cloud bill of a system without hurting it?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Staff Site Reliability Operations at Nvidia interview free →For over 25 years, NVIDIA has been at the forefront of transforming computer graphics, PC gaming, and accelerated computing, driven by a legacy of continuous innovation and exceptional talent! We are now bringing to bear the immense potential of AI to usher in the next era of computing, where our GPUs power the "brains" of computers, robots, and autonomous vehicles that can comprehend the world. This pioneering work demands vision, innovation, and the world's best talent. Join our diverse and supportive environment, where NVIDIANs are inspired to excel and make a profound global impact. We are seeking a Site Reliability Operations Technical Lead to serve as the senior technical individual contributor for reliability and support at the site. This role owns technical service delivery locally, acts as the support point of last resort before regional and global platform teams, and provides technical leadership to site support engineers. Be responsible for the hardest issues across Active Directory, Exchange, database platforms, and compute infrastructure, lead the site through major incidents, and drive out the recurring problems that consume the team’s capacity. You will also lead site-level projects, represent local requirements in global initiatives, and set the technical standard the site support team works to. The successful candidate is equally comfortable running a root cause analysis, supporting an executive before an all-hands, and briefing IT leadership on site risk. What you'll be doing: Own day-to-day site operations — incidents, requests, critical issues, and support coverage — with accountability for queue health, SLA attainment, backlog, and service quality, plus site asset and inventory management across lifecycle, refresh, procurement, and compliance. Serve as Tier 3 escalation owner for the site and AMER across identity (AD, hybrid Entra ID, GPO, Kerberos/LDAP, SSO, MFA), messaging (Exchange hybrid mail flow, mailbox, SMTP relay), compute (Windows, Linux, macOS, virtualization, storage, and hands-on datacenter and lab hardware), and endpoint (M365, Teams, Intune, Autopilot, imaging through migrations) driving root cause and permanent fixes rather than repeat break-fix. Own endpoint compliance, vulnerability remediation, patch management, and hardening; audit readiness and evidence; and partnership with InfoSec on incident response and privileged access. Drive critical issues into global platform teams and vendors with reproduction cases and diagnostic evidence through to a committed fix. Act as technical lead for site SRO engineers setting standards, reviewing work, directing blocking issues, building diagnostic rigor through mentorship, and owning the site knowledge base and runbook library. Serve as the primary technical contact for site IT, partnering with employees, site and executive leadership, Facilities, Security, HR, and Procurement on incidents, planned changes, onboarding and moves, and office and lab expansions. Build automation in PowerShell, Python, or Bash for diagnostics, remediation, health checks, and reporting; analyze ticket and reliability trends to eliminate top recurring drivers; and champion AI-driven and agentic solutions that advance SRO strategy. Represent site and AMER priorities in regional and global IT initiatives, standards, and architecture forums, and lead operational decisions in the manager's absence. What we need to see: 8+ years in enterprise support engineering, infrastructure, or end user services, including 5+ years in a senior, lead, or escalation-tier role in a multi-site environment. Deep hands-on solving across Active Directory and hybrid Entra ID, Exchange hybrid, Windows and Linux server, virtualization, enterprise storage, and datacenter hardware. Enterprise endpoint management (Intune, Autopilot, MECM/SCCM, Jamf, or equivalent), Windows 11, and the Microsoft 365 ecosystem, plus endpoint security and vulnerability remediation. Database operations support and networking fundamentals — DNS, DHCP, VLAN, wireless, firewall policy, and switch-level troubleshooting. Scripting and automation in Python, PowerShell, or Bash applied to real support problems, and ServiceNow or similar ITSM. Demonstrated technical leadership without formal authority, excellent executive-level communication during incidents, and the rigor to pursue root cause over symptom clearing. Willingness to work on-site and hands-on (including lifting and moving equipment), join an on-call rotation, and support after-hours maintenance windows and cutovers. Bachelor's degree in Computer Science, Information Systems, or related field, or equivalent experience. Ways to Stand Out from the crowd: Experience supporting engineering, lab, R&D, or manufacturing environments with specialized equipment and non-standard availability requirements. Local technical lead through a site buildout, relocation, or major migration; or experience influencing global standards and tooling roadmaps for site and regional needs. Executive support programs, AV and hybrid conference room technologies, or build automation with measurable efficiency and experience benefits. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 144,000 USD - 230,000 USD. You will also be eligible for equity and benefits . Applications for this job will be accepted at least until October 5, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Head of Startups - India and South AsiaNvidia · bengaluru
- Manager, AI/HPC Infrastructure Technical Delivery — IndiaNvidia · pune
- Architect – AI-Powered Performance Verification AutomationNvidia · bengaluru
- Data Center Infrastructure SpecialistNvidia · bengaluru
- Senior Developer Relations ManagerNvidia · bengaluru
- Server Performance Architect - HardwareNvidia · bengaluru
- Senior Solutions Architect, Infiniband and Networking Ethernet - NVISNvidia · bengaluru
- Senior Software Engineer, Fabric Networking - GPUNvidia · bengaluru
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on workday · posted 2026-10-02. ApplySarthi collects openings and links to application pages; the role is advertised by Nvidia, not by us.