Senior Site Reliability Engineer, Production Engineering
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
925 open reliability roles across 208 companies are on ApplySarthi right now, most of them in Bengaluru (37), Delhi NCR (13), Pune (9).
- RME Operator with Admin skills, RME (Reliability Maintenance Engineering) Team in ErfurtAmazon Erfurt GmbH - O80
- Senior Site Reliability EngineerCamunda
- Senior Site Reliability Engineer - Hybrid CloudGeniussports
- Site Reliability Engineer / SRE (all genders)Lio
- Site Reliability EngineerDeepJudge
What reliability roles keep asking for: Python (34%), Kubernetes (33%), Observability (31%), AWS (25%), Terraform (22%), Linux (21%), System design (19%), CI/CD (16%) — counted across their open postings here.
Site Reliability Engineer jobs in India · Remote Site Reliability Engineer jobs · CI/CD jobs · Jenkins jobs · Kubernetes jobs · Linux jobs
Jobgether has 3,838 open roles listed here.
- Analista de Gestão de Mudanças
- Analista de Governança e Transparência ESG III - Temporária
- Associate Product Manager, BMO Global Asset Management
- Bilingual Field technology Consultant
- Bilingual Vocational Rehabilitation Specialist
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for reliability roles keep coming back to Python, Kubernetes, Observability, AWS. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- How do you keep secrets and access safe in your infrastructure?
- Walk me through how code gets from a commit to production where you work.
- Tell me about an outage you handled. What did you learn from it?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Site Reliability Engineer, Production Engineering at Jobgether interview free →Accountabilities:: Support production Kubernetes services as part of a global 24/7 production engineering operation, including flexibility to work split-weekend shifts. Administer and maintain large-scale Kubernetes clusters, systems, and infrastructure while protecting service availability, integrity, reliability, and SLAs. Automate operational processes and continuously identify opportunities to reduce manual tasks and improve engineering efficiency. Use monitoring, observability, alerts, and alarms to proactively detect, prevent, investigate, and respond to production incidents. Analyze logs, metrics, system behavior, and infrastructure signals to troubleshoot complex issues and determine root causes. Lead incident management calls, coordinating timely detection, escalation, investigation, and resolution of critical production issues. Engage subject matter experts, service owners, and cross-functional engineering teams to resolve complex incidents efficiently. Develop and improve monitoring, alerting, and reliability mechanisms in collaboration with development teams. Perform systems administration and security monitoring across large-scale infrastructure environments. Apply deep knowledge of Linux, networking, Kubernetes, and cluster infrastructure to maintain reliable production services. Contribute to the architecture, deployment, and ongoing improvement of Kubernetes environments operating at significant scale. Continuously evaluate emerging infrastructure and high-performance computing technologies and identify opportunities for innovation. Requirements: 7+ years of demonstrated experience administering large-scale production Kubernetes environments within high-availability Internet, cloud, or data-center environments, with strong on-premises experience preferred. Bachelor's degree in Computer Science, Engineering, Mathematics, or a related discipline, or equivalent professional experience. Advanced hands-on expertise with Kubernetes, SLURM, and large-scale cluster management. Familiarity with GPU/DPU hardware and high-performance computing cluster environments. Strong Linux systems administration experience, including DNS, DHCP, IP tables, routing, firewalls, and core Linux networking. Proven ability to troubleshoot and maintain services across large-scale bare-metal infrastructure. Experience with CI/CD technologies and tools such as Jenkins and ArgoCD. Scripting or programming experience in Python, Golang, or Rust is preferred but not mandatory. Strong understanding of observability, incident management, reliability engineering, and production operations. Excellent analytical and troubleshooting skills, with the ability to work effectively under pressure during complex incidents. Strong communication and interpersonal skills, including the ability to clearly present technical information and influence cross-functional stakeholders. Ability to learn new technologies quickly and adapt to evolving infrastructure environments. Experience architecting, building, and deploying Kubernetes environments at large scale is highly valuable. Passion for innovation and advanced high-performance cluster technologies is an advantage. Benefits: Full-time opportunity with a remote working option in India. Opportunity to work on large-scale production Kubernetes and infrastructure environments. Exposure to advanced cloud, bare-metal, GPU/DPU, and high-performance computing technologies. Opportunity to work alongside SRE, DevOps, security, development, and other specialized engineering teams. Significant technical ownership across reliability, automation, observability, incident response, and infrastructure operations. Opportunity to solve complex engineering challenges at global scale. Continuous exposure to emerging technologies and opportunities to develop advanced infrastructure expertise. 24/7 production engineering environment offering substantial experience in incident management and high-availability operations. Compensation, healthcare, leave, and other employment benefits are provided according to the applicable employment package and location-specific terms.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-21. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.