Director, Site Reliability Engineering
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
936 open reliability roles across 208 companies are on ApplySarthi right now, most of them in Bengaluru (37), Delhi NCR (13), Pune (10).
- Senior Site Reliability EngineerCamunda
- Senior Site Reliability Engineer - Hybrid CloudGeniussports
- Site Reliability Engineer / SRE (all genders)Lio
- Site Reliability EngineerDeepJudge
- Senior Site Reliability EngineerWorkday
What reliability roles keep asking for: Kubernetes (34%), Python (34%), Observability (32%), AWS (25%), Linux (22%), Terraform (22%), System design (20%), CI/CD (15%) — counted across their open postings here.
Docker jobs · Go jobs · Linux jobs · Perl jobs
Jobgether has 3,942 open roles listed here.
- AI Researcher — Distillation
- AI Researcher — Distillation
- Art Director
- Applied ML Engineer
- AI Science Writer, Nebius Academy (Contract)
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for reliability roles keep coming back to Kubernetes, Python, Observability, AWS. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with Docker? Tell me one thing you learned the hard way.
- How do you decide what to monitor, and what should wake someone up at night?
- How would you cut the cloud bill of a system without hurting it?
- How do you keep secrets and access safe in your infrastructure?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Director, Site Reliability Engineering at Jobgether interview free →Accountabilities:: Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems. Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices. Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem. Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation. Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks. Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions. Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability. Leverage cloud-native architectures and services to strengthen system resilience and support continued growth. Help ensure products and infrastructure meet established reliability standards while minimizing user impact during failures and incidents. Identify emerging technical needs and opportunities to guide the long-term evolution of deployment and infrastructure architecture. Support a culture of ownership, continuous improvement, measurable outcomes, and effective post-incident learning. Requirements: 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields. 4+ years of experience leading SRE or comparable engineering teams. Experience participating in or managing 24/7 on-call operations for large-scale production environments. Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems. Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments. Demonstrated ability to lead complex technical projects from ambiguous initial requirements through execution and postmortem. Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities. Strong investigative and root-cause analysis skills, particularly within distributed and high-scale systems. Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions. Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose. Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies. Experience with AI-driven software development, including the design and implementation of agentic workflows. Strong ability to turn ambiguous or complex problems into practical, innovative solutions with measurable outcomes. Strategic thinking and technical foresight, with the ability to anticipate future infrastructure and reliability requirements. Excellent communication and collaboration skills, with the ability to work effectively across engineering teams and technical disciplines. Strong sense of ownership, autonomy, and accountability in a remote-first working environment. Benefits: Annual compensation of $243,800 USD , plus stock options. Transparent compensation structure, with team members at the same professional level and within the same global region receiving the same compensation regardless of functional team, location, gender, educational background, or years of experience. Fully remote, flexible working arrangement with no core working hours. Average full-time commitment of approximately 40 hours per week. Company-sponsored health benefits for eligible team members based in the United States; these benefits do not extend to team members based in Canada or other countries. Paid parental leave. Support for home-office setup. Co-working allowances. Opportunities to participate in company-wide and team gatherings, with travel expected at least twice per year for an all-hands meeting and a team retreat. Remote-first environment centered on trust, inclusivity, ownership, and empowered project management. Equal employment opportunities and a commitment to an accessible, inclusive hiring process. Reasonable accommodations are available for candidates who require support during the application process. Successful candidates must complete a background check as a condition of employment. The role requires participation in video meetings with cameras enabled.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-23. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.