Lead Software Engineer, Cloud Site Reliability (SRE)
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
935 open reliability roles across 208 companies are on ApplySarthi right now, most of them in Bengaluru (37), Delhi NCR (13), Pune (9).
- RME Operator with Admin skills, RME (Reliability Maintenance Engineering) Team in ErfurtAmazon Erfurt GmbH - O80
- Senior Site Reliability EngineerCamunda
- Senior Site Reliability Engineer - Hybrid CloudGeniussports
- Site Reliability Engineer / SRE (all genders)Lio
- Site Reliability EngineerDeepJudge
What reliability roles keep asking for: Kubernetes (34%), Python (34%), Observability (32%), AWS (25%), Linux (22%), Terraform (22%), System design (20%), CI/CD (15%) — counted across their open postings here.
Software Engineer jobs in India · Remote Software Engineer jobs · AWS jobs · Azure jobs · Docker jobs · Kubernetes jobs
Jobgether has 3,942 open roles listed here.
- AI Researcher — Distillation
- AI Researcher — Distillation
- Art Director
- Applied ML Engineer
- AI Science Writer, Nebius Academy (Contract)
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for reliability roles keep coming back to Kubernetes, Python, Observability, AWS. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with Azure? Tell me one thing you learned the hard way.
- How would you cut the cloud bill of a system without hurting it?
- How do you keep secrets and access safe in your infrastructure?
- Walk me through how code gets from a commit to production where you work.
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Lead Software Engineer, Cloud Site Reliability (SRE) at Jobgether interview free →Accountabilities:: Lead 24x7 NOC and site reliability operations through rotational shifts, ensuring system availability, operational stability, and adherence to service-level agreements. Act as Major Incident Manager for P1 and P2 incidents, coordinating triage activities, war rooms, technical teams, and stakeholder communications through resolution. Manage and troubleshoot Azure infrastructure, including virtual machines, networking, storage, and related cloud services. Administer Azure Kubernetes Service (AKS), Kubernetes, and Docker environments, including scaling, troubleshooting, performance optimization, and reliability improvements. Establish and enhance observability practices across logs, metrics, and traces using platforms such as Datadog and Azure Monitor. Drive proactive monitoring, alert optimization, anomaly detection, and AIOps initiatives to identify and address potential reliability issues before they impact users. Build automation and self-healing workflows using Terraform, ARM templates, Helm, Power Automate, PowerShell, Python, Bash, and other appropriate technologies. Collaborate with engineering teams to strengthen deployment pipelines, improve system reliability, and advance cloud-native architecture and operational practices. Develop operational dashboards and reports using Power BI and ServiceNow to provide visibility into incidents, reliability, and service performance. Lead monthly business reviews and provide clear operational reporting and insights to leadership and stakeholders. Mentor team members, promote knowledge sharing, standardize operational processes, and drive continuous improvements across the reliability function. Support broader initiatives involving multi-cloud environments, predictive monitoring, and self-healing systems where appropriate. Requirements 7–12 years of professional experience in CloudOps, Site Reliability Engineering, NOC, or comparable 24x7 operations environments. Strong hands-on expertise with Azure infrastructure, particularly virtual machines, networking, storage, and related IaaS services. Extensive experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker, including troubleshooting, scaling, and performance tuning. Strong experience with monitoring and observability platforms such as Datadog and Azure Monitor, including integrations, alerting, dashboards, and operational analysis. Proven experience in incident management and major incident handling, including P1/P2 coordination, stakeholder communication, root-cause analysis, and operational reporting. Experience with Infrastructure as Code technologies such as Terraform, ARM templates, and Helm. Strong scripting capabilities using PowerShell, Python, Bash, or similar automation technologies. Experience working with ServiceNow, particularly Incident, Problem, and Change Management modules and associated dashboards. Good understanding of distributed systems, cloud-native architectures, and reliability engineering principles. Excellent communication, leadership, coordination, and problem-solving skills, with the ability to work effectively across technical and business teams. Experience working in multi-cloud environments, particularly Azure and AWS, is a plus. Exposure to AIOps, predictive monitoring, anomaly detection, or self-healing systems is desirable. Relevant certifications in Azure, Datadog, Kubernetes, or related cloud and reliability technologies are advantageous. Bachelor’s degree or equivalent technical education and professional experience. Benefits Fully remote opportunity based in India, with a rotational shift structure supporting 24x7 operations. Leadership responsibility across cloud reliability, infrastructure operations, incident management, and operational excellence. Hands-on exposure to Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, and Infrastructure as Code. Opportunity to develop advanced observability, AIOps, predictive monitoring, automation, and self-healing capabilities. Collaboration with engineering teams on cloud-native architecture, deployment pipelines, scalability, and reliability initiatives. Opportunities to mentor team members and influence operational standards and engineering practices. Exposure to multi-cloud technologies and large-scale distributed systems. An inclusive work environment focused on teamwork, openness, respect, fairness, and continuous improvement. Support for professional development through exposure to modern cloud and reliability technologies and relevant certification paths. Opportunities to participate in leadership reporting, business reviews, and cross-functional initiatives with broad organizational visibility.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-24. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.