ApplySarthi

Staff Site Reliability Engineer

Jobgether

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

925 open reliability roles across 171 companies are on ApplySarthi right now, most of them in Bengaluru (40), Delhi NCR (13), Hyderabad (11).

What reliability roles keep asking for: Python (35%), Kubernetes (33%), Observability (31%), AWS (27%), Terraform (22%), Linux (21%), System design (19%), CI/CD (15%) — counted across their open postings here.

Remote Site Reliability Engineer jobs · AWS jobs · Go jobs · Kubernetes jobs · Observability jobs

Jobgether has 4,481 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for reliability roles keep coming back to Python, Kubernetes, Observability, AWS. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with Observability? Tell me one thing you learned the hard way.
  3. How do you keep secrets and access safe in your infrastructure?
  4. Walk me through how code gets from a commit to production where you work.
  5. Tell me about an outage you handled. What did you learn from it?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Staff Site Reliability Engineer at Jobgether interview free →

Accountabilities:: Take full ownership of a core infrastructure product or subsystem end to end, including design, development, deployment, operation, and production performance. Define project goals and success metrics, align technical work with organizational objectives, and proactively identify and mitigate risks. Translate product requirements and technical specifications into practical designs that address critical edge cases without unnecessary complexity. Build secure, reliable, resilient, high-performing, and cost-efficient infrastructure for diverse applications and workloads. Design, develop, and deploy production software and developer-facing tools that improve engineering workflows and reduce operational toil. Manage infrastructure through code and configuration, primarily using Terraform and established architectural patterns. Partner with product engineering teams to design services for scale and resolve ambiguous technical requirements with stakeholders. Participate in incident response and apply systematic debugging techniques to diagnose infrastructure and service issues. Develop and improve monitoring and observability practices, using operational data to identify stability, performance, and reliability improvements. Apply a security-focused mindset across infrastructure development, implementation, and peer reviews by proactively identifying potential vulnerabilities. Serve as a technical resource for complex infrastructure challenges and mentor engineers through code reviews, pairing, and design feedback. Drive collaboration across engineering and other stakeholder groups, facilitating discussions around technical decisions, processes, and infrastructure strategy. Requirements: 6–10 years of experience in infrastructure, platform, or backend engineering, primarily within cloud-based environments; AWS experience is preferred. Proven experience owning significant infrastructure products or subsystems through their full lifecycle, including design, implementation, deployment, and production operations. T-shaped technical expertise, with deep specialization in one or two areas and sufficient breadth to navigate and contribute across wider systems with limited guidance. Strong hands-on experience managing infrastructure through code and configuration using Terraform or an equivalent technology. Deep understanding of cloud infrastructure fundamentals, including networking, load balancing, containerization, Kubernetes/EKS, and distributed systems. Strong programming skills in Go, Python, or a comparable language, with the ability to develop production-ready software. Hands-on production experience operating Redis or ElastiCache, including cluster and shard management, failover behavior, memory eviction policies, and scaling strategies. Experience with observability and monitoring technologies such as Prometheus, Grafana, OpenTelemetry, or similar tools. Strong understanding of performance tuning, incident management, reliability engineering, and production troubleshooting. Fluency in software engineering best practices, including source control, code reviews, comprehensive testing, edge-case handling, and safe deployment practices. High degree of ownership and autonomy, with demonstrated ability to make progress when requirements are ambiguous or not fully defined. Strong written and verbal English communication skills, including the ability to produce clear technical documentation, participate in effective code reviews, and communicate decisions across engineering teams. Ability to collaborate effectively with product engineering, security, DevOps, and other stakeholders in a distributed environment. Benefits: Fully remote work arrangement. Opportunity to work on large-scale cloud infrastructure and systems supporting customer-facing technology products. Significant ownership over infrastructure products and subsystems from design through production operations. Exposure to cloud architecture, distributed systems, Kubernetes, Terraform, observability, reliability engineering, and developer tooling. Opportunity to work with technologies including AWS, Redis/ElastiCache, Prometheus, Grafana, OpenTelemetry, Go, and Python. Strong focus on engineering quality, security, scalability, and operational excellence. Opportunity to mentor engineers and influence technical standards and infrastructure practices. Collaboration with globally distributed engineering, product, security, and DevOps teams. High-autonomy environment suited to engineers who enjoy solving complex and ambiguous technical problems. US-based cash compensation range of $152,000–$205,000; compensation varies by hiring location and this range is not directly applicable to all locations. Visa sponsorship is not provided; candidates must be authorized to work from their home location. Specific India-based salary, healthcare, retirement, paid time off, and other benefits were not specified in the source description.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-29. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.