Staff Site Reliability Engineer
Sumo Logic
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
938 open reliability roles across 207 companies are on ApplySarthi right now, most of them in Bengaluru (37), Delhi NCR (13), Pune (10).
- Senior Site Reliability EngineerCamunda
- Senior Site Reliability Engineer - Hybrid CloudGeniussports
- Site Reliability Engineer / SRE (all genders)Lio
- Site Reliability EngineerDeepJudge
- Senior Site Reliability EngineerJobgether
What reliability roles keep asking for: Kubernetes (34%), Python (33%), Observability (31%), AWS (24%), Linux (22%), Terraform (22%), System design (20%), CI/CD (15%) — counted across their open postings here.
Site Reliability Engineer jobs in India · Remote Site Reliability Engineer jobs · AWS jobs · C++ jobs · CI/CD jobs · Go jobs
Sumo Logic has 9 open roles listed here.
- Senior Partner Sales Manager
- Senior Solutions Engineer
- Senior Solutions Engineer
- Senior Sales Engineer
- Staff Software Engineer - Testing & Automationdelhi ncr
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for reliability roles keep coming back to Kubernetes, Python, Observability, AWS. Practise those questions before you sit with Sumo Logic.
Questions you are likely to be asked
- Why do you want to join Sumo Logic?
- What is your experience with AWS? Tell me one thing you learned the hard way.
- How do you decide what to monitor, and what should wake someone up at night?
- How would you cut the cloud bill of a system without hurting it?
- How do you keep secrets and access safe in your infrastructure?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Staff Site Reliability Engineer at Sumo Logic interview free →Title: Staff Site Reliability Engineer, Product Area Focus
Location: Noida/ Bangalore (Hybrid)
Summary of role
Sumo Logic's microservices architecture, hosted on AWS, ingests petabytes of data daily across many geographic regions in support of our planet-scale observability and security products, serving hundreds of millions of queries a day against thousands of petabytes of data. At that scale, every inefficiency — in code, architecture, or infrastructure — compounds into real cost.
This role sits within the Product SRE organization, working alongside your global SRE team on your product area's reliability roadmap — with a mandate that leads with code as much as operations. You'll find where Sumo's systems are spending more compute, storage, or engineering time than necessary, and fix it in the code and architecture, not just the infra config, while also carrying the fuller SRE mandate — reliability, security posture, and improving the day-to-day experience of the engineers within your product area. You'll be part of a team that blends SRE and backend software engineering skillsets, partnering closely with product engineering teams across your product area.
This is an engineering role — the work is about shipping code, architecture, and system-level changes that improve unit economics, not managing cost dashboards, tagging, or reserved-instance/savings-plan purchasing.
What you’ll do
- Continuously discover cost and efficiency opportunities through production profiling, telemetry, cost data, capacity trends, and system-level analysis — across algorithmic inefficiencies, resource-heavy code paths, and architectural decisions — and turn ambiguous problems into prioritized engineering initiatives.
- Apply performance and capacity engineering techniques to understand CPU, memory, storage, network, and I/O behavior under real production workloads, and optimize the resulting resource footprint.
- Write production-grade code to implement the optimizations you identify — JVM/GC tuning, algorithmic and resource-efficiency improvements, re-architecting inefficient services — in systems that process petabytes of data daily.
- Define and track engineering efficiency metrics such as cost per GB ingested, cost per query, cost per event, resource utilization, or cost per customer workload, and translate the work into measurable impact for engineering and business stakeholders.
- Lead complex, cross-team engineering initiatives from problem discovery through design, implementation, rollout, and measurement — influencing teams where you don't have direct ownership — and help establish engineering patterns and practices that make cost and efficiency a continuous part of the development lifecycle.
- Partner with engineering teams in your product area to prioritize changes, and with developer infrastructure and Global SRE to align with the broader reliability roadmap.
- Participate in the SRE responsibilities for different product areas — SLOs, on-call, incident response, and blameless RCA — using those experiences to identify systemic reliability, performance, and efficiency improvements.
What you’ll have
- B.Tech, M.Tech, or equivalent degree in Computer Science or a related discipline.
- 8+ years of industry experience with a demonstrated track record of ownership.
- Strong CS fundamentals — comfortable with algorithmic complexity, data-structure performance characteristics, and system design at scale.
- Ability to author production-ready code in at least one OO/systems language (Java, Scala, Go, C++, or similar) — depth of engineering ability matters more than which language.
- Experience with distributed systems and microservice architectures in production.
- Demonstrated track record of independently identifying ambiguous performance, scalability, or cost problems and driving engineering changes that produced measurable improvements in production.
- Strong ability to reason quantitatively about system behavior, capacity, performance, and cost, and use production data to validate hypotheses and measure outcomes.
- Working fluency with cloud infrastructure (AWS compute, storage, networking) — enough to reason about cost and architectural tradeoffs.
- Comfort moving across the stack, from application code to the infrastructure it runs on, to find root causes of inefficiency.
Nice to have
- JVM tuning and GC optimization experience at scale.
- Exposure to cost-attribution/FinOps practices or tooling.
- Experience with Kubernetes, Terraform, or modern CI/CD tooling.
- Prior SRE experience — on-call, SLOs, incident response.
- Experience with streaming technologies (Kafka, Kafka Streams) or observability/security platforms.
Why this role
- Visibility — Cost/efficiency work maps directly to metrics the business already tracks.
- Scope you define — you identify where the opportunities are rather than executing a fixed backlog.
- Real scale — systems ingesting petabytes of data daily and serving hundreds of millions of queries.
About Us
Sumo Logic, Inc. helps make the digital world secure, fast, and reliable by unifying critical security and operational data through its Intelligent Operations Platform. Built to address the increasing complexity of modern cybersecurity and cloud operations challenges, we empower digital teams to move from reaction to readiness—combining agentic AI-powered SIEM and log analytics into a single platform to detect, investigate, and resolve modern challenges. Customers around the world rely on Sumo Logic for trusted insights to protect against security threats, ensure reliability, and gain powerful insights into their digital environments. For more information, visit www.sumologic.com.
Sumo Logic Privacy Policy. Employees will be responsible for complying with applicable federal privacy laws and regulations, as well as organizational policies related to data protection.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Principal Product ManagerSumo Logic · delhi ncr
- Principal Product ManagerSumo Logic · bengaluru
- Staff Site Reliability Engineer Sumo Logic · bengaluru
- Staff Software Engineer - Testing & AutomationSumo Logic · delhi ncr
- Senior Partner Sales ManagerSumo Logic
- Senior Sales EngineerSumo Logic
- Senior Solutions EngineerSumo Logic
- Senior Solutions EngineerSumo Logic
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on greenhouse · posted 2026-02-10. ApplySarthi collects openings and links to application pages; the role is advertised by Sumo Logic, not by us.