Senior Site Reliability Engineer
Platform.sh
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
940 open reliability roles across 173 companies are on ApplySarthi right now, most of them in Bengaluru (39), Delhi NCR (12), Pune (8).
- Application Consultant-Site ReliabilityIBM
- Staff Site Reliability EngineerVeeam Software
- Staff Site Reliability EngineerSentinelOne
- Staff Site Reliability EngineerOkta · bengaluru
- Hardware Reliability Engineer (Build Reliability)SpaceX
What reliability roles keep asking for: Python (36%), Kubernetes (34%), Observability (33%), AWS (27%), Terraform (22%), Linux (21%), System design (19%), CI/CD (16%) — counted across their open postings here.
Remote Site Reliability Engineer jobs · AWS jobs · Ansible jobs · Azure jobs · CI/CD jobs
Platform.sh has 10 open roles listed here.
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for reliability roles keep coming back to Python, Kubernetes, Observability, AWS. Practise those questions before you sit with Platform.sh.
Questions you are likely to be asked
- Why do you want to join Platform.sh?
- What is your experience with Observability? Tell me one thing you learned the hard way.
- Walk me through how code gets from a commit to production where you work.
- Tell me about an outage you handled. What did you learn from it?
- How do you decide what to monitor, and what should wake someone up at night?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Site Reliability Engineer at Platform.sh interview free →Upsun is the software factory for AI-human workflows. It is built for today’s hybrid teams, where AI agents write and test code and humans focus on solving the problems that really matter. Developers, DevOps engineers, and platform teams use Upsun to build, ship, and scale confidently without wrestling with backend infrastructure. We give you your time back. You get:
- Predictable performance, even at scale
- Secure, compliant environments by default
- Real-time observability and profiling built in
- Cloning, configuration, and provisioning in seconds
- AI-ready features that plug directly into your stack
The name says it all. "Up" means uptime, reliability, and acceleration. "Sun" reflects our follow-the-sun-support, a 24x7, globally distributed support team keeping the lights on while you rest. Our core belief is that software should power brighter solutions and greater innovation.
Upsunners are a remote, global workforce, and we thrive in a multicultural team. We are committed to open source and an open, welcoming environment. Our team spans the globe and the experience spectrum.
What's our commonality, our cultural fabric? A curious spirit and a thirst for knowledge; an eagerness for innovative ideas and cultures. We believe we can build anything together in an environment that frees you to do your best work.
Our values:
🌿 We make a positive impact.
✨ We aim for the stars.
💚 We care for each other.
Impact of a Senior Site Reliability EngineerAs a Senior Site Reliability Engineer at Upsun, you will lead the evolution of our cloud application platform from traditional cloud operations into a proactive, automation-driven SRE model. You will own critical engineering workstreams that enhance system reliability, scalability, and operational efficiency across multi-cloud environments. Partnering closely with engineering, product, and platform teams, you will embed reliability and performance into every stage of the software delivery lifecycle. In this role, you will anticipate architectural bottlenecks, drive infrastructure-as-code practices, and establish robust observability standards that ensure long-term system stability and uptime for our global users.
-
Drive reliability & observability strategy: Architect and elevate system monitoring, alerting, and logging using Prometheus, Grafana, and ELK Stack, establishing actionable SLIs/SLOs aligned with core business metrics.
-
Automate infrastructure & workflows: Eliminate operational toil by designing and implementing resilient, automated solutions using IaC tools like Terraform and Ansible across AWS, GCP, and Azure.
-
Scale CI/CD & delivery pipelines: Optimize pipeline architectures for fast, secure, and zero-downtime releases, ensuring infrastructure resilience during high-volume deployment cycles.
-
Lead incident response & post-mortems: Guide high-priority incident triage, drive blameless post-mortem analysis, and implement preventative measures to continuously improve system resiliency.
-
Cross-functional leadership: Partner with product and software engineering teams to incorporate SRE best practices into product roadmaps.
-
Champion technical innovation: Proactively identify performance bottlenecks and evaluate emerging technologies (e.g., eBPF, container orchestration) to optimize platform stability and performance.
-
Time distribution: Follow a 4-week rotation balancing engineering and operations to focus on reliability, automation, and scalability through hands-on troubleshooting and engineering innovation.
-
Senior SRE & Cloud Expertise: 5+ years of experience in Site Reliability Engineering, Cloud Operations, or DevOps, with proven experience owning reliability for production platforms at scale.
-
Software Engineering & Tooling: Strong proficiency in Go or Python to build custom automation tools, custom controllers, or SRE platform components (beyond basic shell scripting).
-
Deep Linux Internals Proficiency: Advanced hands-on knowledge of Linux operating system internals, kernel parameters, networking protocols, performance profiling, and system troubleshooting.
-
Infrastructure as Code & Cloud Platforms: Deep expertise with cloud providers (AWS, GCP, Azure, or Openstack) with custom tooling built around cloud SDKs, and declarative infrastructure tools (e.g., Terraform) to manage distributed systems.
-
Autonomous Ownership & Systems Thinking: Proven ability to anticipate operational risks, make architectural trade-offs, and lead technical infrastructure initiatives with minimal guidance.
-
Collaborative Communication: Outstanding cross-functional communication skills with a track record of building alignment, and fostering an inclusive engineering culture.
Bonus
-
Experience with custom-built orchestration, edge, storage, and operational tooling in a dynamic environment.
-
Experience with Docker and production Kubernetes cluster management or containerized deployment architectures.
-
Familiarity with Platform-as-a-Service (PaaS) architectures or developer-facing cloud platforms.
At Upsun, remote work isn't just a trend - it's our foundation. The freedom of remote work with the support of a diverse, global team has been our successful model for over a decade. Our culture celebrates flexibility and collaboration, and while we have team members in over 30 countries around the globe, we are currently focused on hiring for this role in Western Australia. Although we’re unable to provide visa sponsorship at this time, we welcome applications from all qualified candidates who are legally authorized to work in these countries.
This role includes on-call hours: One week every 4-5 weeks, 02:00 AM – 10:00 AM UTC (or 02:00 – 10:00 UTC). Weekend shift included in the one-week of on-call.
We know that a great hire won’t meet every requirement that we’ve outlined. If you can see yourself elevating the team, we want to hear your story. Few of us would be here had we not taken a chance.
You can expect 4 interviews on Google Meet to follow the order below. Should you successfully move through the entire process you will have the opportunity to meet with a variety of Upsunners. Our goal is to ensure you can make the most informed decision on whether this role, and our culture aligns with what you’re looking for in your future working environment.
- 45 Minutes with Talent Acquisition
- 60 Minutes with Hiring Manager
- 60 Minutes with Team (ICs)
- 60 Minutes with Senior Director, SRE
All roles require background checks.
What we offer💡 A product you can believe in - Join us in transforming how businesses build and manage web applications, driven making a positive impact as a proud B Corp.
🏆 An Award-Winning Workplace - We’ve been recognized by Forbes’ Top 30 Companies for Remote Jobs and France’s Best Workplaces for Women.
🗣️ A culture that values your voice - Join a flexible, open, and inclusive work environment where your voice is encouraged, and your ideas shape our growth and evolution.
🌎 A global team - Collaborate with colleagues from diverse backgrounds across the world, embracing different perspectives
🎉 Benefits and perks - Make the most of what matters to you
You belong here🏝 Flexible PTO
📈 Company stock options
🧠 Professional development budget
💻 Office equipment budget
💆♀️ Wellness budget
🧳 Annual team gatherings
🛜 Internet reimbursement
👶 Inclusive parental leave
✈️ Remote work travel program
At Upsun, we celebrate diversity in all its forms and are committed to fostering an inclusive, equitable, and supportive workplace where everyone can thrive. We embrace and value different perspectives, backgrounds, and experiences, because they make us stronger as a team. Whoever you are, wherever you're from, and whatever path you've taken, you are welcome here. We encourage you to bring your whole self to work, connect with others, and share your passion. If you need accommodations at any stage of our hiring process, please let us know. We're here to ensure an accessible and comfortable experience for you.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Billing Systems Architect Platform.sh
- Customer Retention ManagerPlatform.sh
- Senior Analytics EngineerPlatform.sh
- Senior Manager, SupportPlatform.sh
- Senior Solutions ArchitectPlatform.sh
- Product ManagerPlatform.sh
- Senior Software EngineerPlatform.sh
- Software EngineerPlatform.sh
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on greenhouse · posted 2026-09-30. ApplySarthi collects openings and links to application pages; the role is advertised by Platform.sh, not by us.