ApplySarthi Match jobs to your CV

Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

Fact Finder

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

938 open reliability roles across 207 companies are on ApplySarthi right now, most of them in Bengaluru (37), Delhi NCR (13), Pune (10).

What reliability roles keep asking for: Kubernetes (34%), Python (33%), Observability (31%), AWS (24%), Linux (22%), Terraform (22%), System design (20%), CI/CD (15%) — counted across their open postings here.

Site Reliability Engineer jobs in Germany · Remote Site Reliability Engineer jobs · Kubernetes jobs · Observability jobs

Fact Finder has 2 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for reliability roles keep coming back to Kubernetes, Python, Observability, AWS. Practise those questions before you sit with Fact Finder.

Questions you are likely to be asked

  1. Why do you want to join Fact Finder?
  2. What is your experience with Kubernetes? Tell me one thing you learned the hard way.
  3. Walk me through how code gets from a commit to production where you work.
  4. Tell me about an outage you handled. What did you learn from it?
  5. How do you decide what to monitor, and what should wake someone up at night?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d) at Fact Finder interview free →

Introduction At a glance Location & work model : Berlin, hybrid Tech stack: Kubernetes on our own servers, Harvester ( KubeVirt ), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph Team: A growing SRE team – you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardware Process: Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team Languages: Fluent English required; German is a plus, not a must Why this role is special Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next. SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time. Your first 90 days You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next. Your mission Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case Eliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hard Plan capacity, performance and cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations Your profile Must-haves: Kubernetes in production – built, not just used : you've set up and maintained clusters on your own servers (e.g. kubeadm , RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experience isn't enough for this role Lived SRE practice : SLOs, error budgets, incident management, on-call Hands-on experience with GitOps or comparable infrastructure/deployment automation – experience with Argo CD or Flux is a strong plus Solid observability skills – metrics, logs, traces, alerting that people trust A strong automation instinct – you'd rather fix a problem's cause than repeat its workaround A collaborative, enabling mindset – you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solution Nice-to-haves (genuinely optional – we'll teach you the rest): Harvester, KubeVirt , vSphere/ ESXi , OpenStack or similar virtualization/HCI platforms Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN) Auto-scaling (HPA, VPA, KEDA, cluster auto scaler ) and capacity/cost planning Experience building Kubernetes operators/CRDs German language skills Certifications (CKA, CKS) are welcome but no substitute for hands-on experience – in the tech interview we'll ask about what you've actually built and operated. You don't tick every box – or your title was never “SRE”? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you – production experience and engineering mindset matter more to us than titles or buzzwords. THE JOY OF WORKING WITH US Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe. Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right. AI-first mindset: We use AI as a real part of our daily work, not as a buzzword. Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role. Flexible work: Hybrid work model three office days per week with a focus on outcomes. Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline. Job Location Berlin, Munich, Pforzheim or Stockholm (all Hybrid) Find more English Speaking Jobs in Germany on Arbeitnow

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on arbeitnow · posted 2026-09-26. ApplySarthi collects openings and links to application pages; the role is advertised by Fact Finder, not by us.