Infrastructure Site Reliability Engineer
Radiant
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
2,026 open infrastructure roles across 290 companies are on ApplySarthi right now, most of them in Bengaluru (66), Hyderabad (28), Delhi NCR (12).
- Lead Infrastructure EngineerJPMorgan
- Member of technical staff (Infrastructure) - ParisHcompany
- Platforms and Infrastructure EngineerEnactintelligence
- Engineering Manager - InfrastructureCamunda
- Senior Backend Engineer: Machine Learning InfrastructureConstructor
What infrastructure roles keep asking for: AWS (30%), System design (22%), Kubernetes (21%), Python (21%), Observability (19%), Terraform (15%), CI/CD (13%), Linux (12%) — counted across their open postings here.
Ansible jobs · Kubernetes jobs · Linux jobs · Observability jobs
Radiant has 29 open roles listed here.
- GRC Analyst
- Group Reporting Manager
- Senior Accountant
- Accountant
- Principal Procurement Lead, Compute and Storage
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for infrastructure roles keep coming back to AWS, System design, Kubernetes, Python. Practise those questions before you sit with Radiant.
Questions you are likely to be asked
- Why do you want to join Radiant?
- What is your experience with Observability? Tell me one thing you learned the hard way.
- How do you decide what to monitor, and what should wake someone up at night?
- How would you cut the cloud bill of a system without hurting it?
- How do you keep secrets and access safe in your infrastructure?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Infrastructure Site Reliability Engineer at Radiant interview free →About Radiant Radiant is redefining how AI infrastructure is built. We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure powers GPU-native workloads, multi-tenant control planes, and high-performance AI systems designed for the most demanding environments. We are not building a generic cloud. We are building purpose-built AI infrastructure - from powered land, to compute, to software . As we scale our platform and expand our engineering organisation, we are looking for leaders who can build strong teams, uphold high standards, and deliver reliably at pace. Job Summary: We’re looking for an experienced Infrastructure Site Reliability Engineer to run and evolve our infrastructure stack. You’ll contribute across bare-metal, virtualization, and orchestration layers, keeping things stable and secure 24/7 x 365 — all while mentoring teammates, improving process and automation as well as helping translate deep technical concepts for a wide range of collaborators and customers. What You’ll Do : Deploy and operate resilient, scalable infrastructure supporting AI/HPC workloads Optimize Linux system configuration, BIOS/firmware, kernel, and disk subsystem for performance Configure, monitor and manage bare-metal infrastructure using IPMI, Redfish, etc Build and maintain automation scripts and infrastructure as code to support platform lifecycle, as well as simplifying troubleshooting for Incident resolution and provision of tooling for our support organisation Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. Maintain and enhance ’s observability stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis, document learnings, and automate remediations Mentor junior engineers and act as an Operational requirements consultant to other departments Communicate technical decisions clearly to non-technical stakeholders and customers Uphold a culture of: do, document, automate Willingness to cross train with Platform Engineering/Platform SRE to fully support both our infrastructure and platform stacks. Willingness to cross train with HPC Engineering, supported by NVIDIA to enhance our HPC supportability offering What you bring: 5+ Years Proven experience in globally scaled, performance-intensive environments operating to a 24/7 support model Expert-level Linux administration, especially Ubuntu distributions Proficiency in system tuning, disk I/O optimization, and hardware-level performance tweaks Familiarity with Out of Band management tools (IPMI, Redfish, PXE, etc.) Strong networking fundamentals: TCP/IP, DNS, DHCP, VLANs, routing, switching Strong experience with infrastructure scripting and automation (Bash, Python, Ansible) Deep understanding of observability principles and tools (Prometheus, Grafana) Hands-on experience operating orchestration platforms (Kubernetes, MAAS, Tinkerbell) Strong grasp of ITSM and service operation best practices Excellent communication and mentorship skills Comfortable interfacing with internal stakeholders and external customers Bonus: Knowledge of HPC workloads and GPU-based infrastructure Bonus: Experience with InfiniBand networks and HPC performance tuning Nice to have: Bachelor or Masters Level degree in Computer Science, Engineering or related field, or equivalent experience. LPIC Certifications ITIL Foundation level qualification or equivalent experience How you work: You approach problems with a systems mindset - balancing practical execution with long-term scalability You elevate the team, setting high standards for technical quality and engineering excellence. You hold yourself and others accountable - giving direct feedback and expecting the same You take initiative, owning challenges end-to-end and proactively driving solutions. You invest in others, mentoring to build both capability and confidence. You communicate clearly - translating complexity into clarity across engineering and business audiences Why should you join us? What sets us apart is our blend of modern technology, competitive benefits, and an open, welcoming work culture that enables our people to thrive. Here are just some of the great things you can expect from us: 25 days of annual leave A culture that emphasises results over hierarchy, process & ego: we place great emphasis on the quality, ingenuity and creativity of work. Open communication, regular feedback: we value smooth collaboration, direct and actionable feedback, and believe that leading with empathy and a growth mindset makes us better together. Learning Time : we all have dedicated learning time to focus on new skills, projects or interests that lay outside of your day-to-day job. Health & Wellbeing: we want everyone to feel healthy and happy, so we offer private medical insurance via Bupa. Cycle to Work Scheme: we're committed to building a sustainable business, so we encourage cycling to work. Gympass subscription to a variety of gyms and wellbeing apps Participation in the company shares program Enhanced parental pay & leave Diversity, Equality, Inclusion and Belonging We are an equal opportunity employer and we strive to reduce unconscious bias throughout our hiring process. All applicants will be considered for employment without attention to ethnicity, religion, sexual orientation, gender identity, family or parental status, national origin, veteran, neurodiversity status or disability status. To ensure our recruitment processes provide an equal opportunity for all applicants to succeed, we encourage you to let us know if there are any adjustments that we can make.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Senior Network EngineerRadiant
- Senior Backend Software Engineer - Core ServicesRadiant
- Platform Site Reliability EngineerRadiant
- Cloud Infrastructure Support EngineerRadiant
- HPC Infrastructure Site Reliability EngineerRadiant
- Infrastructure Tooling & Observability Engineer( UK)Radiant
- Cluster ArchitectRadiant
- Director, Technical Product MarketingRadiant
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on ashby · posted 2026-04-07. ApplySarthi collects openings and links to application pages; the role is advertised by Radiant, not by us.