Orchestration Workload Engineer - ACE - AI Factory
Roche
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
32 open orchestration roles across 24 companies are on ApplySarthi right now, most of them in Bengaluru (3), Pune (1), Hyderabad (1).
- Associate Director, Orchestration EngineAmgen
- Principal AI Engineer Optimisation Intelligence and Agent OrchestrationMaersk · bengaluru
- Senior Software Engineer - OrchestrationFivetran
- Software Dev Engineer II, Control Plane & Orchestration, Prime VideoAmazon.com Services LLC
- Director of Sales, AI and Experience OrchestrationJobgether
What orchestration roles keep asking for: Java (28%), Observability (28%), Python (28%), AWS (22%), C++ (22%), Kubernetes (22%), Azure (19%), GCP (19%) — counted across their open postings here.
Accounting jobs · Ansible jobs · Docker jobs · Kubernetes jobs
Roche has 1,160 open roles listed here.
- ERP Solution Consultant - Clinical Supplyhyderabad
- APAC RCSC Application Specialist
- APAC RCSC Core Lab Team Manager
- APAC RCSC Engineering Specialist
- Technical Innovation Lead
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for orchestration roles keep coming back to Java, Observability, Python, AWS. Practise those questions before you sit with Roche.
Questions you are likely to be asked
- Why do you want to join Roche?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Orchestration Workload Engineer - ACE - AI Factory at Roche interview free →At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters. The Position As a Workload Orchestration Engineer within the Accelerated Compute Engineering (ACE) team, you will be recognised internally as an expert in workload orchestration, owning and advancing our scheduler tech stack across our High-Performance Computing (HPC) platforms. With the rapid expansion of our compute infrastructure, your broad expertise will drive the efficient scheduling, policy management, and resource optimization of our multi-node CPU and GPU environments. In this role, you will use your expertise to bridge traditional scientific computing with modern AI paradigms, while acting as a coach and mentor to help colleagues develop technical expertise. You will solve unique, unprecedented scheduling and infrastructure challenges that directly impact Roche’s compute architecture, ensuring our researchers, data scientists, and engineers can execute compute workloads reliably, efficiently, and successfully. Hosting and Infrastructure (HI) provides mission-critical on-premises infrastructure, cloud hosting, connectivity, and technology products that enable all functions at every Roche site to develop, innovate, connect, and deliver compliant digital products across the Roche Enterprise. The Value Streams - Accelerated Compute Engineering (ACE) Team acts as a center of excellence and delivery for High Performance Compute and AI Infrastructure across Roche. This team facilitates seamless onboarding and adoption for business vertical customers needing accelerated compute—helping infrastructure consumers optimize for high availability, seamless data transfer, flexibility, speed, and the rapidly changing needs of AI to achieve rapid time-to-value. The Opportunity: SLURM Architecture & Ecosystem Leadership Serve as the internal expert on the SLURM Workload Manager, architecting, scaling, and maintaining SLURM across heterogeneous HPC (and AI environments) to ensure high availability and dynamic resource distribution. Design and tune advanced SLURM configurations, including custom plugin integration, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and complex QoS/fair-share policies. Bridge HPC and cloud-native ecosystems by evaluating and implementing integrations between SLURM, Kubernetes, and orchestration platforms (e.g., SLURM Slinky or Run:ai) to streamline job submission workflows across architectures. Hybrid Workload & Kubernetes Integration Integrate containerization standards across SLURM (using Singularity/Apptainer) while maintaining operational familiarity with Kubernetes container orchestration to support hybrid AI/HPC workloads. Solve unique, unprecedented multi-tenant bottlenecks, such as GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures. Technical Leadership, Mentorship & Governance Lead large, global cross-functional initiatives across ACE, infrastructure, platform, scientific computing, and AI teams to establish workload orchestration standards, policies, and architectural patterns across Roche compute environments. Act as a technical mentor and coach for junior and mid-level engineers, driving skill development and continuous learning across the chapter. Partner with Observability Engineers to establish deep telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization, utilizing configuration-as-code to deploy policies uniformly. Who You Are: Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline. Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization. Demonstrated track record of leading complex technical initiatives and mentoring engineering peers. Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments. SLURM Architecture & Optimization : Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources. SLURM Operations, Accounting & Observability : Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions. Kubernetes & Container Knowledge: Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context. AI Infrastructure & Interconnects: Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL). Automation: Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines. Broad Platform Expertise : Apply broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management to solve orchestration problems spanning multiple technology domains. Domain Expertise & Problem Solving: Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads. Coaching & Collaboration: Strong leadership presence with a dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams. Cross-Organizational Coordination: Collaborative team player with demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders. Strategic Vision: Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows. #RDT2026 Where pay transparency applies, details are provided based on the primary posting location. For this role, the primary location is Kaiseraugst. If you are interested in additional locations where the role may be available, we will provide the relevant compensation details later in the hiring process. Who we are A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact. Let’s build a healthier future, together. Roche is an Equal Opportunity Employer.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Clinical Site Manager for Near Patient CareRoche
- Patient Journey PartnerRoche · mumbai
- 2027 Summer Intern - Quality, Regulatory, and External Affairs (QR&E)Roche
- Global Industrial Hygiene and Biosafety ExpertRoche
- Regulatory AI & Automation SpecialistRoche · hyderabad
- Executive AssistantRoche · pune
- Customer Success Specialist (Pathology & Sequencing Labs)Roche · delhi ncr
- Zonal Manager - Corelab (Bangalore)Roche · bengaluru
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on workday · posted 2026-09-21. ApplySarthi collects openings and links to application pages; the role is advertised by Roche, not by us.