Research Scientist / Engineer – Reinforcement Learning Infrastructure
Lumaai
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
2,149 open infrastructure roles across 301 companies are on ApplySarthi right now, most of them in Bengaluru (70), Hyderabad (24), Delhi NCR (13).
- Principal Core Infrastructure EngineerOracle
- Infrastructure EngineerTailscale
- Infrastructure EngineerElevenlabs
- Senior Cloud Infrastructure EngineerApplied
- Senior Cloud & Infrastructure Engineer - Project Delivery & Implementation (m/f/d)Cuculus Gmbh
What infrastructure roles keep asking for: AWS (32%), Python (27%), Kubernetes (22%), Observability (22%), System design (21%), CI/CD (17%), Linux (17%), Terraform (16%) — counted across their open postings here.
Research Scientist jobs in the United Kingdom · Remote Research Scientist jobs · Kubernetes jobs · LLMs jobs · PyTorch jobs
Lumaai has 7 open roles listed here.
- Software Engineer, Inference
- Copy of Research Scientist / Engineer – Performance Optimization
- Copy of Research Scientist / Engineer – Training Infrastructure
- Director of Customer Success [EMEA]
- Forward Deployed Creative [UK, French Speaking]
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for infrastructure roles keep coming back to AWS, Python, Kubernetes, Observability. Practise those questions before you sit with Lumaai.
Questions you are likely to be asked
- Why do you want to join Lumaai?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- Walk me through how code gets from a commit to production where you work.
- Tell me about an outage you handled. What did you learn from it?
- How do you decide what to monitor, and what should wake someone up at night?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Research Scientist / Engineer – Reinforcement Learning Infrastructure at Lumaai interview free →You'll build the systems that make reinforcement learning work at frontier scale — coupling policy optimization with large fleets of inference workers, agentic environments, and the reward and verification systems that turn model behavior into learning signal. RL is how Luma's models go from capable to useful. RL at scale is a full-loop systems problem: training, rollout generation, environment execution, and reward computation running concurrently across thousands of GPUs, all needing to stay fast, stable, and correct together. It fits someone who has lived this — post-trained LLMs with RL, built environments and verifiers, and debugged asynchronous rollout pipelines at scale. If you haven't operated RL at real scale, this will be deep water. What You'll Own Design, build, and scale distributed RL post-training systems, orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs. Build high-throughput rollout generation, integrating inference engines (vLLM, SGLang), weight synchronization, and asynchronous/off-policy schemes. Design RL environments for agentic, multi-step tasks — sandboxed code execution, tool use, computer use, multimodal interaction — reproducible and scalable to millions of episodes. Build reward infrastructure: verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking. Develop the evaluation, monitoring, and debugging tooling that keeps large RL runs stable. Advance training efficiency and stability, and turn new post-training ideas into production runs with researchers. First 90 Days One way the first 90 could unfold. Days 1–30 — Immerse & Diagnose: Learn the current RL stack and where throughput, stability, or correctness break. Days 30–60 — Ship & Validate: Improve a piece of the loop (rollout throughput, reward infra, or an environment) and prove it on a real run. Days 60–90 — Scale & Systemize: Harden the full loop across thousands of GPUs and asynchronous architectures. What You Bring Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale. Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models. Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use. Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang). Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads. Nice to Have Running RL training across 100+ GPUs, including asynchronous or disaggregated trainer/rollout architectures. Containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads. Research contributions in RL for LLMs, or open-source contributions to RL training frameworks. About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer. Find more English Speaking Jobs in United Kingdom on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Senior Customer Success Manager (German speaking)Lumaai
- Forward Deployed Creative [UK, French Speaking]Lumaai
- Copy of Research Scientist / Engineer – Performance OptimizationLumaai
- Copy of Research Scientist / Engineer – Training InfrastructureLumaai
- Director of Customer Success [EMEA]Lumaai
- Software Engineer, InferenceLumaai
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-10-10. ApplySarthi collects openings and links to application pages; the role is advertised by Lumaai, not by us.