Software Engineer, Inference
Lumaai
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
149 open inference roles across 44 companies are on ApplySarthi right now, most of them in Delhi NCR (2), Bengaluru (2).
- (Senior) Sales Manager DACH - AI & InferenceImpossiblecloud
- Senior Deep Learning Software Engineer, Inference and Model OptimizationNvidia
- AI Engineer 5 (FM Hosting, LLM Inference)Capitalone
- Forward Deployed Engineer – AI Inference (Intern)Lyceum
- Staff Software Engineer, Console GPU Inferencesonyinteractiveentertainmentglobal
What inference roles keep asking for: LLMs (54%), Python (52%), Machine learning (44%), PyTorch (30%), Kubernetes (28%), System design (28%), C++ (26%), Observability (26%) — counted across their open postings here.
Software Engineer jobs in the United Kingdom · Software Engineer jobs in London · Remote Software Engineer jobs · CI/CD jobs · Docker jobs · Hugging Face jobs · Kubernetes jobs
Lumaai has 7 open roles listed here.
- Research Scientist / Engineer – Reinforcement Learning Infrastructure
- Copy of Research Scientist / Engineer – Performance Optimization
- Copy of Research Scientist / Engineer – Training Infrastructure
- Director of Customer Success [EMEA]
- Forward Deployed Creative [UK, French Speaking]
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Lumaai.
Questions you are likely to be asked
- Why do you want to join Lumaai?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- How would you design an API for a feature you have worked on?
- What do you do when a production issue happens on your code?
- Walk me through a system you built. How was it designed, and what would you change now?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Software Engineer, Inference at Lumaai interview free →You'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs. This is large-scale inference systems work: scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. It fits a strong systems engineer comfortable with model serving and Kubernetes at scale. If you want pure modeling rather than the systems that run models, this is firmly the systems side. What You'll Own Ship new model architectures by integrating them into the inference engine. Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments. Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows. Automate, test, and maintain inference services for maximum uptime and reliability. Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines. Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs. First 90 Days One way the first 90 could unfold. Days 1–30 — Immerse & Diagnose: Learn the inference stack, the fleets, and where reliability or utilization break. Days 30–60 — Ship & Validate: Integrate a model or ship tooling/scheduling that improves uptime or GPU utilization. Days 60–90 — Scale & Systemize: Harden deployment pipelines and scheduling across clusters and providers. What You Bring Strong Python and system-architecture skills. Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar. Experience with queues, scheduling, traffic control, and fleet management at scale. Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling. Familiarity with Redis and S3-compatible storage. Nice to Have Modern networking stacks including RDMA (RoCE, InfiniBand, NVLink). High-performance large-scale ML systems (100+ GPUs). CUDA, and FFmpeg or multimedia processing. About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer. Find more English Speaking Jobs in United Kingdom on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Senior Customer Success Manager (German speaking)Lumaai
- Forward Deployed Creative [UK, French Speaking]Lumaai
- Copy of Research Scientist / Engineer – Performance OptimizationLumaai
- Copy of Research Scientist / Engineer – Training InfrastructureLumaai
- Director of Customer Success [EMEA]Lumaai
- Research Scientist / Engineer – Reinforcement Learning InfrastructureLumaai
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-10-10. ApplySarthi collects openings and links to application pages; the role is advertised by Lumaai, not by us.