Performance Engineer, Inference
Sarvam
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
612 open performance roles across 173 companies are on ApplySarthi right now, most of them in Bengaluru (47), Hyderabad (22), Pune (8).
- Associate Director Capital Reporting and PerformanceGsk · bengaluru
- Power & Performance EngineerIntel
- Acquisition Lead, Performance MarketingZefir
- Director of Performance Marketing (m/w/d)Pammys™ (dieseo GmbH)
- Product Manager, Performance AIWpp
What performance roles keep asking for: C++ (15%), Python (14%) — counted across their open postings here.
Sarvam has 61 open roles listed here.
- Growth Marketing Internbengaluru
- Marketing Internbengaluru
- Staff Software Engineer - Frontendbengaluru
- GTM Strategy - Chanakyadelhi ncr
- DevOps Engineerbengaluru
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for performance roles keep coming back to C++, Python. Practise those questions before you sit with Sarvam.
Questions you are likely to be asked
- Why do you want to join Sarvam?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- What do you do when a production issue happens on your code?
- Walk me through a system you built. How was it designed, and what would you change now?
- Tell me about a hard bug you tracked down. How did you find the cause?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Performance Engineer, Inference at Sarvam interview free →Performance Engineer, Inference Part of Sarvam's Performance Engineering team. We are hiring two specialized performance roles - Inference (this posting) and Kernels (companion posting). They are a vertical stack: the kernels team authors the µs-level GPU code, and the inference team integrates it into a running serving stack and owns the system-level numbers. If your depth genuinely spans both, apply to either and tell us - but most candidates are strongest in one, and we hire for that depth. Location: [Bengaluru / Chennai / Hybrid / On-site] · Team: Performance Engineering · Level: Senior About the team Sarvam serves multiple model families - small and large LLMs, Mixture-of-Experts, Indic ASR & TTS, streaming models and multimodal models - across a multi-node, multi-tenant fleet of Hoppers and Blackwells. The Performance Engineering team owns the numbers the rest of the company plans against: how fast we serve, how much it costs, and how much we get out of every GPU. This team works at the intersection of the serving runtime, the kernel layer, and the SRE org that keeps the fleet alive. About the role You will own Sarvam's production serving path for large distributed models end to end. You should be source-level fluent in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM - able to read and modify it where stock behavior does not fit our workloads - and you should operate and extend a distributed-serving stack at depth: disaggregated prefill-decode across nodes, distributed KV/cache transfer, and the routing and scheduling that span them. You will integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack, and you will build and train your own speculators - draft models, distillation from the target, acceptance-rate tuning against the live serving distribution - rather than only wiring in stock implementations. You will produce, and defend, the latency and throughput numbers the company plans against, and you will spend significant time in cross-team work with architecture co-design, the kernels team, the model team, and SRE. Your scoreboard: TTFT (p50 / p95 / p99), TPOT, throughput, GPU utilization, and cost per million tokens. What we're looking for 5+ years in ML systems, with 2+ years on inference serving at production scale. Your record shows concrete outcomes - tokens per day, throughput wins, p99 reductions - rather than "deployed a model." Experience serving 100B+ parameter models in production across multi-node tensor, pipeline, or expert parallelism. Source-level fluency in one of SGLang, vLLM, Dynamo, or TensorRT-LLM - you have modified the scheduler, the KV allocator, or the disaggregation path - and reading-level familiarity with the other three. Distributed serving at operating-and-extending depth: a disaggregated prefill-decode stack, distributed KV/cache transfer (Mooncake or equivalent), and cross-node routing and scheduling. You have run one of these in production and modified it where it didn't fit. Speculative decoding as a build-and-train competency: you have trained your own draft models or speculators (EAGLE / DFlash or otherwise), distilled them from a target model, measured and tuned acceptance rate against a real serving distribution, and composed speculation with the rest of the stack - not only integrated a stock implementation. Deep understanding of KV cache internals: block tables, copy-on-write, prefix sharing, and fragmentation. Working command of TP / PP / EP, NCCL primitives, and how they interact with the scheduler. Multi-tenant serving: model co-location and MIG / MPS isolation. Profiling fluency with Nsight Systems, framework tracing, and py-spy / perf. C++ and CUDA at a read-and-modify level. On-call ownership of an inference SLO. Strong pluses Upstream contributions to SGLang, vLLM, Dynamo, llm-d, TensorRT-LLM, or LMDeploy on non-trivial code paths. Direct production experience with Dynamo or llm-d at scale. Having operated a forked runtime in production. Published or shipped speculator work - a draft model or speculative-decoding technique you trained and measured. MoE serving at scale, long-context (128K+), multi-model serving, or Indic and multilingual workloads.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Solution SpecialistSarvam · bengaluru
- Engagement Manager, ChanakyaSarvam · delhi ncr
- Backend Engineer, ChanakyaSarvam · bengaluru
- Data Scientist - Evaluations, ChanakyaSarvam · bengaluru
- Embedded Data Scientist, ChanakyaSarvam · delhi ncr
- Embedded Infrastructure Engineer, ChanakyaSarvam · delhi ncr
- Product Manager (Models)Sarvam · bengaluru
- Partnerships & Alliances Lead (GSIs) - IndiaSarvam · bengaluru
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on ashby · posted 2026-08-10. ApplySarthi collects openings and links to application pages; the role is advertised by Sarvam, not by us.