Forward Deployed Engineer AI Inference
Lyceum
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
150 open inference roles across 44 companies are on ApplySarthi right now, most of them in Delhi NCR (2), Bengaluru (2).
- (Senior) Sales Manager DACH - AI & InferenceImpossiblecloud
- Software Engineer, InferenceLumaai
- Senior Deep Learning Software Engineer, Inference and Model OptimizationNvidia
- AI Engineer 5 (FM Hosting, LLM Inference)Capitalone
- Staff Software Engineer, Console GPU Inferencesonyinteractiveentertainmentglobal
What inference roles keep asking for: LLMs (54%), Python (51%), Machine learning (43%), PyTorch (30%), Kubernetes (27%), System design (27%), C++ (26%), Observability (26%) — counted across their open postings here.
Lyceum has 8 open roles listed here.
- Forward Deployed Engineer – AI Inference (Intern)
- Customer Success Associate
- Infrastructure Support Engineer
- Head of Customer Success – GPU Infrastructure
- Founder Associate to the CTO
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Lyceum.
Questions you are likely to be asked
- Why do you want to join Lyceum?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- What would you check first if a model's accuracy dropped after going live?
- When would you not use machine learning for a problem?
- Walk me through a model you built, from the data to how it was used.
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Forward Deployed Engineer AI Inference at Lyceum interview free →The Role You will own the technical side of our largest and most complex inference deals, from first discovery call to go-live. You're the trusted technical counterpart for our customers' CTOs and ML leads, you design the setups that run their models, and you work hand in hand with our commercial team to get deals closed. You'll be one of two senior technical owners of our inference deals. Together with our Inference Sales Engineering lead, you'll also turn what we learn in the field into a repeatable playbook and product. What You'll Do Own the technical side of our largest and most complex inference deals end to end, from discovery to go-live Be the senior technical counterpart for customers' CTOs and ML leads on architecture, sizing, SLAs, latency/throughput trade-offs and cost Design dedicated inference setups: model choice, GPU type and count, parallelism, quantization, inference engine and configuration Lead benchmarks and proofs of concept, and turn the results into clear recommendations Scope customer customizations with our engineering team and decide what becomes product Build the playbook and tooling for matching workloads to GPUs, and feed our roadmap with what you learn in the field Work in tandem with our commercial team on pricing, proposals and the technical parts of contracts What We're Looking For A degree in computer science, data science or a closely related field 4+ years in a customer-facing technical role, e.g. solutions or sales engineer, forward deployed engineer, ML engineer working closely with customers, or technical consultant Hands-on experience with LLM inference and model serving (e.g. vLLM, SGLang, TensorRT-LLM, Triton) and GPU sizing A track record of owning the technical side of complex B2B deals, ideally with enterprise customers You're credible with customer CTOs and engineers alike, and you explain trade-offs clearly Entrepreneurial mindset: give you an outcome, and you find a way without getting blocked Fluent English Bonus Points Experience at an inference provider, GPU cloud or AI infrastructure company Benchmarking and performance optimization: throughput, latency, cost per token Experience with quantization, parallelism strategies, KV-cache and batching Startup experience German Why us? Outstanding team: Work with some of the best engineers in the world, coming from hedge funds, big tech, AI startups and top universities Once in a lifetime opportunity: Early-stage company in the fastest-growing market in the world Ownership: Shape how European AI companies access GPU compute European mission: Build sovereign, GDPR-compliant AI infrastructure for the next generation of deep-tech Find Jobs in Germany on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Member of Technical Staff – AI Inference platform, featuresLyceum
- Member of Technical Staff – AI Inference platformLyceum
- Customer Success AssociateLyceum
- Infrastructure Support EngineerLyceum
- Head of Customer Success – GPU InfrastructureLyceum
- Founder Associate to the CTOLyceum
- Forward Deployed Engineer – AI Inference (Intern)Lyceum
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-10-11. ApplySarthi collects openings and links to application pages; the role is advertised by Lyceum, not by us.