Member of Technical Staff – AI Inference platform
Lyceum
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
132 open inference roles across 34 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).
- Senior Machine Learning Engineer, LLM Inference OptimizationNebius
- Senior Machine Learning Engineer, LLM Inference OptimizationJobgether
- AI Engineer 5 (FM Hosting, LLM Inference)Capitalone
- Software Engineer II - AI/ML, Neuron InferenceAnnapurna Labs (U.S.) Inc.
- Senior Software Engineer, AI Inference SystemsNvidia
What inference roles keep asking for: LLMs (49%), Python (49%), Machine learning (36%), PyTorch (26%), System design (26%), Kubernetes (24%), AWS (23%), Observability (20%) — counted across their open postings here.
Remote Member of Technical Staff jobs · Kubernetes jobs · LLMs jobs · Observability jobs · Python jobs
Lyceum has 2 open roles listed here.
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Lyceum.
Questions you are likely to be asked
- Why do you want to join Lyceum?
- What is your experience with Observability? Tell me one thing you learned the hard way.
- What would you check first if a model's accuracy dropped after going live?
- When would you not use machine learning for a problem?
- Walk me through a model you built, from the data to how it was used.
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Member of Technical Staff – AI Inference platform at Lyceum interview free →Your mission You will make Lyceum's AI inference platform reliable, secure, and scalable - ensuring it performs under pressure as we grow to thousands of concurrent users. While others on the team expand what the platform can do, your job is to make sure it keeps working, fails gracefully, and gets faster over time. Your focus Scalability: Architect and implement the systems that allow our inference platform to scale to thousands of concurrent users. This includes request routing, load balancing, autoscaling, and resource scheduling across GPU clusters. Reliability and observability: Build robust monitoring, alerting, and incident response tooling. Design for graceful degradation, automatic recovery, and minimal downtime. Performance engineering: Profile and optimise the full inference path from request ingestion through model execution to response delivery. Identify and eliminate bottlenecks at every layer. Infrastructure evolution: Evaluate and integrate open-source inference frameworks and tooling (Dynamo, vLLM, Triton, etc.) where they improve throughput, latency, or stability of the serving stack. Your KPIs Platform uptime and availability (SLA adherence) P50/P95/P99 latency and throughput under load Time-to-detection and time-to-resolution for incidents Scalability milestones (concurrent users, requests per second, GPU utilisation) Your profile We consider candidates from diverse backgrounds, with a deep love for technical challenges and the desire to take on ownership beyond what's reasonably expected. Requirements 3+ years of experience in backend, infrastructure, or systems engineering Strong proficiency in Go and Python Experience building or operating a model serving platform or ML platform Solid understanding of systems performance - profiling, benchmarking, and optimising latency and throughput Familiarity with observability tooling (Prometheus, Grafana, OpenTelemetry, or similar) Understanding of security fundamentals - network isolation, authentication, encryption, multi-tenancy Nice to have Experience with NVIDIA Dynamo or similar inference orchestration/routing frameworks Hands-on experience with GPU serving infrastructure (vLLM, Triton, TensorRT-LLM) Experience with Kubernetes in a production environment (deployment, networking, resource management) Experience operating at scale (10k+ RPS, multi-region, multi-cluster) Comfortable working on-call or in incident response when things break Why us? Outstanding team: Work with some of the best engineers in the world, coming from hedge funds, big tech, AI startups and top universities Once in a lifetime opportunity: Early-stage company in the fastest-growing market in the world Ownership: Shape how European AI companies access GPU compute European mission: Build sovereign, GDPR-compliant AI infrastructure for the next generation of deep-tech Find more English Speaking Jobs in Switzerland on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-09-26. ApplySarthi collects openings and links to application pages; the role is advertised by Lyceum, not by us.