Member of Technical Staff | Inference Platform
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
132 open inference roles across 34 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).
- Member of Technical Staff – AI Inference platform, featuresLyceum
- Senior Machine Learning Engineer, LLM Inference OptimizationNebius
- AI Engineer 5 (FM Hosting, LLM Inference)Capitalone
- Software Engineer II - AI/ML, Neuron InferenceAnnapurna Labs (U.S.) Inc.
- Senior Software Engineer, AI Inference SystemsNvidia
What inference roles keep asking for: LLMs (49%), Python (49%), Machine learning (36%), PyTorch (26%), System design (26%), Kubernetes (24%), AWS (23%), Observability (20%) — counted across their open postings here.
Remote Member of Technical Staff jobs · AWS jobs · GCP jobs · Kubernetes jobs · Machine learning jobs
Jobgether has 3,942 open roles listed here.
- AI Researcher — Distillation
- AI Researcher — Distillation
- Art Director
- Applied ML Engineer
- AI Science Writer, Nebius Academy (Contract)
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- Tell me about a time you disagreed with your manager. What happened?
- Where do you want to be in three years?
- What is a weakness you are working on, and how?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Member of Technical Staff | Inference Platform at Jobgether interview free →Accountabilities:: Evolve and operate a Kubernetes-based online and batch inference runtime supporting production machine learning workloads. Run large-scale batch inference through ephemeral jobs, implementing multi-dimensional admission control across CPU, memory, and GPU resources. Build and extend Kubernetes controllers and custom resources to support reliable model execution and scheduling. Optimize model inference engines and feature-processing pipelines using efficient, vectorized, and columnar operations. Develop efficient mechanisms for serving graphs and data from Lance-based storage. Own the execution of training, post-training, and fine-tuning jobs across both cloud infrastructure and customer-hosted Kubernetes environments. Design and improve autoscaling strategies, GPU serving, inference performance, and infrastructure cost efficiency. Implement comprehensive telemetry and monitoring for model execution, enabling performance, reliability, and cost optimization. Solve complex infrastructure challenges including deterministic job sizing, checkpointing, recovery of batch workloads, difficult input files, and automated profiling of newly accepted models. Design robust approaches to per-customer encryption, workload isolation, and secure execution across shared and customer environments. Improve serving availability, online inference latency, batch throughput, GPU utilization, and cost per prediction or training job. Ensure training and batch workloads complete reliably and on schedule without requiring manual intervention or repeated retries. Write production-quality code, participate in rigorous code reviews, and take operational ownership of the systems you build. Requirements: Professional experience operating model-serving infrastructure or large-scale batch compute workloads on Kubernetes. Strong experience building Kubernetes controllers, operators, or comparable Kubernetes-native infrastructure. Strong software engineering skills with production-quality Python and experience developing reliable, maintainable systems. Demonstrated ability to profile and optimize data-intensive Python pipelines and identify performance bottlenecks. Understanding of distributed systems, workload scheduling, resource allocation, and production infrastructure. Experience with performance optimization across compute, memory, storage, and GPU resources. Strong cost-awareness and the ability to treat infrastructure efficiency as an important product requirement. Experience operating production systems and willingness to take responsibility for the reliability and performance of systems you build. Strong engineering judgment, problem-solving skills, and ability to work effectively on open-ended infrastructure challenges. Familiarity with ML inference workloads and the operational requirements of serving models at scale. Experience with Ray, Ray Serve, or KubeRay in production is a strong advantage. Knowledge of Kueue or other batch scheduling and admission-control technologies is beneficial. Experience with GPU serving and performance optimization is a plus. Familiarity with Arrow, Parquet, Lance, or other columnar data formats is advantageous. Experience shipping and operating software on customer-hosted Kubernetes environments is valuable. Experience with GCP or AWS and platforms such as GKE or EKS is a plus. Experience working in financial services or other regulated environments is beneficial. Benefits: Full-time, fully remote position based in Brazil. Opportunity to own critical production infrastructure powering both real-time and large-scale batch machine learning workloads. Broad technical scope across Kubernetes, ML inference, distributed systems, GPU infrastructure, data processing, and cloud platforms. Hands-on opportunity to build and evolve Kubernetes controllers, scheduling systems, autoscaling infrastructure, and model-serving platforms. Direct impact on measurable engineering and business outcomes, including availability, latency, throughput, GPU utilization, and cost per prediction. Opportunity to solve complex infrastructure challenges involving reliability, recovery, security, encryption, isolation, and multi-tenant execution. Exposure to cloud and customer-hosted environments, including production Kubernetes deployments. Engineering culture centered on ownership, production quality, operational responsibility, and measurable outcomes. Opportunity to work on infrastructure where compute efficiency is treated as a core product capability.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-23. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.