ApplySarthi Match jobs to your CV

Member of Technical Staff | Inference Platform

Jobgether

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

132 open inference roles across 34 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).

What inference roles keep asking for: LLMs (49%), Python (49%), Machine learning (36%), PyTorch (26%), System design (26%), Kubernetes (24%), AWS (23%), Observability (20%) — counted across their open postings here.

Remote Member of Technical Staff jobs · AWS jobs · GCP jobs · Kubernetes jobs · Machine learning jobs

Jobgether has 3,942 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with Kubernetes? Tell me one thing you learned the hard way.
  3. Tell me about a time you disagreed with your manager. What happened?
  4. Where do you want to be in three years?
  5. What is a weakness you are working on, and how?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Member of Technical Staff | Inference Platform at Jobgether interview free →

Accountabilities:: Evolve and operate a Kubernetes-based online and batch inference runtime supporting production machine learning workloads. Run large-scale batch inference through ephemeral jobs, implementing multi-dimensional admission control across CPU, memory, and GPU resources. Build and extend Kubernetes controllers and custom resources to support reliable model execution and scheduling. Optimize model inference engines and feature-processing pipelines using efficient, vectorized, and columnar operations. Develop efficient mechanisms for serving graphs and data from Lance-based storage. Own the execution of training, post-training, and fine-tuning jobs across both cloud infrastructure and customer-hosted Kubernetes environments. Design and improve autoscaling strategies, GPU serving, inference performance, and infrastructure cost efficiency. Implement comprehensive telemetry and monitoring for model execution, enabling performance, reliability, and cost optimization. Solve complex infrastructure challenges including deterministic job sizing, checkpointing, recovery of batch workloads, difficult input files, and automated profiling of newly accepted models. Design robust approaches to per-customer encryption, workload isolation, and secure execution across shared and customer environments. Improve serving availability, online inference latency, batch throughput, GPU utilization, and cost per prediction or training job. Ensure training and batch workloads complete reliably and on schedule without requiring manual intervention or repeated retries. Write production-quality code, participate in rigorous code reviews, and take operational ownership of the systems you build. Requirements: Professional experience operating model-serving infrastructure or large-scale batch compute workloads on Kubernetes. Strong experience building Kubernetes controllers, operators, or comparable Kubernetes-native infrastructure. Strong software engineering skills with production-quality Python and experience developing reliable, maintainable systems. Demonstrated ability to profile and optimize data-intensive Python pipelines and identify performance bottlenecks. Understanding of distributed systems, workload scheduling, resource allocation, and production infrastructure. Experience with performance optimization across compute, memory, storage, and GPU resources. Strong cost-awareness and the ability to treat infrastructure efficiency as an important product requirement. Experience operating production systems and willingness to take responsibility for the reliability and performance of systems you build. Strong engineering judgment, problem-solving skills, and ability to work effectively on open-ended infrastructure challenges. Familiarity with ML inference workloads and the operational requirements of serving models at scale. Experience with Ray, Ray Serve, or KubeRay in production is a strong advantage. Knowledge of Kueue or other batch scheduling and admission-control technologies is beneficial. Experience with GPU serving and performance optimization is a plus. Familiarity with Arrow, Parquet, Lance, or other columnar data formats is advantageous. Experience shipping and operating software on customer-hosted Kubernetes environments is valuable. Experience with GCP or AWS and platforms such as GKE or EKS is a plus. Experience working in financial services or other regulated environments is beneficial. Benefits: Full-time, fully remote position based in Brazil. Opportunity to own critical production infrastructure powering both real-time and large-scale batch machine learning workloads. Broad technical scope across Kubernetes, ML inference, distributed systems, GPU infrastructure, data processing, and cloud platforms. Hands-on opportunity to build and evolve Kubernetes controllers, scheduling systems, autoscaling infrastructure, and model-serving platforms. Direct impact on measurable engineering and business outcomes, including availability, latency, throughput, GPU utilization, and cost per prediction. Opportunity to solve complex infrastructure challenges involving reliability, recovery, security, encryption, isolation, and multi-tenant execution. Exposure to cloud and customer-hosted environments, including production Kubernetes deployments. Engineering culture centered on ownership, production quality, operational responsibility, and measurable outcomes. Opportunity to work on infrastructure where compute efficiency is treated as a core product capability.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-23. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.