ApplySarthi Match jobs to your CV

Lead Machine Learning Engineer - ML Infrastructure

Jobgether

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

2,026 open infrastructure roles across 290 companies are on ApplySarthi right now, most of them in Bengaluru (66), Hyderabad (28), Delhi NCR (12).

What infrastructure roles keep asking for: AWS (30%), System design (22%), Kubernetes (21%), Python (21%), Observability (19%), Terraform (15%), CI/CD (13%), Linux (12%) — counted across their open postings here.

Machine Learning Engineer jobs in Canada · Remote Machine Learning Engineer jobs · Kubernetes jobs · Machine learning jobs · Observability jobs · Spark jobs

Jobgether has 3,942 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for infrastructure roles keep coming back to AWS, System design, Kubernetes, Python. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with Machine learning? Tell me one thing you learned the hard way.
  3. How would you explain your model's result to someone who is not technical?
  4. What would you check first if a model's accuracy dropped after going live?
  5. When would you not use machine learning for a problem?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Lead Machine Learning Engineer - ML Infrastructure at Jobgether interview free →

Accountabilities:: Set the technical strategy and own end-to-end delivery of the machine learning platform across training, experimentation, batch and online inference, and edge deployment. Make architectural decisions across ML infrastructure layers and serve as the primary technical accountability point for multiple AI product teams. Design, launch, and continuously improve ML-powered features while co-owning production outcomes such as safety metrics, reliability, performance, and cost. Design and operate scalable online and batch inference systems using technologies such as Ray and Spark , including deployment patterns, observability, service-level objectives (SLOs), and unified training-to-production workflows. Partner with firmware and edge engineering teams to package, validate, and deploy machine learning models to connected devices. Build feedback loops between edge deployments and cloud infrastructure to support continuous model and system improvement. Own reliability, observability, security, and operational practices for ML systems spanning cloud and edge environments. Establish and improve on-call practices, incident response processes, infrastructure hardening, and production reliability standards. Own or co-own end-to-end technical delivery for high-priority and high-risk initiatives, from modeling and system architecture through production rollout. Serve as the technical authority for ML infrastructure architecture and establish direction for applied ML, firmware, security, and data platform teams. Mentor senior engineers and applied scientists while helping teams make sound technical trade-offs at the appropriate level of abstraction. Improve developer experience through documentation, engineering standards, reusable practices, and clear platform guidance. Contribute to and represent the organization within relevant open-source communities, including Ray, Spark, RayDP, and Kubernetes . Balance research velocity with platform stability and communicate technical trade-offs effectively across science and engineering teams. Champion customer-focused, long-term, inclusive, collaborative, and growth-oriented engineering practices. Requirements 10+ years of experience in machine learning engineering , with demonstrated technical leadership across at least two major ML platform domains such as distributed training, data or research infrastructure, cloud inference, or feature engineering. Proven track record of delivering ML-powered products or features end-to-end, from technical design through production deployment and iteration, with measurable product or business impact. Strong hands-on expertise with Ray and Kubernetes in production environments. Strong experience with Spark is highly preferred. Deep understanding of machine learning fundamentals beyond pipelines, including evaluation methodology, dataset design, ablation, model drift, and the ability to review and redirect modeling approaches. Ability to bridge research and engineering teams and translate ML requirements into scalable, reliable production systems. Demonstrated cross-organizational technical leadership, including influencing platform decisions, roadmaps, and go/no-go decisions based on throughput, latency, reliability, and cost trade-offs. Strong understanding of distributed ML infrastructure and production systems at scale. Experience navigating the trade-offs between scientific experimentation and platform stability, with the communication skills needed to align both sides. Strong architectural judgment and ability to operate as the senior technical authority for complex ML infrastructure initiatives. Excellent communication and collaboration skills, with the ability to influence senior technical stakeholders without relying solely on formal authority. Strong mentoring and technical leadership capabilities. Prior contributions to open-source projects such as Ray, Spark, RayDP, or Kubernetes are a plus. Experience with enterprise security and compliance requirements in ML environments is a plus. Experience with edge or on-device machine learning and collaboration with firmware or embedded engineering teams is a plus. Benefits Annual base salary range of $196,000–$269,500 CAD , with actual compensation varying based on factors such as location, knowledge, skills, and experience. Eligibility for an initial RSU grant with no vesting cliff , subject to applicable plan terms. Ongoing equity refresh opportunities tied to performance, subject to plan terms and conditions. Performance-based bonus or variable compensation for eligible roles. Flexible, employee-led remote working model. Comprehensive health benefits. Parental leave programs. Professional development stipend. Opportunities to work on high-impact AI and machine learning infrastructure at significant scale. Opportunity to influence platform architecture and technical direction across multiple product teams. Exposure to cloud, edge, distributed computing, machine learning, and open-source technologies. Collaborative environment emphasizing long-term ownership, growth, inclusion, and customer outcomes. Reasonable accommodations available throughout the recruitment process for qualified candidates.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-24. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.