ApplySarthi Match jobs to your CV

Sr Platform Engineer, ML Infrastructure

Jobgether

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

2,026 open infrastructure roles across 290 companies are on ApplySarthi right now, most of them in Bengaluru (66), Hyderabad (28), Delhi NCR (12).

What infrastructure roles keep asking for: AWS (30%), System design (22%), Kubernetes (21%), Python (21%), Observability (19%), Terraform (15%), CI/CD (13%), Linux (12%) — counted across their open postings here.

Platform Engineer jobs in the United States · Remote Platform Engineer jobs · AWS jobs · Airflow jobs · Computer vision jobs · Kubernetes jobs

Jobgether has 3,942 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for infrastructure roles keep coming back to AWS, System design, Kubernetes, Python. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with Machine learning? Tell me one thing you learned the hard way.
  3. Walk me through a model you built, from the data to how it was used.
  4. How did you know your model was actually good, and not just good on your test set?
  5. Tell me about a time the data was messy or wrong. What did you do?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Sr Platform Engineer, ML Infrastructure at Jobgether interview free →

Accountabilities:: Design, build, and operate scalable ML infrastructure and platform capabilities supporting experimentation, training, deployment, and production operations. Develop developer tooling, services, automation, and infrastructure that help ML and engineering teams build and operate production systems more efficiently. Lead complex technical initiatives independently, from problem definition and architecture through implementation, rollout, and operational ownership. Make architectural decisions that balance immediate delivery needs with long-term scalability, reliability, maintainability, and developer experience. Partner with ML engineers, infrastructure teams, and other stakeholders to understand needs and deliver effective platform solutions. Identify and solve challenging infrastructure problems involving performance, reliability, scalability, and operational efficiency. Drive adoption and continuous improvement by incorporating feedback from engineering teams using the platform. Maintain high standards for software quality, production readiness, observability, and operational excellence. Deliver platform capabilities that create measurable engineering and business impact across multiple teams and use cases. Requirements: 5+ years of professional software engineering experience, particularly in platform engineering, infrastructure, or distributed systems. Strong Python engineering skills, including experience developing production services, SDKs, automation, or platform tooling. Proven experience designing, building, and operating production platforms used by multiple engineering teams. Solid understanding of ML platform architecture and the end-to-end machine learning lifecycle, including experimentation, distributed training, model deployment, and production operations. Experience building and operating applications on Kubernetes and cloud platforms, with AWS experience preferred. Strong understanding of production reliability, observability, scalability, and operational best practices. Strong technical judgment and the ability to independently drive complex initiatives from discovery through production while collaborating across technical teams. Experience with developer platforms, internal tooling, or services that improve engineering productivity and reduce operational complexity is preferred. Familiarity with workflow orchestration or distributed computing technologies such as Airflow, Kubeflow, Ray, Spark, or similar systems is a plus. Experience designing or optimizing distributed, GPU-intensive compute platforms for ML training, inference, or large-scale image processing is preferred. Experience supporting production ML platforms in computer vision, robotics, or related technical domains is advantageous. Demonstrated technical leadership through architecture, mentoring, or influencing technical direction across teams. Benefits: Base salary range of $160,000–$287,000 per year, depending on experience, qualifications, education, location, and skills. Eligibility for an annual performance bonus. Competitive benefits package. Full-time, remote position within the United States. Visa sponsorship may be available for this position. Opportunities for career development, mentorship, and learning and development. Inclusive and collaborative work environment focused on meaningful, technically challenging work. Opportunity to work on advanced machine learning, robotics, and intelligent machinery technologies with cross-disciplinary teams.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-24. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.