Member of Technical Staff | ML Systems
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
2,476 open systems roles across 334 companies are on ApplySarthi right now, most of them in Bengaluru (88), Hyderabad (50), Pune (15).
- Systems EngineerPhilips
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - FederalServiceNow
- Backend Engineer, AI (Agent Systems)Bjak
- Electrical Technician, Magnet SystemsProxima Fusion GmbH
- Junior Mechanical Engineer – Quadcopter SystemsHarmattan Ai
What systems roles keep asking for: Python (24%), System design (14%), Linux (13%), C++ (13%) — counted across their open postings here.
Remote Member of Technical Staff jobs · Machine learning jobs · PyTorch jobs · Python jobs
Jobgether has 3,838 open roles listed here.
- Analista de Gestão de Mudanças
- Analista de Governança e Transparência ESG III - Temporária
- Associate Product Manager, BMO Global Asset Management
- Bilingual Field technology Consultant
- Bilingual Vocational Rehabilitation Specialist
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for systems roles keep coming back to Python, System design, Linux, C++. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with Machine learning? Tell me one thing you learned the hard way.
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Member of Technical Staff | ML Systems at Jobgether interview free →Accountabilities:: Build high-performance CUDA kernels and compute primitives supporting the training and serving of graph neural networks. Evolve distributed sampling and training infrastructure, including neighbor sampling and performance improvements for multi-node workloads. Develop and maintain the systems used to efficiently train and serve machine learning models at scale. Define binary and columnar data formats, including Lance, Arrow, and CSR/CSC representations, and own data materialization and feature backfills for training and evaluation. Establish data contracts and consumption requirements in collaboration with teams responsible for customer and proprietary datasets. Build and operate experiment tracking, checkpointing, and evaluation infrastructure with reproducibility as a default requirement. Own model registry, lineage, versioning, and compatibility across models, embeddings, and downstream models. Define and operate release gates that ensure every production, batch, or on-premise model corresponds to an authorized and governed release. Make model releases fully auditable by maintaining traceability across the data, code, configuration, and supporting evidence used to produce them. Improve time-to-experiment by making data, compute, and experiment tracking rapidly accessible to research teams. Improve time-to-governed-release by creating efficient and reliable paths from validated candidate models to production-ready releases. Optimize training throughput and GPU utilization across large-scale foundation model workloads. Build internal platform capabilities with the mindset of a product, focusing on the needs and productivity of researchers and platform engineers. Maintain high standards for reliability, reproducibility, performance, and operational quality across ML systems. Requirements: Strong systems engineering background combined with production-quality Python development skills. Professional experience with distributed machine learning training, including technologies such as Ray or PyTorch Distributed. Experience operating or developing multi-node GPU workloads and understanding the challenges of distributed compute. Experience working with columnar data formats and large-scale data materialization pipelines. Familiarity with ML lifecycle infrastructure, including experiment tracking, model registries, evaluation systems, checkpoints, and reproducibility tooling. Strong understanding of software engineering principles for building reliable, maintainable, production-grade infrastructure. Ability to think of internal platforms as products with real users, requirements, feedback loops, and measurable outcomes. Ability to take ownership of systems and outcomes rather than focusing narrowly on individual implementation tasks. Strong analytical and problem-solving skills, with the ability to work across data, compute, model, and infrastructure layers. A product-oriented mindset and interest in enabling researchers and engineers to work more effectively. Data science expertise is not required, provided you have strong systems engineering capabilities and an understanding of ML infrastructure. Experience developing CUDA kernels or optimizing GPU performance is a strong advantage. Knowledge of graph neural networks, graph sampling, or large-scale graph workloads is a plus. Experience with Lance, Arrow, or comparable columnar or indexed storage technologies is beneficial. Familiarity with multi-cloud GPU infrastructure, including tools such as SkyPilot, is advantageous. Experience with model governance, auditability, or regulated production environments—particularly financial services—is a plus. Benefits: Fully remote position based in Brazil. Full-time opportunity within an engineering organization focused on high-impact ML infrastructure. Broad technical ownership across ML systems, distributed computing, GPU performance, data infrastructure, and model governance. Opportunity to work on foundational machine learning systems supporting research and production workloads. Direct impact on experiment velocity, training performance, model reproducibility, and governed releases. Opportunity to work with advanced technologies including CUDA, distributed GPU workloads, graph neural networks, columnar data systems, and ML lifecycle infrastructure. Environment that values technical depth while giving engineers ownership of systems and outcomes. Internal-platform mindset, with researchers and platform engineers treated as real users whose productivity drives priorities. Opportunity to shape reliable ML infrastructure designed to operate across cloud, production, batch, and customer environments.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-23. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.