ML / RAG / Inference Engineer
Federis
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
132 open inference roles across 34 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).
- Member of Technical Staff – AI Inference platform, featuresLyceum
- Senior Machine Learning Engineer, LLM Inference OptimizationNebius
- Senior Machine Learning Engineer, LLM Inference OptimizationJobgether
- AI Engineer 5 (FM Hosting, LLM Inference)Capitalone
- Software Engineer II - AI/ML, Neuron InferenceAnnapurna Labs (U.S.) Inc.
What inference roles keep asking for: LLMs (49%), Python (49%), Machine learning (36%), PyTorch (26%), System design (26%), Kubernetes (24%), AWS (23%), Observability (20%) — counted across their open postings here.
Customer success jobs · Elasticsearch jobs · Go jobs · Java jobs
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Federis.
Questions you are likely to be asked
- Why do you want to join Federis?
- What is your experience with RAG? Tell me one thing you learned the hard way.
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the ML / RAG / Inference Engineer at Federis interview free →**Mission** Build the AI orchestration layer for Federis so customers can govern retrieval, prompts, model access, inference routing, evaluation, and observability across replaceable external AI and vector systems. **Key Responsibilities** • Implement RAG workflows, retrieval configuration, prompt/version governance, inference routing, evaluation hooks, and model usage telemetry. • Build provider adapters for vector databases, model serving systems, embedding services, and customer-managed inference endpoints. • Create quality, latency, cost, grounding, and safety evaluation pipelines that support enterprise release gates and customer reporting. • Partner with security and product teams on policy enforcement for data access, prompt execution, model selection, and audit capture. • Design scalable data and inference paths that work across SaaS, private cloud, and air-gapped customer environments. • Document model and retrieval behavior clearly enough for solution engineers, customers, and auditors to understand. Required Experience • 4+ years in ML engineering, search, data platforms, RAG systems, model serving, or AI product engineering. • Strong software engineering ability in Python, TypeScript, Go, Java, or similar production languages. • Practical experience with embeddings, vector search, ranking, chunking, retrieval evaluation, prompt/version management, and model APIs. • Understanding of latency, throughput, cost, reliability, privacy, and observability tradeoffs in AI systems. • Ability to build adapter-based integrations without binding the core product to a single AI vendor or database. **Useful Differentiators** • Experience with vLLM, Triton, Milvus, Qdrant, Weaviate, OpenSearch, Apache Solr, or customer-hosted LLM stacks. • Experience creating AI evaluation harnesses, guardrails, prompt governance, red-team tests, or model risk workflows. • Prior work in regulated enterprise AI, data governance, knowledge management, or secure document intelligence. **Success Scorecard** • AI integrations are replaceable and measurable across quality, latency, cost, safety, and audit dimensions. • RAG workflows produce reproducible evidence for what data was used, what model was called, and what policy applied. • Customer deployments can use their preferred inference and vector systems without core rewrites. • Evaluation results inform roadmap priorities and customer success plans.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on wellfound · posted 2026-09-15. ApplySarthi collects openings and links to application pages; the role is advertised by Federis, not by us.