Senior Data Engineer (AI/ML)
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
7,274 open data roles across 649 companies are on ApplySarthi right now, most of them in Bengaluru (415), Hyderabad (313), Mumbai (155).
- Data Architect – Data Products & Data AnalysisCallista Group AG
- RE/RS, Data Understanding - FoundationsOpenAI
- Junior Data Engineer & MarTech Specialist (Mobile Apps) (m/w/d)Trg
- Senior Software Engineer, Backend - Data LayerCamunda
- Senior AI Data Expert (f/m/x)exmox
What data roles keep asking for: AWS (24%), SQL (23%), Python (22%) — counted across their open postings here.
Data Engineer jobs in India · Remote Data Engineer jobs · Airflow jobs · CI/CD jobs · Databricks jobs · ETL jobs
Jobgether has 3,935 open roles listed here.
- AI Researcher — Distillation
- AI Researcher — Distillation
- Accounting & Regulatory Reporting
- Accounts Receivable Coordinator
- AI Security Analyst
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for data roles keep coming back to AWS, SQL, Python. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Data Engineer (AI/ML) at Jobgether interview free →Accountabilities:: Design and build AI/LLM data pipelines supporting training, inference, evaluation, embeddings, and retrieval workloads. Build production-grade RAG systems covering ingestion, chunking, embedding generation, indexing, retrieval, reranking, and context construction. Develop AI applications using LLMs, structured outputs, function and tool calling, and agentic workflows. Build and optimize semantic search and vector retrieval systems. Develop frameworks for LLM evaluation, monitoring, tracing, quality measurement, latency analysis, and cost optimization. Design scalable batch and streaming pipelines using Databricks, Apache Spark, Delta Lake, Snowflake, and Airflow. Build data products and platforms that make structured and unstructured enterprise data accessible to AI applications. Develop reliable ETL/ELT pipelines and optimize large-scale distributed workloads for performance and cost. Establish data quality, governance, lineage, security, and observability practices. Partner with ML and application engineering teams to transition AI prototypes into reliable, production-ready systems. Support large-scale data platforms, real-time processing, event-driven architectures, and complex orchestration workflows. Contribute to AI evaluation datasets and pipelines that measure quality, accuracy, relevance, latency, and cost. Monitor production AI systems, including token usage, model performance, failures, latency, and overall system health. Requirements: 5+ years of experience in data engineering, software engineering, distributed systems, or a related field. Strong programming skills in Python and/or Scala/Java, combined with advanced SQL capabilities. Hands-on experience with Databricks, Snowflake, Apache Spark, Delta Lake, and Airflow. Strong experience working with cloud-based data platforms and scalable data architectures. Practical experience building applications using LLMs or Generative AI. Strong understanding of RAG architectures, embeddings, vector databases, semantic search, and retrieval systems. Familiarity with prompting, structured outputs, tool calling, model evaluation, and other modern LLM concepts. Experience designing scalable, reliable, observable production data systems. Strong knowledge of large-scale data platforms, distributed processing, and complex data workflows. Experience with real-time and streaming architectures using technologies such as Kafka or Spark Structured Streaming. Experience designing low-latency pipelines and event-driven architectures. Strong experience with multi-stage ETL/ELT and data orchestration workflows using Airflow or similar platforms. Experience optimizing Spark or Databricks workloads through partitioning, clustering, caching, joins, and compute optimization. Experience supporting both batch and real-time AI/ML workloads. Experience with LLM/AI evaluation frameworks, automated evaluations, experimentation, quality metrics, and evaluation datasets. Familiarity with AI observability and tracing, including token usage, model performance, latency, failures, and production monitoring. Experience with LangGraph, LangChain, LlamaIndex, or similar AI orchestration frameworks is preferred. Experience with vector databases such as Qdrant, Pinecone, Weaviate, or Databricks Vector Search is preferred. Familiarity with Kafka, MLflow, Unity Catalog, Databricks Mosaic AI, or model-serving platforms is a plus. Strong understanding of distributed systems, cloud architecture, APIs, CI/CD, data governance, and production operations. Benefits: 100% remote position across India. Work from almost anywhere for up to 20 days per year. Generous paid vacation and time off for your birthday. Paid parental leave. Company-paid therapy sessions through SpringHealth. Company-paid Headspace subscription. Annual company-wide week off to support rest and well-being. Generous health insurance and pension fund. Tax optimization options. Development Dollars and leadership development opportunities. Access to thousands of on-demand learning resources. Paid volunteer time. Travel discounts. Employee Resource Groups. Quarterly team offsites. Global and collaborative working environment. Opportunities to work with large-scale data engineering, Generative AI, distributed systems, and modern AI infrastructure. Flexible collaboration across international teams and time zones, with local laws and regulations taken into consideration.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-21. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.