Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC)
Jobgether
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
14 open evaluations roles across 10 companies are on ApplySarthi right now, most of them in Bengaluru (3).
- AI Engineer 5 (Gen AI Platform Services: Agentic AI, Guardrails, Evaluations)Capitalone
- Principal AI Evaluations Platform EngineerMicrosoftcorporation
- Physical Security Specialist, SPEAR Evaluations, US Amazon Dedicated CloudAmazon
- Research Intern - AI EvaluationsKarya · bengaluru
- Researcher, Agent Safety, Training and EvaluationsOpenai
What evaluations roles keep asking for: LLMs (43%), Python (43%), Generative AI (21%), Machine learning (21%), Observability (21%), Supply chain (21%), C++ (14%), Go (14%) — counted across their open postings here.
Airflow jobs · LLMs jobs · Python jobs · Spark jobs
Jobgether has 4,083 open roles listed here.
- Artificial Intelligence (AI) Technician
- Artificial Intelligence (AI) Technician
- .Net Technical Lead
- Account Executive/ Sr. Account Executive
- Account Operations & Farming Specialist
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for evaluations roles keep coming back to LLMs, Python, Generative AI, Machine learning. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- How would you design an API for a feature you have worked on?
- What do you do when a production issue happens on your code?
- Walk me through a system you built. How was it designed, and what would you change now?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC) at Jobgether interview free →Accountabilities: Design and build reusable data curation infrastructure, including shared libraries, components, and workflows that replace fragmented notebook-based processes. Develop scalable LLM-powered pipelines that transform raw catalog, metadata, and other data sources into training and evaluation datasets such as question-answer pairs and synthetic scenarios. Implement large-scale batch inference workflows while balancing data quality, computational efficiency, token usage, and cost. Develop sampling strategies that optimize coverage, diversity, difficulty, and representation across relevant content and member segments. Create data-quality and filtering systems using techniques such as LLM-as-judge scoring, evaluation-model-based ranking, deduplication, validation, and other quality controls. Partner closely with researchers to design experiments that measure how data curation choices affect downstream model behavior and performance. Establish curated datasets as discoverable, reusable artifacts with clear versioning, lineage, documentation, and reproducibility. Drive adoption of standardized data curation practices across engineering, research, and modeling teams. At the L6 level, provide technical leadership across data and evaluation infrastructure and help define technical direction for multi-engineer initiatives. Requirements Strong software engineering expertise in Python, including experience developing reusable infrastructure, libraries, frameworks, or platforms used by other engineers and researchers. Hands-on experience building LLM-driven data generation or transformation pipelines, including synthetic data generation, structured outputs, or large-scale batch inference. Practical experience with data quality techniques such as sampling, filtering, deduplication, validation, and model-based quality scoring, including LLM-as-judge approaches. Strong modeling intuition and an understanding of how dataset composition and curation decisions can influence model behavior and performance. Experience designing experiments or evaluation approaches to measure the impact of data and modeling decisions. Experience with distributed data processing technologies such as Spark, Ray, or comparable frameworks. Excellent collaboration and communication skills, particularly when partnering with researchers, data scientists, and platform engineering teams. For L6 roles, demonstrated experience with LLM evaluation systems is required. For L6 roles, demonstrated technical leadership across data or evaluation infrastructure, including setting technical direction for multi-engineer initiatives, is required. Experience with dataset versioning, lineage, artifact management, experiment tracking, or model registries is highly valued. Experience optimizing large-scale LLM inference for cost, throughput, or operational efficiency is a plus. Familiarity with human annotation workflows and methods for calibrating LLM judges against human ratings is beneficial. Experience with pipeline orchestration frameworks such as Metaflow, Airflow, or similar tools is advantageous. Background in recommendation systems, personalization, search, content catalogs, or metadata-driven applications is a plus. Benefits Annual compensation range of $600,000–$1,066,000 , with the range varying based on location and individual market factors. Compensation is structured primarily around annual salary, with the flexibility to determine the desired balance between salary and stock options each year. Comprehensive health insurance plans and mental health support. 401(k) retirement plan with employer matching. Stock option program. Health Savings Accounts and Flexible Spending Accounts. Family-forming benefits. Life and serious injury benefits. Disability programs. Paid leave of absence programs. Flexible paid time off for full-time salaried employees. Remote work opportunity within the United States. Opportunity to work on high-impact AI infrastructure spanning foundation models, evaluation, and data curation. Collaborative environment with substantial technical autonomy and opportunities for senior-level technical leadership.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- AI Researcher — DistillationJobgether
- AI Harness EngineerJobgether
- AI Process Forward Deployed EngineerJobgether
- Account Director, Client SolutionsJobgether
- Account DirectorJobgether
- [People] Hunter / Talent Acquisition AssistantJobgether
- (Senior or Staff) Backend Engineer, AI toolingJobgether
- AI Process Forward Deployed EngineerJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-10-07. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.