Senior Data Scientist AI Evaluation
Jobgether
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
118 open evaluation roles across 37 companies are on ApplySarthi right now, most of them in Hyderabad (9), Bengaluru (5), Delhi NCR (3).
- Robotics Simulation and Evaluation PhD Intern - Summer 2027Nvidia
- Executive Director, Search and Evaluation OncologyAstrazeneca
- Senior Data Analyst - AI Evaluation (Life Sciences).Relx · chennai
- Legal Expert - AI Training & EvaluationWeekdayworks
- Associate Test & Evaluation TechnicianBoeing
What evaluation roles keep asking for: Python (45%), Machine learning (25%), LLMs (24%), C++ (23%), Observability (19%), SQL (15%), Java (13%), PyTorch (13%) — counted across their open postings here.
LLMs jobs · Machine learning jobs · Python jobs · SQL jobs
Jobgether has 4,311 open roles listed here.
- Account Operations & Farming Specialist
- .NET Backend Developer Pleno
- .Net Developer
- .Net Developer
- .Net Developer
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for evaluation roles keep coming back to Python, Machine learning, LLMs, C++. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- When would you not use machine learning for a problem?
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Data Scientist AI Evaluation at Jobgether interview free →Accountabilities: AI Evaluation Design: Define ground truth, quality metrics, scoring methodologies, and evaluation frameworks to assess the accuracy, consistency, and reliability of AI models and agents. Evaluation Pipeline Development: Build repeatable evaluation loops that monitor model quality over time, identify performance regressions, and support reliable release decisions. Statistical Measurement and Validation: Apply rigorous statistical methods to evaluation design, including sample sizing, confidence intervals, significance testing, and techniques for handling non-deterministic model outputs. Automated Grader Validation: Compare automated evaluation methods with human assessments to establish reliability, identify limitations, and improve scoring accuracy. Evaluation Infrastructure Collaboration: Partner with Engineering and Analytics Engineering to operationalize evaluation frameworks, integrate testing into development workflows, and support scalable evaluation infrastructure. Performance Analysis and Improvement: Interpret evaluation results, identify weaknesses and quality gaps, and provide actionable recommendations to improve model and agent performance. Quality Standards and Documentation: Establish evaluation guidelines, documentation practices, review processes, and consistent quality standards across AI development initiatives. Cross-Functional Partnership: Collaborate with Product, Engineering, Analytics Engineering, and business stakeholders to define success criteria and align evaluation priorities with business needs. Continuous Improvement and Mentorship: Promote evaluation best practices, share technical insights, and help foster a culture of measurable AI quality, accountability, and evidence-based decision-making. Safe AI Deployment: Support teams in making informed release decisions through independent, reliable assessments of model performance and readiness. Requirements: Approximately 6–10 years of experience in quantitative data science, machine learning, or a related technical discipline, with focused experience in measurement, evaluation, experimentation, or model validation. Strong quantitative and statistical foundations, including experience designing rigorous experiments and interpreting results with appropriate uncertainty measures. Demonstrated experience defining meaningful evaluation metrics, establishing ground truth, and assessing the quality of complex or ambiguous model outputs. Proficiency in Python and SQL, with experience applying these skills to data analysis, model assessment, and evaluation workflows. Experience evaluating machine learning models in production environments and translating findings into practical improvements. Ability to validate automated graders against human judgments and account for variability, non-determinism, and potential measurement bias. Strong problem-solving and analytical judgment, particularly when developing evaluation approaches for new or evolving AI systems. Excellent communication skills, with the ability to explain evaluation results, trade-offs, and recommendations to technical teams and business stakeholders. Proven ability to collaborate across functions while independently owning complex analytical projects in a fast-paced environment. A quantitative degree in data science, statistics, mathematics, computer science, or a related field is an advantage; equivalent practical experience is also welcome. Hands-on experience evaluating large language models (LLMs) or AI agents in production, including evaluation harnesses, LLM-as-judge calibration, and continuous integration regression gates, is a plus. Experience evaluating text-to-SQL systems, analytics agents, or other AI applications where outputs can be verified against underlying data is an advantage. Background in fintech, brokerage, financial services, or other domains where incorrect outputs can create significant business or risk consequences is desirable. Familiarity with AI tools used in research, analysis, and engineering workflows is beneficial. Benefits: Competitive compensation: Salary package designed to reflect experience and expertise. Stock options: Opportunity to participate in the company's long-term growth. Health benefits: Benefits designed to support employee health and well-being. Home-office setup allowance: One-time allowance of USD $500 to support your remote workspace. Monthly stipend: USD $150 per month provided through a Brex card. Remote work environment: Opportunity to work remotely within the Americas, with Canada as the target location for this listing. Technical ownership: Lead the development of evaluation practices and establish standards that influence AI quality across the organization. Cross-functional impact: Work closely with product, engineering, analytics, and business teams to turn rigorous measurement into better AI systems. Professional development: Build expertise in AI evaluation, model validation, and the responsible deployment of intelligent systems in a growing technology environment.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- AI Process Forward Deployed EngineerJobgether
- Account DirectorJobgether
- [People] Hunter / Talent Acquisition AssistantJobgether
- AI Research Engineer (Kernel & Inference Optimization)Jobgether
- Account Manager (Email Marketing)Jobgether
- Advogado(a) | BancárioJobgether
- Agentic Workforce Adoption ManagerJobgether
- AI EngineerJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-10-09. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.