Staff Machine Learning Engineer, Agent Eval Platform
ServiceNow
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
1,084 open learning roles across 267 companies are on ApplySarthi right now, most of them in Bengaluru (63), Hyderabad (24), Delhi NCR (14).
- AI Engineer - Imitation Learning (Senior)Rivr
- Machine Learning EngineerDeepJudge
- Machine Learning Platform EngineerBjak
- PhD Autonomy Engineer Intern - Computer Vision / Deep Learning Summer 2027Skydio
- Senior Backend Engineer: Machine Learning InfrastructureConstructor
What learning roles keep asking for: Machine learning (48%), Python (36%), LLMs (22%), PyTorch (22%), Deep learning (16%), AWS (13%), Generative AI (13%) — counted across their open postings here.
Machine Learning Engineer jobs in the United States · Machine Learning Engineer jobs in Mountain View · Remote Machine Learning Engineer jobs
ServiceNow has 704 open roles listed here.
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
- APAC Director, Strategic Partnership for C&I
- Sr. Staff Product Designer, Mobile Experience Strategy & Systems
- Staff Software Engineerhyderabad
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for learning roles keep coming back to Machine learning, Python, LLMs, PyTorch. Practise those questions before you sit with ServiceNow.
Questions you are likely to be asked
- Why do you want to join ServiceNow?
- What is your experience with ServiceNow? Tell me one thing you learned the hard way.
- How would you explain your model's result to someone who is not technical?
- What would you check first if a model's accuracy dropped after going live?
- When would you not use machine learning for a problem?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Staff Machine Learning Engineer, Agent Eval Platform at ServiceNow interview free →The Role Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better? That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against . The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing. This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments. What you get to do in this role: Judge design and calibration A shared base judge with per-item rubrics expressed as configuration next to the dataset — so eval authors express intent, rather than forking a prompt per eval Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy — was the clarifying question appropriate, was policy followed, was the path efficient Scoring that reports its own confidence , so uncertain judgements route to a human instead of quietly becoming training data A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team — they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job Fine-tuning a small judge model where an off-the-shelf one isn't good enough Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit Self-learning for the agent harness This is where the pillar is headed, and a large part of why the seat exists. A calibrated trajectory judge is, functionally, a reward model . Turning ours into a process reward model — a dense, step-level signal for what a good agent trajectory looks like — is the unlock Using that signal to optimize the agent itself : prompts, tool selection, planner behavior, retrieval, routing — tuned against simulation rather than against production traffic Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success To be successful in this role you have: 8+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used Experience turning subjective human judgement into a measurement that holds up — one that other people, and ideally other models, can act on. This is the core of the job Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain Strong Python, and the discipline to ship production-grade code rather than notebooks Ability to think and communicate clearly about complex problems — a large part of this job is convincing engineers that a number means what you say it means, and being right A high degree of ownership and a bias toward shipping at startup pace Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on Experience in at least 3 of these: LLM-as-judge or automated evaluation design, and calibrating it against human judgement Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint Search ranking, recsys, or online experimentation evaluation — golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly Fine-tuning and evaluating small models: SFT, preference tuning, distillation Reward modeling, RLHF/RLAIF, or process reward models Agent trajectory analysis and step-level fault attribution Prompt engineering as an engineering discipline — versioned, tested, and measured, not tuned by vibes Work Personas We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance. Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Manager, Product DesignServiceNow · hyderabad
- Director, UX ResearchServiceNow · hyderabad
- Principal Applications Dev EngineerServiceNow · hyderabad
- Staff Data EngineerServiceNow · hyderabad
- Staff Technical Product Manager – AI/LLM expertise + AI Evaluation ScienceServiceNow · hyderabad
- Director - India Reseller ChannelServiceNow · bengaluru
- Senior DevOps EngineerServiceNow · bengaluru
- Staff Data EngineerServiceNow · hyderabad
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on smartrecruiters · posted 2026-08-26. ApplySarthi collects openings and links to application pages; the role is advertised by ServiceNow, not by us.