Senior Software Engineer, Agent Eval Platform
ServiceNow
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
11,666 open software roles across 746 companies are on ApplySarthi right now, most of them in Bengaluru (794), Hyderabad (300), Pune (179).
- Software Development Engineer, AWS SecurityAmazon Development Centre (London) Limited
- Lead Software Engineer - JavaJPMorgan · bengaluru
- Software Engineer (New Graduates)DevsUnite · delhi ncr
- ETIC, Software Testing, ManagerPwc
- Staff Software Engineer, Web Application ServicesMozilla
What software roles keep asking for: AWS (25%), Python (23%), Java (23%), System design (20%), Kubernetes (18%), C++ (17%), Observability (16%), CI/CD (13%) — counted across their open postings here.
Software Engineer jobs in the United States · Software Engineer jobs in Mountain View · Remote Software Engineer jobs
ServiceNow has 704 open roles listed here.
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
- APAC Director, Strategic Partnership for C&I
- Sr. Staff Product Designer, Mobile Experience Strategy & Systems
- Staff Software Engineerhyderabad
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for software roles keep coming back to AWS, Python, Java, System design. Practise those questions before you sit with ServiceNow.
Questions you are likely to be asked
- Why do you want to join ServiceNow?
- What is your experience with ServiceNow? Tell me one thing you learned the hard way.
- Walk me through a system you built. How was it designed, and what would you change now?
- Tell me about a hard bug you tracked down. How did you find the cause?
- How do you decide what to test, and what does good code review look like to you?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Software Engineer, Agent Eval Platform at ServiceNow interview free →The Role Moveworks' AI agents don't just generate text — they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did — across a multi-step trajectory through a world it changed — precisely enough that the score can teach it to do better? That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card — a judge good enough to grade a trajectory is a judge good enough to train against . The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing. This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments. What you get to do in this role: We're hiring across three areas. You'll anchor on one and touch the others; which one is a conversation we have with you, not a slot we drop you into. Eval orchestration at scale The runtime that executes multi-turn agent scenarios end-to-end — stand up the environment and user simulator, drive the user↔agent↔world loop, collect transcripts, traces, and final state, run validators and scoring, tear down Scheduling, retries, high-concurrency execution, and run isolation at production dataset sizes Versioned specs, datasets, and reports, with run-to-run comparison as a first-class operation Consolidating evals that run today as one-off workflows onto a single orchestration service — one source of truth, one place to schedule and retry Establishing a reliability floor and an SLO for the harness itself Getting to self-serve, so any team runs an eval without bespoke integration Agent observability and tracing Leading the move to OpenTelemetry-native observability for the agent platform, replacing the parallel per-service logging, correlation, and redaction mechanisms in use today The span data model for agent trajectories — prompts, tool calls, plan updates, outcomes — so a trajectory is queryable , not reconstructed by hand from log files Trace context propagation across async boundaries and sessions that stay alive for minutes or hours Making full prompts and completions survive the pipeline intact, and keeping eval traffic from contaminating its own data Fault attribution and cross-run diffing: which component actually broke, and what changed since the last green run The debug surface support and harness engineers use, and the tracing contract with the team that builds the agent Stateful simulation The simulation environment itself: stateful fakes of the enterprise systems agents call — ITSM, HR, knowledge bases, inventory — backed by a real datastore that persists changes during a run, so a created ticket is visible to a later read Per-run data injection and programmatic setup/teardown so every run is hermetic and repeatable LLM-driven user simulators for open-ended personas, and scripted state-machine simulators for deterministic flows Contract-testing mocks against real API schemas in CI, so simulation fidelity can't quietly drift as vendor APIs change Ahead of us: isolated sandbox environments reproducing the config, identity, search content, and permissions an agent actually reads — provisioned from an identical baseline and torn down every run And across all three: laying the foundation for using eval signal to optimize the agent, not just measure it. To be successful in this role you have: Experience in at least 3 of these: Distributed systems: idempotency, delivery guarantees, isolation, and — unusually central here — determinism and reproducibility Orchestration and workflow runtimes: DAG execution, scheduling, retries, backfills, high-concurrency job systems (Temporal, Airflow, Argo, or something you built yourself) Observability internals as a builder , not just a user: OpenTelemetry SDKs and collectors, semantic conventions, span context propagation, high-cardinality trace data Concurrent and async programming: Python asyncio, Go concurrency, structured cancellation Data-intensive pipelines: high-volume ingest, schema evolution, sampling and retention trade-offs gRPC/protobuf service and interface design Required: 5+ years building production backend or infrastructure systems Strong in Python or Go (ideally both) Experience designing and operating systems that handle real traffic at scale Comfort making a non-deterministic system measurable. You don't need an ML background — but you should find it interesting to turn fuzzy agent behavior into a signal engineers are willing to gate releases on Comfort with ambiguity; these are novel problems without textbook solutions For positions in this location, we offer a base pay of $161,300-274,200, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location. Work Personas We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, national origin or nationality, ancestry, age, disability, gender identity or expression, marital status, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance. Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. From Fortune. ©2025 Fortune Media IP Limited. All rights reserved. Used under license.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Manager, Product DesignServiceNow · hyderabad
- Director, UX ResearchServiceNow · hyderabad
- Principal Applications Dev EngineerServiceNow · hyderabad
- Staff Data EngineerServiceNow · hyderabad
- Staff Technical Product Manager – AI/LLM expertise + AI Evaluation ScienceServiceNow · hyderabad
- Director - India Reseller ChannelServiceNow · bengaluru
- Senior DevOps EngineerServiceNow · bengaluru
- Staff Data EngineerServiceNow · hyderabad
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on smartrecruiters · posted 2026-08-13. ApplySarthi collects openings and links to application pages; the role is advertised by ServiceNow, not by us.