Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations
ServiceNow
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
13 open benchmarking roles across 9 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).
- Benchmarking Program Manager, Cloud Economics ScaleAmazon Web Services, Inc.
- BIE, Speed Benchmarking, CXBT Capability TeamASSPL - Karnataka · bengaluru
- Senior Software Development Engineer in Test (SDET) - BenchmarkingBitwarden · delhi ncr
- Senior Data Center Performance Engineer - Benchmarking and OptimizationNvidia
- Senior Software Engineer, Competitive BenchmarkingMongoDB
What benchmarking roles keep asking for: Data modelling (38%), Excel (31%), Python (31%), Java (23%), Power BI (23%), R (23%), Redshift (23%), Tableau (23%) — counted across their open postings here.
Engineering Manager jobs in the United States · Remote Engineering Manager jobs
ServiceNow has 704 open roles listed here.
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal
- APAC Director, Strategic Partnership for C&I
- Sr. Staff Product Designer, Mobile Experience Strategy & Systems
- Staff Software Engineerhyderabad
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for benchmarking roles keep coming back to Data modelling, Excel, Python, Java. Practise those questions before you sit with ServiceNow.
Questions you are likely to be asked
- Why do you want to join ServiceNow?
- What is your experience with ServiceNow? Tell me one thing you learned the hard way.
- How would you explain your model's result to someone who is not technical?
- What would you check first if a model's accuracy dropped after going live?
- When would you not use machine learning for a problem?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Engineering Manager, Agentic & Generative AI Benchmarking and Evaluations at ServiceNow interview free →The Advanced Technology Group (ATG) at ServiceNow is a customer-focused innovation group building intelligent software and smart user experiences using existing and latest advanced technologies to enable end-to-end, industry-leading work experiences for customers. We are a group of researchers, applied scientists, engineers, and product managers with a dual mission. We build and evolve the AI platform, and partner with teams to build products and end-to-end AI-powered work experiences. In equal measure, we lay the foundations, research, experiment, and de-risk AI technologies that unlock new work experiences in the future. Job Description We are seeking an exceptional, data-driven Senior Engineering Manager, Agentic & GenAI Benchmarking and Evaluations to establish and lead AI evaluation practices for both ServiceNow and our customers. As ServiceNow shifts enterprise workflows from simple generation to complex, autonomous agents, ensuring system reliability, safety, and accuracy is paramount. In this role, you will lead a specialized team of AI evaluation engineers and data scientists. Your team will build the infrastructure, rigorous validation frameworks, and benchmarks that quantify the performance of Now Assist agentic workflows across multi-step orchestration, tool-calling, and enterprise-grounded reasoning. You will bridge the gap between frontier AI research and hard production metrics, directly impacting the trust and adoption of autonomous workflows for millions of enterprise users. What You Get To Do In This Role Build the Evaluation Infrastructure : Design, own, and scale automated testing and evaluation harnesses (unit evals, integration evals, and production drift monitors) to measure agent quality and eliminate regressions. Define Enterprise AI Benchmarks : Create standard, repeatable evaluation frameworks tailored to complex business workflows—assessing multi-agent orchestration, intent routing, multi-step planning loops, and long-term memory accuracy. Validate Grounding & RAG Pipelines : Partner with search and data fabric teams to systematically evaluate Retrieval-Augmented Generation (RAG) pipelines, hybrid search, and semantic re-ranking systems. Model Selection Optimization : Rigorously benchmark frontier LLMs (e.g., OpenAI, Anthropic, Google, and proprietary ServiceNow models) to evaluate trade-offs across execution capabilities, latency, context-window efficiency, and inference costs. Lead a High-Performing Team : Recruit, mentor, and foster an AI-native engineering team, driving engineering best practices, prompt-infrastructure stability, and production-grade rigor. Cross-Functional Leadership : Collaborate with Core Product, Machine Learning Platforms, and Engineering leads to translate baseline performance statistics into actionable product improvements and model fine-tuning targets. To be successful in this role you have: 8+ years of professional software engineering or machine learning experience, including 3+ years managing or technically leading high-performing AI/ML teams. Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry. Strong foundational knowledge of frontier AI SDKs and deep experience deploying or testing agentic/probabilistic software architectures (multi-agent orchestration, tool execution, and probabilistic feature deployment). Demonstrated experience implementing rigorous AI metrics (e.g., ROUGE, BLEU, G-Eval, LLM-as-a-judge patterns, and custom deterministic evaluation code) at an enterprise scale. Proficiency in Python and familiarity with data analytics infrastructures (SQL, Pandas, NumPy) alongside standard MLOps tracking platforms. Experience with complex knowledge infrastructure, SaaS platform architectures, or relational datasets (e.g., Knowledge Graphs, CMDBs). Ability to translate deeply technical evaluation data into executive-level risk assessments, ROI summaries, and strategic roadmap recommendations. Bachelor’s or higher degree in Computer Science, Data Science, Machine Learning, or a highly quantitative field (Master's or Ph.D. is a plus). For positions in this location, we offer a base pay of $201,300 - $352,300 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location. Work Personas We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements. Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance. Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Manager, Product DesignServiceNow · hyderabad
- Director, UX ResearchServiceNow · hyderabad
- Principal Applications Dev EngineerServiceNow · hyderabad
- Staff Data EngineerServiceNow · hyderabad
- Staff Technical Product Manager – AI/LLM expertise + AI Evaluation ScienceServiceNow · hyderabad
- Director - India Reseller ChannelServiceNow · bengaluru
- Senior DevOps EngineerServiceNow · bengaluru
- Staff Data EngineerServiceNow · hyderabad
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on smartrecruiters · posted 2026-07-13. ApplySarthi collects openings and links to application pages; the role is advertised by ServiceNow, not by us.