ApplySarthi Match jobs to your CV

Software Engineering Director, Agentic Evaluations

Jobgether

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

5,603 open engineering roles across 519 companies are on ApplySarthi right now, most of them in Bengaluru (299), Hyderabad (110), Chennai (55).

What engineering roles keep asking for: AWS (17%), Python (13%) — counted across their open postings here.

AWS jobs · FastAPI jobs · Go jobs · Java jobs

Jobgether has 3,934 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for engineering roles keep coming back to AWS, Python. Practise those questions before you sit with Jobgether.

Questions you are likely to be asked

  1. Why do you want to join Jobgether?
  2. What is your experience with AWS? Tell me one thing you learned the hard way.
  3. Walk me through a system you built. How was it designed, and what would you change now?
  4. Tell me about a hard bug you tracked down. How did you find the cause?
  5. How do you decide what to test, and what does good code review look like to you?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Software Engineering Director, Agentic Evaluations at Jobgether interview free →

Accountabilities:: Lead the development and continuous improvement of evaluations that run against live software agents, identifying practical approaches for producing useful, credible, and repeatable results. Own the full technology stack supporting the evaluation platform, including administrative tooling, APIs, evaluation workflows, orchestration, and production systems. Design reusable evaluation primitives and architecture that can be applied across multiple software categories and reduce the effort required to onboard new agent types. Turn recurring integration, onboarding, and maintenance processes into reusable AI skills, agents, or automated workflows that increase evaluation velocity. Monitor emerging frameworks, methodologies, and industry practices for agent evaluation, identifying opportunities to improve measurement quality and distribution of results. Partner closely with data science teams to develop proprietary benchmarks and richer approaches to evaluating agent performance. Lead, mentor, and develop a core engineering team, providing technical direction and subject-matter expertise in AI agent evaluations. Promote the effective use of evaluations across agent-focused products and initiatives throughout the wider organization. Requirements 10+ years of professional software development experience, primarily in backend or full-stack environments. 2+ years of direct engineering management experience, including team leadership, mentoring, and technical direction. Expert-level backend development skills in languages such as Python, Java/Kotlin, TypeScript/JavaScript, or Go, with strong experience in frameworks such as FastAPI or Node.js. Hands-on experience building evaluations for customer-facing AI agents, including the use of agent trajectory traces and evaluation rubrics to measure task completion, accuracy, correctness, and/or policy adherence. Direct experience using frontier models from providers such as OpenAI, Anthropic, or Google in LLM-as-a-judge applications. Regular experience using coding-agent tools such as Claude Code, Codex, OpenCode, or Pi as part of a modern software development workflow. Bachelor’s degree in computer science, engineering, or a related field. Experience integrating agent tool use directly or through MCP servers is a plus. Familiarity with benchmark frameworks such as STATE-Bench, tau2-bench, or similar evaluation approaches is desirable. Experience using browser automation and agent tooling such as Playwright, browser-use, or Chrome DevTools MCP is advantageous. Experience with durable execution and workflow platforms such as Temporal, DBOS, Cloudflare Workflows, or Vercel Workflows is helpful. Familiarity with agent sandboxing technologies such as AWS E2B, Daytona, Cloudflare Containers, or Vercel Containers is also valuable. Strong communication, technical leadership, mentoring, and cross-functional collaboration skills are essential. Benefits Total earnings of approximately $260,000–$320,000, combining base salary and bonus. Equity participation. Performance-based bonus opportunities. Fully remote position within the United States. Flexible working environment designed to support distributed teams. Unlimited paid time off. Generous parental leave. Comprehensive benefits designed to support employee well-being and flexibility. Inclusive and diverse workplace with employee-led community initiatives and professional growth opportunities.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on lever · posted 2026-09-26. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.