Member of Technical Staff, AI Evaluation
P-1 AI
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
This role on the market
113 open evaluation roles across 36 companies are on ApplySarthi right now, most of them in Hyderabad (9), Bengaluru (5), Delhi NCR (3).
- Legal Expert - AI Training & EvaluationWeekdayworks
- Associate Director, Clinical AI Evaluation and Responsible DeploymentAstrazeneca
- Associate Design Evaluation EngineerAnalogdevices
- Manager II, Program Management, AI Model Evaluation, International Seller Growth - ISGAmazon · bengaluru
- AI Engineer 5 (Gen AI Platform Services: Agentic AI, Guardrails, Evaluation)Capitalone
What evaluation roles keep asking for: Python (44%), Machine learning (24%), C++ (23%), LLMs (22%), Observability (19%), SQL (15%), Java (12%), System design (12%) — counted across their open postings here.
Member of Technical Staff jobs in the United States · Remote Member of Technical Staff jobs
P-1 AI has 4 open roles listed here.
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for evaluation roles keep coming back to Python, Machine learning, C++, LLMs. Practise those questions before you sit with P-1 AI.
Questions you are likely to be asked
- Why do you want to join P-1 AI?
- When would you not use machine learning for a problem?
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Member of Technical Staff, AI Evaluation at P-1 AI interview free →**TL;DR:** If you: - have mastered extreme systems in a physical engineering domain to know what peak engineering actually looks like; - have built or worked with AI evaluations, including judges and human evaluation; - can work comfortably with customers, subject matter experts, partners, and technical teams; - are motivated by the goal of building superintelligence for engineering… … you should apply for this role! **About P-1 AI:** At P-1 AI, we are building an AI engineer agent for the physical world named Archie. We maximize Archie’s anthropomorphism so that he fits seamlessly into existing engineering teams and workflows in the form factor of a human engineer. Archie today is at the level of a junior mechanical and electrical engineer, with a quantitative intuition over the product design space and the ability to use complex engineering tools—the same tools his human teammates use. Archie's tech stack includes a custom agentic harness, structured design representation, continual skills learning, and small custom post-trained models (SFT and RLVR) using proprietary semi-synthetic training data sets and environments which create a deep competitive moat. Our ultimate aim is to build engineering ASI. We recently announced a $50 million Series A financing led by NEA, which added former General Electric CEO Jeff Immelt to our board. The round also included the addition of several AI luminaries from Anthropic and Nominal to our existing angel investors from Google and OpenAI. **About the opportunity:** In this role, you’ll sit at the forefront of AI for physical engineering, working with subject matter experts, customers, partners, and our internal teams to challenge and evaluate Archie on tough engineering work, understand where it fails, and capture those failures in evaluations that guide product development. You need to understand not only whether outlook is plausible, but whether the underlying engineering reasoning and decisions make sense. **About the role:** - Study how engineers and customers use Archie in the field, identify where Archie succeeds or fails from an engineering perspective, and turn those observations into actionable evaluations. - Recreate real-world engineering tasks and failure modes in environments that can be repeatedly used by our development teams. - Build and refine evaluation methodologies for Archie, including automated judges and human evaluation processes. - Assess whether automated judges actually align with expert human judgment and improve them when they do not. - Work across engineering, product, customers, SMEs, and external partners to make sure our evaluations reflect real engineering workflows rather than abstract benchmarks. - Use evaluation results to help the team understand where Archie needs to improve and which problems matter most to users. **About you:** - Technical background in a physical engineering domain such as power systems, power electronics, mechanical engineering, aerospace, robotics, or a closely related field. - Strong enough understanding of agentic AI systems to design and reason about evaluations. - Experience with AI evaluations, including creating evaluation tasks, designing or crafting judges, and comparing automated evaluation against human judgment. - Comfortable working directly with customers, engineers, subject matter experts, partners, and internal technical teams to extract what matters from ambiguous workflows. - Strong product judgment, with the ability to turn observations about how a system is being used or failing into concrete evaluation work. - Motivated by the mission of building superintelligence for engineering and excited by the challenge of making AI systems genuinely capable in the physical world. **Location:** Remote (US/Canada) or San Mateo, CA. Remote employees spend one week out of six working together on-site in our San Mateo office. Relocation support available. **Benefits:** Competitive salary, meaningful equity ownership, healthcare, dental, vision, 401(k) match, and unlimited PTO. **Interview process:** - Introductory call (30 mins) - Biographical/behavioural interview (45 mins) - Technical interview (60 mins) - CEO interview (30 mins)
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Full Stack Super SWEP-1 AI
- Customer Success ManagerP-1 AI
- Director of ProductP-1 AI
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on wellfound · posted 2026-10-06. ApplySarthi collects openings and links to application pages; the role is advertised by P-1 AI, not by us.