AI Model Policy Trainer, Image Evaluation - Seattle Onsite
Handshake
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
87 open evaluation roles across 33 companies are on ApplySarthi right now, most of them in Hyderabad (7), Bengaluru (3), Delhi NCR (1).
- Senior Applied AI Engineer (Agents & Evaluation)Gengis AI
- Staff Machine Learning Engineer (TLM), Driver Understanding and EvaluationWaymo
- Ready to Hire Evaluation Analyst, Data center learning Amazon Data Services Ireland Limited
- Senior Medical Writer / Senior Project Manager, Clinical EvaluationAbbott
- Senior Test & Evaluation Engineer (Test Program Req & Planning)Boeing
What evaluation roles keep asking for: Python (20%) — counted across their open postings here.
Handshake has 72 open roles listed here.
- Senior Software Engineer, Coding
- Mid-Market Customer Success Manager
- Director of Internal Communications
- Senior Product Marketing Manager, Handshake AI
- Member of Technical Staff, Post-Training
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for evaluation roles keep coming back to Python. Practise those questions before you sit with Handshake.
Questions you are likely to be asked
- Why do you want to join Handshake?
- What is your experience with Workday? Tell me one thing you learned the hard way.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
- How would you explain your model's result to someone who is not technical?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the AI Model Policy Trainer, Image Evaluation - Seattle Onsite at Handshake interview free →About Handshake Handshake was founded on a simple belief that everyone deserves a path to a great career, regardless of where they went to school or who they know. Today, we power 25 million job seekers, 1 million+ employers, and 1,600 educational institutions. Handshake AI works directly with frontier AI lab researchers to create evaluations, publish benchmarks, and improve AI models through human expertise. Role Details Location: Onsite in Seattle, WA, Monday-Friday. Compensation: $36-$72/hr. Placement within the range depends on experience. Employer: TCWGlobal. This is a W-2 assignment supporting Handshake AI. Employment: Full time, 40 hours per week, non-exempt and eligible for overtime pay. Schedule: Monday-Friday, 8 a.m.-5 p.m. PT. About the Role As an AI Image Evaluator, you will help image generation models learn two things at once: what a good image is, and what an acceptable image is. You will look at prompts and the images a model produced from them, then answer questions like: Did the image do what the prompt asked? Is it well made, or does it have the hands, lighting, text, and anatomy problems that give generated images away? Which of two images is better, and why? Does the image violate the customer's content policy, and if so, which category and how severely? Does it depict a real person, a protected brand, or a minor in a way the policy does not allow? The interesting cases are the close ones. Two images that look nearly identical until you notice one has a logo in the background. A stylized nude that is fine as figure study and not fine with one change of pose. A prompt that asked for "a realistic photo of a senator" and a model that complied. A beautiful image that ignored half the prompt, next to an ugly one that nailed it. We are looking for people who already see images critically, whether that came from photography, illustration, design, years inside Midjourney and Stable Diffusion, or moderating visual content at scale. You do not need all of these. You need one deep, and the judgment to learn the rest. This is not rote annotation. Rubrics cannot anticipate every image, and good evaluators do not apply them mechanically. You will balance the rubric's text and intent with customer expectations, precedent, and team calibration, and you will explain your reasoning clearly enough that it can train a model. What You Will Do Evaluate generated images against their prompts for adherence, composition, realism, style consistency, and technical defects; maintain accuracy and consistency across repeated evaluations Compare images side by side and select the stronger one with a clear, evidence-based rationale Classify images against customer content policies covering sexual content, violence, hate symbols, real-person likeness, intellectual property, and depictions of minors Select the most defensible classification when an image is genuinely ambiguous, and write concise rationales that cite rubric language and specific visual details Distinguish "I do not like this" from "this fails the prompt" from "this violates policy," and keep those judgments separate while applying customer policy consistently Write and refine prompts that probe where a model's quality or safety behavior breaks down Identify rubric gaps, contradictions, and emerging edge cases, and raise them with project leads and policy teams Participate actively in calibration discussions; challenge interpretations respectfully and update your judgment when stronger reasoning emerges You May Be a Fit If You have a trained eye from photography, illustration, concept art, art direction, retouching, photo editing, VFX, or visual design, and you can say precisely why one image is better than another You use generative image tools heavily (Midjourney, Stable Diffusion, ComfyUI, Flux, DALL-E, Ideogram) and know their failure modes, their prompt quirks, and how their safety filters get bypassed You have moderated or reviewed visual content at scale and have applied a policy taxonomy to borderline images under time pressure You notice small details: an extra finger, a mismatched shadow, a brand mark, a face that is a little too familiar You can hold a rubric steady across a long session of near-identical images and treat sensitive material with maturity and sound judgment You can hold a strong opinion without becoming attached to being right You explain judgment calls clearly and precisely in writing so another person can audit your reasoning You can separate your personal taste from the standard a customer has asked you to apply Strong candidates may come from photography, illustration, graphic or UX design, art direction, photo editing, animation or VFX, game art, trust and safety, content moderation, brand or IP enforcement, ad review, or art education. We care more about how you see and how you reason than where you learned to do it. A degree and a technical background are not required. Nice to Have A public portfolio, publication credits, or a body of generative work (Civitai, Discord communities, LoRA or model training, published prompt work) Experience judging images comparatively: portfolio review, photo competition judging, creative A/B testing, art school critique Formal training in anatomy, color, lighting, or composition Content moderation or trust and safety experience on an image-heavy platform Working knowledge of copyright, trademark, and right-of-publicity basics Prior work in AI evaluation, RLHF, image labeling, or data annotation Familiarity with calibration sessions, inter-rater agreement, or adjudication workflows Sensitive-Content Notice This role involves regular and deliberate engagement with sensitive imagery. Depending on the project, evaluations may include sexual content and nudity, graphic violence and gore, hate symbols, self-harm, and depictions of real people and of minors in contexts that must be assessed against policy. Some of this material is disturbing by design, because the purpose of the work is to teach models not to produce it. The work is conducted within structured evaluation frameworks and professional guidelines, with exposure limits, content rotation, mandatory reporting protocols for illegal material, and access to mental health support. Candidates must be able to engage with this material carefully, responsibly, and sustainably while maintaining sound judgment and consistent work quality. Benefits Benefits are provided through TCWGlobal. Eligible employees can enroll in medical coverage, including prescription and mental health benefits, dental and vision insurance, healthcare and dependent care flexible spending accounts, and pretax commuter benefits. A 401(k) retirement plan with an employer match is available to eligible participants. Additional offerings include voluntary accident, critical illness, and term life insurance, wellness and pet-related reimbursements, charitable matching, and employee discounts. Eligibility requirements, waiting periods, employee contributions, and plan terms apply. Full-time employees scheduled for at least 30 hours per week are eligible for health coverage beginning the first of the month following at least 30 days of employment. Retirement plan eligibility follows a separate schedule. Paid time off: PTO accrues from the first day of the assignment at one hour per 40 hours worked, with no waiting period to use accrued PTO. PTO may be used for vacation, personal days, or illness, subject to policy approval requirements. Accrual is capped at 80 hours per year and balances at 100 hours, unless otherwise required by law. Additional paid sick leave is provided in accordance with applicable Washington and Seattle requirements. Paid holidays: The holiday policy lists 13 paid holidays. Eligible employees receive eight hours of holiday pay at their base hourly rate for observed holidays that fall on a scheduled workday, subject to the holiday policy. See the TCWGlobal 2026 Benefits Guide for coverage options, costs, and enrollment details. Interview Process The process includes application review, a recruiter screen, a skills assessment, a hiring manager screen, onsite interviews, and a final round before the offer stage. Your recruiter will share the next steps as you move through the process. Equal Opportunity and Accommodations TCWGlobal is an equal opportunity employer. Hiring decisions are based on qualifications and abilities, without discrimination based on characteristics protected by applicable law. Reasonable accommodations are available during the application and interview process. If you need an accommodation, please let your recruiter know.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Manager, Strategic ProjectsHandshake
- Strategic Projects LeadHandshake
- Software Engineer II, Reinforcement Learning EnvironmentsHandshake
- Data Analyst, Finance and PaymentsHandshake
- Member of Technical Staff, Post-TrainingHandshake
- Forward Deployed Engineer IHandshake
- Senior Software Engineer, International ExpansionHandshake
- Senior Forward Deployed EngineerHandshake
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on ashby · posted 2026-09-02. ApplySarthi collects openings and links to application pages; the role is advertised by Handshake, not by us.