Content Specialist III | AI Evaluation & Prompting
Jobgether
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
106 open evaluation roles across 34 companies are on ApplySarthi right now, most of them in Hyderabad (9), Delhi NCR (3), Bengaluru (3).
- Research Fellowship: Agent Intelligence & Evaluationixigo · delhi ncr
- Data selection and quality evaluation for biological foundation modelsInceptive
- Senior Risk Manager - Inventory Trust, Inventory EvaluationAmazon
- Senior Test and Evaluation Laboratory TechnicianBoeing
- AI Engineer 4 (AI Foundations: Benchmarking, Evaluation, and Explainability)Capitalone
What evaluation roles keep asking for: Python (38%), Machine learning (20%), C++ (18%), LLMs (17%), Observability (15%), SQL (15%) — counted across their open postings here.
Jobgether has 4,431 open roles listed here.
- Account Coordinator (Performance Marketing Agency)
- Acute Product Consultant
- Admin/Business Support
- Admin/Business Support
- Admin/Business Support
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for evaluation roles keep coming back to Python, Machine learning, C++, LLMs. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- Walk me through a model you built, from the data to how it was used.
- How did you know your model was actually good, and not just good on your test set?
- Tell me about a time the data was messy or wrong. What did you do?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Content Specialist III | AI Evaluation & Prompting at Jobgether interview free →Accountabilities: Test new AI model versions across a broad range of topics, use cases, and conversation scenarios. Evaluate AI-generated responses against established rubrics, guidelines, and quality standards. Identify, document, and communicate examples of both successful and unsuccessful model behavior. Write, test, and refine system prompts designed to influence model personality, tone, responses, and behavior. Assess whether evaluation rubrics and quality frameworks accurately measure model performance and identify opportunities for improvement. Conduct quality reviews of human- and agent-based evaluations to ensure accuracy, consistency, and compliance. Perform hands-on experiments with AI models and products to investigate capabilities, limitations, and unexpected behaviors. Investigate model failures, identify potential causes or patterns, and document findings clearly. Apply detailed instructions and evaluation criteria consistently across a high volume of work. Perform fact checking and verify model claims against reliable original sources when required. Translate observations and research findings into clear, actionable feedback for product and AI teams. Adapt evaluation approaches as model behavior, product priorities, and quality requirements evolve. Requirements 5+ years of professional experience in writing, editing, journalism, production, linguistics, STEM, coding, policy, or another relevant subject-matter discipline. Bachelor’s degree or equivalent professional experience. At least 1 year of hands-on AI experience is preferred, including prompting, annotation, model evaluation, red teaming, or related work. Strong experience evaluating AI-generated content, large language model outputs, or other complex digital content is highly desirable. Deep expertise in at least one subject area, with the ability to evaluate content accurately as a subject-matter expert. Strong writing, editing, fact-checking, and research capabilities. Experience writing and refining prompts and assessing how prompt changes influence AI behavior. Strong analytical judgment and exceptional attention to detail. Ability to consistently apply detailed rubrics, guidelines, and evaluation criteria across high-volume work. Ability to identify subtle differences in AI responses and determine why a response succeeds or fails. Strong documentation and communication skills, with the ability to clearly explain findings, questions, blockers, and recommendations. Ability to verify claims against original or authoritative sources. Experience conducting detailed failure investigations is preferred. Comfortable working independently while escalating important questions, risks, and blockers early. Highly adaptable and comfortable shifting priorities as AI products and business needs evolve. Professional fluency in English is required. Must be authorized to work in the United States without requiring current or future sponsorship. Benefits 6-month remote contract based in the United States. Full-time schedule of approximately 40 hours per week . Standard schedule of Monday through Friday with 8-hour shifts. Health insurance. Dental insurance. Vision insurance. Health savings account (HSA). Life insurance. Paid time off. Retirement plan. Opportunity to work directly on the evaluation and improvement of advanced AI systems. Hands-on exposure to AI model evaluation, prompt engineering, quality frameworks, and emerging AI products. Opportunity to apply specialized subject-matter expertise to real-world AI development and product improvement.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- AI Researcher — DistillationJobgether
- Art DirectorJobgether
- Applied ML EngineerJobgether
- AI Researcher — DistillationJobgether
- Associate Product Manager, BMO Global Asset ManagementJobgether
- Bilingual Field technology ConsultantJobgether
- Billing Operations LeadJobgether
- Billing Operations LeadJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-10-05. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.