ApplySarthi

Agentic AI Engineer - Gandiva.ai + Colosseum

Gilgamesh

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

CRM jobs · LLMs jobs · Observability jobs · Python jobs

Gilgamesh has 2 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for colosseum roles keep coming back to CRM, LLMs, Observability, Python. Practise those questions before you sit with Gilgamesh.

Questions you are likely to be asked

  1. Why do you want to join Gilgamesh?
  2. What is your experience with CRM? Tell me one thing you learned the hard way.
  3. When would you not use machine learning for a problem?
  4. Walk me through a model you built, from the data to how it was used.
  5. How did you know your model was actually good, and not just good on your test set?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Agentic AI Engineer - Gandiva.ai + Colosseum at Gilgamesh interview free →

About **Gilgamesh**: Gilgamesh builds systems for operators to plan, execute, compete, and prove results. Gandiva.ai is a go-to-market operating system that connects strategy to advertising, CRM, and analytics tools, then adapts execution to performance signals. Colosseum is a competitive platform where operators complete real-world briefs, receive AI- and practitioner-assisted verification, and earn current rankings through demonstrated work. ## The role We’re looking for an Agentic AI Engineer to build the systems that turn goals into reliable, observable work across Gandiva.ai and Colosseum. You’ll design AI workflows that can interpret intent, plan bounded tasks, use external tools, evaluate results, and recover safely when something goes wrong. You’ll work across the agent runtime and product features: growth workflows for Gandiva.ai, and evidence-based assessment workflows for Colosseum. This is a hands-on engineering role for someone who can take an ambiguous product problem from prototype to a system people can trust. ## What you’ll do - Design and ship agentic workflows that combine language models, business rules, APIs, and deterministic code. - Build the orchestration layer: typed tools, task state, retries, checkpoints, approval steps, and resumable workflows. - For Gandiva.ai, connect agents to advertising, CRM, and analytics systems so they can turn GTM goals into plans, prepare or execute changes within granted permissions, read performance signals, and recommend or take the next action. - Put clear controls around consequential actions such as publishing content, changing budgets, or contacting prospects. Define permissions, approval thresholds, audit trails, and recovery paths with the product team. - For Colosseum, build AI-assisted workflows that interpret contest briefs, inspect submissions and supporting evidence, apply explicit rubrics, and present useful reasoning to human reviewers. - Make verification and ranking systems evidence-linked and reviewable. Measure where AI assessments agree or disagree with practitioner judgments, and help reviewers resolve edge cases. - Build evaluation datasets and repeatable tests for task completion, tool-call correctness, rubric consistency, instruction following, latency, and cost. - Instrument agents with traces and useful operational signals so the team can find failed steps, understand decisions, and improve quality over time. - Protect user and company data across prompts, model calls, logs, memory, and integrations. Account for prompt injection, untrusted submissions, secret handling, and tenant boundaries. - Work closely with product, GTM, creative, and practitioner teams to turn real workflows into clear requirements and ship improvements in short cycles. ## What we’re looking for - You have built and shipped software that uses LLMs or other AI models to complete multi-step tasks, call tools, or automate real workflows. - You are strong in Python or TypeScript and comfortable building backend services, APIs, asynchronous jobs, and integrations. - You understand the practical parts of agent systems: tool schemas, context and state, retries, failure handling, human review, and choosing deterministic code where it is more reliable. - You know how to evaluate AI features with task-specific examples and measurable criteria, not just prompt intuition. - You can work with relational data and production infrastructure, and you pay attention to access control, observability, reliability, and cost. - You communicate clearly with both technical teammates and operators, and can explain what an AI workflow did, why it did it, and where it needs human judgment. - You take ownership from problem definition through deployment and iteration. ## Helpful experience - Integrating advertising platforms, CRMs, analytics, outbound, or content tools. - Building evaluation, review, ranking, or scoring workflows where outputs need to be checked against evidence and a rubric. - LLM tracing and evaluation tools, retrieval systems, durable workflow engines, queues, or event-driven architectures. - Designing safe permission models for AI systems that can affect external accounts or customer-facing content. - Working in an early-stage product team where you’ve had to make sound trade-offs and deliver with incomplete information. ## What success looks like - Gandiva.ai users can move from a GTM objective to useful, measurable execution with appropriate visibility and control over agent actions. - Colosseum reviewers can assess submissions faster while seeing the evidence behind AI-assisted evaluations and retaining a clear path to correct them. - Agent workflows have repeatable evaluations, useful failure signals, and clear recovery behavior. - The team can expand integrations and capabilities without losing control of reliability, data access, or operating cost. ## How to apply Share your résumé or profile, links to relevant work, and a short example of an agent you shipped. Tell us what the system was allowed to do, how you evaluated it, and what you learned from failures.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on wellfound · posted 2026-09-28. ApplySarthi collects openings and links to application pages; the role is advertised by Gilgamesh, not by us.