ApplySarthi Match jobs to your CV

Founding AI Research Engineer: Agent Harness & Model Systems

Astoria AI

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

114 open founding roles across 67 companies are on ApplySarthi right now, most of them in Bengaluru (13), Delhi NCR (3), Mumbai (2).

What founding roles keep asking for: LLMs (34%), Python (27%), PostgreSQL (25%), TypeScript (24%), React (23%), SaaS (23%), CI/CD (21%), Node.js (19%) — counted across their open postings here.

Observability jobs · Workday jobs

Astoria AI has 3 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for founding roles keep coming back to LLMs, Python, PostgreSQL, TypeScript. Practise those questions before you sit with Astoria AI.

Questions you are likely to be asked

  1. Why do you want to join Astoria AI?
  2. What is your experience with LLMs? Tell me one thing you learned the hard way.
  3. Tell me about a time the data was messy or wrong. What did you do?
  4. How would you explain your model's result to someone who is not technical?
  5. What would you check first if a model's accuracy dropped after going live?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Founding AI Research Engineer: Agent Harness & Model Systems at Astoria AI interview free →

**COMPENSATION — READ THIS FIRST ** Before launch and seed round (approx. months 4–6): equity-only. A founding-team equity grant negotiable based on experience, scope, and contribution. Standard vest and cliff. After launch and seed round: competitive base salary in the $220k–$320k range, ongoing equity participation through refresh grants, and other benefits. Why we lead with this: we value your and our time. We are looking for people who want to make their mark. Astoria is building the most ambitious product in the category. The pre-launch and pre-seed round phase is founder-mode: small team, high intensity, outsized equity. This role is for someone who sees the equity as the primary upside — not a supplement — because they believe what we're building will define the category. **The 90-Second Version ** We are in stealth and will help you with your diligence. Astoria AI is building an agentic runtime for human potential. Multiple specialized AI agents serve candidates and companies across the entire hiring and talent lifecycle: career strategy, networking intelligence, verified matching, compliance-native hiring, onboarding, engagement, retention, and workforce planning. We are attacking a market that is structurally broken in ways that damage real people. 44% of resumes contain fabrications. ATS keyword filters reject 88% of qualified candidates. 27% of posted jobs are ghost jobs. Meanwhile the regulatory wave — Mobley v. Workday, the EU AI Act, FCRA lawsuits against Eightfold — is forcing every enterprise buyer to ask whether their hiring AI creates legal risk or eliminates it. Nobody has fixed this. The incumbents can't, because their business models depend on the broken parts. That is the opening. Here is the technical thesis behind this role. Multiple agents making consequential decisions about people's careers do not become trustworthy because the underlying model is clever. They become trustworthy because of the harness they run inside — the runtime that governs how they call tools, assemble context, hand off to one another, fail, retry, escalate to a human, and leave a record you can replay in front of a regulator. Agent quality is dominated by harness quality. A good model in an excellent harness beats an excellent model in a bad one, every time, and we intend to have the evaluations that prove it. You own that harness. You also own the second half that feeds it: proprietary domain models that natively speak the language of work, and that are trained, tuned, and evaluated specifically for how they behave inside the harness rather than on a static benchmark. **WHAT WE MEAN BY "THE HARNESS" ** The harness is everything between raw model weights and an agent that can be trusted with someone's career. Concretely: tool-call contracts and schema-validated structured outputs; context assembly, compaction, and long-session memory; state and handoff between agents; routing decisions about which model serves which seat; retries, fallbacks, and a real failure taxonomy; latency and cost budgets enforced per call; guardrails and human-review gates on consequential actions; tracing and replay; and the in-loop evaluation that scores all of it continuously rather than at release time. In a hiring product, the harness is also where compliance lives. Under the EU AI Act, hiring is a high-risk category; post-Mobley, every enterprise counsel asks the same question. The ability to say "here is the exact trace of why this candidate was surfaced, which tools ran, what the model saw, and where a human signed off" is not a logging feature — it is the Compliance Shield, and it is built in the harness. What you build is simultaneously engineering infrastructure and the thing that closes enterprise deals. Most teams let the harness accrete as glue code between a framework and a prompt. We treat it as the core system of the company, and we are hiring a founding engineer to own it as such. What You'll Own **The harness ** The orchestration runtime: planner/executor and router/specialist patterns, multi-agent handoff, escalation paths, and the state model that makes long-running agent sessions coherent. Tool-call contracts: schema-validated structured outputs, function-calling reliability, a documented failure taxonomy, and retry/fallback behavior that degrades gracefully instead of hallucinating through an error. Context engineering: assembly, compaction, and retrieval policy for long sessions — deciding what the model sees, in what order, under what token budget, and proving those choices with evals rather than intuition. Model routing: which seat gets which model, with latency and cost budgets enforced per call, and the policy for escalating from the fast specialist to the deliberative model when the stakes warrant it. Guardrails and human-in-the-loop gates on consequential actions, with confidence thresholds that are measured, not guessed. Observability: tracing, replay, and structured decision records — so any agent action can be reconstructed months later for an audit, a customer escalation, or a regression hunt. In-harness evaluation as the company's quality bar: agents scored inside the live loop of tools, retries, and guardrails — not on static prompts — with harness-level regression gates in CI that block releases. The Compliance Shield surface: auditable traces and fairness instrumentation that ship as enterprise-facing product artifacts, documented and re-run on every release. The models Run the base-model bake-off: evaluate open-weight candidates (Qwen-, Mistral-, DeepSeek-class and peers) on our tasks, our license constraints, and — decisively — their behavior inside our harness. Write the recommendation memo the company acts on. Tool-use and agentic fine-tuning: train models that call tools correctly, respect schemas, know when to abstain, and behave predictably across multi-step sessions. This is where the two halves of the job meet. Build the data engine: corpus curation, deduplication, decontamination, PII scrubbing, quality filtering, and synthetic data generation for HR and recruiting domains. Run continued pretraining on the curated HR corpus, and instruction-tune (SFT) on Astoria's task formats — matching, screening, requisition drafting, structured extraction, career reasoning. Preference-tune on real hiring-outcome data (DPO-class methods and successors); own the choice of method and the honesty of the measurement. Distill and quantize the recruiting specialist to its serving budget; own tokens-per-second and cost-per-call as first-class metrics, and the inference stack (vLLM-class) it runs on. Own training infrastructure on rented GPUs: multi-node runs (FSDP/DeepSpeed-class), checkpointing, spot-interruption recovery, experiment tracking, and the training-cost budget. Maintain the model registry and versioning discipline: every deployed model has a card, an eval report, and a rollback path. Write it down: internal tech reports, runtime contracts, and decision memos the founders act on. When it serves the product, publish — we will support it. **A NOTE ON THE WORD "PROPRIETARY" ** Proprietary here does not mean pretraining a frontier model from scratch. Bloomberg spent millions training a 50B-parameter model from scratch and watched a general-purpose API model beat it on most public benchmarks within a year — while open-weight LoRA fine-tunes matched it for a few hundred dollars of compute. Proprietary means the things that actually compound: our harness (the runtime, contracts, and traces no competitor can copy off a model card), our data (a curated HR corpus and outcome-labeled hiring data nobody can buy), our post-training (recipes that survive any base-model swap), and our evaluations (task suites and fairness slices that define what "better" means in this domain). If that framing reads as obvious to you, you are who this posting is for. What You Bring **The must-haves ** - 4+ years of professional ML or software engineering, with recent hands-on work on agentic systems that ran in production and that real users or systems depended on. - AI harness experience — required. You have built or worked deep inside the harness that turns a model into a working agent: tool-calling loops, structured outputs, context management, state, retries, guardrails, human-review gates. You know from experience that a model is only as good as its behavior in-harness, and you evaluate accordingly. - Agentic architecture — required. You have designed, or trained models for, multi-agent systems: orchestration patterns, planner/executor and router/specialist splits, agent memory and state, tool routing, escalation between agents. You understand that a model is trained for the seat it occupies — a planner, a router, and a high-volume executor make different demands on the same weights. - Agent evaluation in the loop: you have built or run evaluations that score agents on multi-step trajectories, not single-turn outputs, and you have opinions about what makes such an eval trustworthy. - Production instincts for reliability: tracing, failure taxonomies, graceful degradation, latency and cost budgets. You have been paged for an agent that silently did the wrong thing, and you fixed the system, not the prompt. - Hands-on LLM post-training in the last two years — SFT at minimum; preference tuning in anger a strong plus; tool-use or function-calling fine-tuning ideal. - PyTorch fluency and the open post-training stack: Hugging Face transformers/datasets/PEFT and at least one of TRL, Axolotl, or LLaMA-Factory — or your own equivalent you can defend. - Real distributed-training experience: you have run multi-GPU jobs, debugged OOMs and loss spikes, survived checkpoint corruption and spot preemption, and can tell the story. - Data seriousness: you treat the corpus as part of the product — dedup, decontamination, PII handling, quality filtering are craft to you, not chores. - Paper-to-code speed: hand you a post-training or agent-architecture paper on Monday, see a working implementation against our system within the week. - Enough infrastructure to be dangerous: Docker, cloud GPU fleets, experiment tracking (W&B/MLflow-class), CI. - Clear async writing. Tech reports, model cards, runtime contracts, and decision memos are half this job's output. **Strong signals ** We are especially interested in candidates who can show: - An agent system you built that ran in production, described in the language of someone who owned its failures — what broke in the harness, what you changed, how you knew it was fixed. - Contributions to harness-layer tooling: agent frameworks, orchestration libraries, lm-evaluation-harness-class eval stacks, SWE-bench-style task harnesses, or tracing and observability tools for LLM systems. - An agent or model evaluation suite other people adopted — especially one that scores multi-step trajectories. - A model you tuned specifically for tool use or agentic behavior, with in-harness numbers to show for it. - A Hugging Face profile with fine-tuned models or adapters that strangers actually download — with model cards carrying honest eval tables. - A public reproduction of a post-training or agent paper, including where your numbers diverged from the authors' and why. - Contributions to the training stack: TRL, Axolotl, vLLM, llama.cpp, LLaMA-Factory. - Strong placements in Kaggle LLM competitions — bounded compute, leaderboard evals, maximum capability squeezed out. - Technical writing others cite: a blog, a tech report, a well-argued issue thread. **Nice to have** - Experience with MCP or comparable tool-integration protocols, and opinions about where they help and where they get in the way. - RLHF/RLAIF beyond DPO-class methods; reward-model experience; RL on agent trajectories. - Distillation and quantization to production (AWQ/GPTQ/GGUF-class) with measured quality-latency trade-offs. - Tokenizer work, corpus-scale data pipelines, or dataset releases. - Safety, red-teaming, or fairness-evaluation experience — especially anything adjacent to regulated domains. - Multilingual training or evaluation. - Workshop or conference publications. Nice. Genuinely optional. **What this role is not ** Not prompt engineering — that craft lives with our agentic engineer, and you will be building the system that makes prompts matter less. Not gluing together an off-the-shelf agent framework and calling it a runtime; we expect you to have opinions about where the frameworks fall short at this level of consequence. Not frontier-scale pretraining — there is no ten-thousand-GPU cluster here, and we think that's the right call. Not publication-first research — we will support publishing when it serves the product, but the product is the point. And not fine-tuning-by-managed-button: if your training experience is exclusively a closed platform's upload form, this will be a hard job. **How to Apply ** A short note — half a page is plenty. Not a cover letter. Tell us what you think of the harness thesis above, or push back on it. We are looking for a point of view, not enthusiasm. An agent system you have owned. Three sentences: what it did, how it failed in the harness, and what you changed in the system to fix it. A link to something we can inspect — a model, adapter, eval suite, or harness-layer contribution, with numbers attached. A model card, a W&B report, or a repo beats a resume line. Astoria AI is an equal opportunity employer. We especially encourage applications from engineers and researchers whose own career paths have been non-linear — you will understand this mission intuitively, because you have lived the problem we are solving. Astoria AI is the operating system for human potential. The place where potential finds its purpose.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on wellfound · posted 2026-09-23. ApplySarthi collects openings and links to application pages; the role is advertised by Astoria AI, not by us.