ApplySarthi Match jobs to your CV

Founding Voice AI Engineer

Libera

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

119 open founding roles across 71 companies are on ApplySarthi right now, most of them in Bengaluru (13), Delhi NCR (3), Mumbai (2).

What founding roles keep asking for: LLMs (33%), Python (27%), SaaS (24%), PostgreSQL (24%), TypeScript (23%), React (22%), CI/CD (20%), Node.js (18%) — counted across their open postings here.

LLMs jobs · Observability jobs · Python jobs

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for founding roles keep coming back to LLMs, Python, SaaS, PostgreSQL. Practise those questions before you sit with Libera.

Questions you are likely to be asked

  1. Why do you want to join Libera?
  2. What is your experience with LLMs? Tell me one thing you learned the hard way.
  3. How did you know your model was actually good, and not just good on your test set?
  4. Tell me about a time the data was messy or wrong. What did you do?
  5. How would you explain your model's result to someone who is not technical?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Founding Voice AI Engineer at Libera interview free →

**Full-time · Remote (anywhere) · Start: as soon as possible, ideally November 2026** ### About Libera Libera builds AI agents for debt collection in Japan. Our agents call and answer debtors in Japanese, in real time; negotiate repayment within Japan's servicer and lending regulations; follow up by SMS and letter; and record every step. We are a seed-stage company, founded in 2026 by a former BCG consultant and backed by two of Japan's leading VCs. A production pilot with a licensed servicer in Tokyo is under way, and pilots with several major Japanese financial institutions are getting started. Japan generates roughly US$270 billion (¥40 trillion) in receivables every year, and the industry that collects them still runs on phone calls, letters, and decades-old software. The agent that can do this work in Japanese does not exist yet. We intend to build it. ### About the role You will own our voice agent end to end: everything between audio in and audio out, and its quality in production. That covers speech recognition and turn detection, LLM control and conversation state, speech synthesis, guardrails, evaluation, and the collection results the agent delivers. You will be our first full-time engineer dedicated to voice, working directly with our Head of Engineering and our CEO, and you will shape how this part of the company grows. ### What you'll do - **Runtime:** build the real-time conversation loop — end-of-turn detection, barge-in, latency budgets under 1.5 seconds, disconnects, and transfers — so the agent can hold a natural phone conversation in Japanese. - **Conversation control:** design how the agent verifies identity, listens, negotiates within approved terms, and secures a promise to pay, reliably across thousands of calls and dozens of turns. - **Channels:** extend the agent beyond voice to SMS, email, and letters, and orchestrate the channel and timing for each debtor so that the right message arrives and the call gets answered. - **Evaluation:** build the simulation and regression suite — LLM-as-judge, noisy audio, accents, interruptions — that tells us whether a change has improved the agent before it reaches a customer. - **Data and post-training:** turn call recordings and outcomes into evaluation sets and training data, and post-train models so the agent learns from the best human collectors. - **Production:** ship to live customer calls, monitor them in real time, and enforce compliance rules in code — calling hours, contact limits, disclosure, stop-contact — with a complete audit trail. ### What you'll bring - You have shipped and operated a voice agent in production: real phone lines, real customers, and conversations that involved money or commitments. - You have post-trained an LLM on real data (SFT, DPO, or RL) and shipped the result, with evaluations that demonstrated the improvement. - You have built evaluation for non-deterministic systems, and you improve agents with data rather than intuition. - You can own the full stack on your own: Python, async and streaming, the basics of WebRTC, SIP, and telephony, cloud deployment, and observability. - You work fluently with AI coding tools and read their output critically. - You are willing to spend time at customer sites in Japan, listening to real calls and fixing what you hear. We do not screen on years of experience or degrees. ### Even better - Experience adapting ASR or TTS to a domain (telephone audio, names, numbers), or training VAD or end-of-turn models. - Experience in collections, lending, loan servicing, insurance, or contact centers — anywhere mistakes are expensive. - Japanese at any level. It is not required; our CEO and Japanese-speaking colleagues will work with you on language quality. - Open-source contributions to Pipecat or LiveKit, or published work in speech (Interspeech, ICASSP). ### How we work - Fully remote, from anywhere. We ask for a few hours of overlap with Japan time and a trip to Tokyo roughly once a quarter. - English is the working language of the engineering team. Japanese is the language of our customers. - Current stack: Pipecat and Pipecat Flows, Soniox, Cartesia, OpenAI models, Twilio, Cloudflare Workers, and GitHub Actions. You are free to change any of it, with evidence. - A small team: a CEO who sits with customers every week, a Head of Engineering, a few part-time engineers, and you. Company accounts for Claude Code, Codex, and model providers. ### Compensation - Base salary: ¥15,000,000–¥21,000,000 (approx. US$100,000–US$140,000), depending on experience. - Equity from 1%, vesting over four years, with additional grants considered at each financing round. - Relocation and visa support if you want to move to Japan. Employment in Japan, or a full-time contractor or EOR arrangement elsewhere. ### Process 1. 30 minutes with our CEO: the business, and your questions. 2. 90 minutes with our Head of Engineering: we play you real calls from our agent and discuss how you would improve them. 3. A paid one-day pair-programming session on our codebase, remote or in Tokyo. 4. A final conversation with our CEO and Head of Engineering, followed by an offer. We move quickly and tell you where you stand after every step. When you apply, tell us in a few lines about a voice or agent system you have built and the role you played in it.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on wellfound · posted 2026-09-24. ApplySarthi collects openings and links to application pages; the role is advertised by Libera, not by us.