ApplySarthi

Machine Learning Engineer (Voice AI)

ViH Labs and Innovation

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

1,094 open learning roles across 229 companies are on ApplySarthi right now, most of them in Bengaluru (63), Hyderabad (25), Delhi NCR (14).

What learning roles keep asking for: Machine learning (48%), Python (36%), LLMs (22%), PyTorch (22%), Deep learning (16%), AWS (13%), Generative AI (13%) — counted across their open postings here.

Machine Learning Engineer jobs in India · Remote Machine Learning Engineer jobs · Docker jobs · Linux jobs · Machine learning jobs · PyTorch jobs

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for learning roles keep coming back to Machine learning, Python, LLMs, PyTorch. Practise those questions before you sit with ViH Labs and Innovation.

Questions you are likely to be asked

  1. Why do you want to join ViH Labs and Innovation?
  2. What is your experience with Machine learning? Tell me one thing you learned the hard way.
  3. Walk me through a model you built, from the data to how it was used.
  4. How did you know your model was actually good, and not just good on your test set?
  5. Tell me about a time the data was messy or wrong. What did you do?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Machine Learning Engineer (Voice AI) at ViH Labs and Innovation interview free →

Machine Learning Engineer – Voice AI (Traditional Speech Architectures) About the Role We are looking for a passionate and highly skilled Machine Learning Engineer – Voice AI to lead the development of our multilingual and multi-dialect speech systems. The ideal candidate should have hands-on experience in traditional speech processing, Text-to-Speech (TTS), Automatic Speech Recognition (ASR), audio signal processing, and Voice AI technologies. This role is critical to our long-term vision of building and self-hosting production-grade Voice AI solutions tailored for multiple industry use cases across India. Since we work extensively with regional Indian languages and dialects, we are looking for someone who can take ownership of the Voice AI domain with strong technical leadership and a research-driven mindset. Key Responsibilities • Design, develop, and optimize Voice AI systems using traditional and modern speech processing techniques. • Build multilingual and multi-dialect speech solutions for Indian regional languages. • Develop and improve Text-to-Speech (TTS), speech enhancement, pronunciation modeling, and voice adaptation pipelines. • Work on Automatic Speech Recognition (ASR), speaker identification, speaker verification, and keyword spotting systems. • Design robust audio preprocessing, feature extraction, and post-processing pipelines. • Improve speech quality, intelligibility, naturalness, and dialect adaptation. • Work with traditional speech processing techniques, digital signal processing (DSP), and statistical speech models where applicable. • Build scalable training and inference pipelines for self-hosted Voice AI systems. • Optimize low-latency inference for production deployments. • Collaborate with product, engineering, and data teams to deploy production-ready Voice AI solutions. • Evaluate models using objective speech quality metrics and human evaluation. • Stay up to date with advancements in speech processing, Voice AI, and audio machine learning. Required Skills & Qualifications • 2+ years of experience in Machine Learning, Speech Processing, Voice AI, or related domains. • Strong understanding of traditional speech processing techniques, including: Digital Signal Processing (DSP) MFCC, LPC, PLP, Mel Spectrograms Fourier Transform (FFT), Short-Time Fourier Transform (STFT), and Filter Banks Speech segmentation and Voice Activity Detection (VAD) • Experience with traditional Text-to-Speech systems, statistical parametric speech synthesis, concatenative synthesis, or modern neural TTS architectures. • Good understanding of speech feature extraction, spectrograms, vocoders, and audio codecs. • Experience with Automatic Speech Recognition (ASR) pipelines and acoustic modeling is preferred. • Strong Python programming skills. • Hands-on experience with PyTorch, TensorFlow, or similar machine learning frameworks. • Experience handling multilingual speech datasets and audio corpora. • Familiarity with Linux, Docker, GPU training, and inference optimization. • Understanding of speech evaluation metrics such as WER, CER, MOS, PESQ, and STOI. • Ability to independently own and drive Voice AI research and production initiatives. Preferred Qualifications • Experience working with Indian language datasets and dialect adaptation. • Experience with speaker recognition, speaker diarization, or voice biometrics. • Familiarity with Kaldi, ESPnet, HTK, CMU Sphinx, OpenSMILE, Praat, or similar speech processing toolkits. • Experience building self-hosted Voice AI infrastructure. • Knowledge of conversational AI, telephony systems, and speech analytics. • Research contributions, open-source projects, or published work in speech processing or speech synthesis. What We Offer • Opportunity to work on cutting-edge multilingual Voice AI systems. • Access to large proprietary speech datasets. • Freedom to experiment, research, and build production-grade Voice AI infrastructure. • High-impact role with ownership and technical leadership opportunities. • Collaborative and innovation-driven work environment. Ideal Candidate We are looking for someone who is deeply interested in speech processing and Voice AI, comfortable taking ownership of the complete speech pipeline—from audio preprocessing and feature engineering to speech synthesis, recognition, and deployment. The ideal candidate is excited about solving challenges in Indian languages and dialects, capable of leading Voice AI initiatives end-to-end, and committed to building scalable, production-ready speech systems using both traditional speech processing techniques and modern machine learning approaches.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on wellfound · posted 2026-08-22. ApplySarthi collects openings and links to application pages; the role is advertised by ViH Labs and Innovation, not by us.