ApplySarthi

Multimodal ML Engineer

Npv

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

27 open multimodal roles across 12 companies are on ApplySarthi right now, most of them in Delhi NCR (1).

What multimodal roles keep asking for: Python (44%), Machine learning (41%), Computer vision (37%), Deep learning (33%), LLMs (33%), PyTorch (33%) — counted across their open postings here.

Hugging Face jobs · LLMs jobs · PyTorch jobs

Npv has 4 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for multimodal roles keep coming back to Python, Machine learning, Computer vision, Deep learning. Practise those questions before you sit with Npv.

Questions you are likely to be asked

  1. Why do you want to join Npv?
  2. What is your experience with Hugging Face? Tell me one thing you learned the hard way.
  3. What would you check first if a model's accuracy dropped after going live?
  4. When would you not use machine learning for a problem?
  5. Walk me through a model you built, from the data to how it was used.

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Multimodal ML Engineer at Npv interview free →

We're looking for a Multimodal ML Engineer to join White Circle , an AI Safety company building the safety, reliability, and optimization layer for AI systems through natural-language policies it automatically tests, enforces, and improves at scale. Backed by $70M (Series A) from top funds and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, and others, White Circle processes 100M+ API calls monthly and fine-tunes and trains its own LLMs to run faster and cheaper than open or proprietary models. You will Train and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints. Design experiments, build multimodal data pipelines, and train MoE architectures. Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end. Define evaluation metrics that actually matter for the product. Requirements 3+ years training large-scale multimodal models. Strong PyTorch and distributed training experience (DeepSpeed, FSDP). Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar. Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling). Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization. Relocation to Paris or London (hybrid) required. Bonus Audio signal processing fundamentals – spectrograms, mel features, noise reduction. MoE architecture experience. We offer Competitive salary + equity. Official employment, visa and relocation help. Find more English Speaking Jobs in France on Arbeitnow

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on arbeitnow · posted 2026-10-09. ApplySarthi collects openings and links to application pages; the role is advertised by Npv, not by us.