Forward Deployed Inference Engineer
Tensordyne
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
168 open inference roles across 39 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).
- Inference Engineer, AGIAmazon
- ML Inference EngineerSpaitial
- Senior ML Engineer | Kimchi (LLM Inference Optimization)Kimchi
- Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs , Elastic CollectivesAnnapurna Labs
- Lead Product Manager, InferenceMistral.ai
What inference roles keep asking for: Machine learning (49%), LLMs (44%), Python (43%), System design (24%), PyTorch (24%), Kubernetes (22%), AWS (20%), NLP (20%) — counted across their open postings here.
Generative AI jobs · Kubernetes jobs · LLMs jobs · PyTorch jobs
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to Machine learning, LLMs, Python, System design. Practise those questions before you sit with Tensordyne.
Questions you are likely to be asked
- Why do you want to join Tensordyne?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- Tell me about a hard bug you tracked down. How did you find the cause?
- How do you decide what to test, and what does good code review look like to you?
- Describe a time a deadline forced a trade-off in quality. What did you choose and why?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Forward Deployed Inference Engineer at Tensordyne interview free →About Tensordyne Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world's most demanding generative AI workloads. Our platform combines purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture into a tightly integrated system purpose built for large-scale AI inference. We work with hyperscalers, Neoclouds, frontier model developers, enterprises, and infrastructure partners operating at the leading edge of AI. As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, we are building the technical customer organization that will sit directly between our engineering teams and the companies deploying the platform. Role summary We are looking for a Forward Deployed Inference Engineer who combines deep AI systems expertise with strong customer instincts. This person will own the path from a customer workload or model request to a technical result and, where needed, to an optimized model running successfully on Tensordyne hardware and software. The role sits at the intersection of model architecture, inference performance, systems optimization, developer tooling, and customer deployment. You will work hands-on with engineering while also acting as a technical bridge to Product, BizDev, Sales, and customers. What you will do Turn customer workloads into fast, credible performance answers through profiling & benchmarking . Define relevant KPIs, compare against competitive baselines, and keep our evaluation methodology current with external benchmarks. Model enablement & optimization: convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify performance bottlenecks, and work with compiler, runtime, kernel, and system teams to improve results. Deployment / forward engineering: work directly with customers and partners on technical PoCs, integration, deployment, and debugging; translate requirements into measurable acceptance criteria for quality, latency, throughput, and other relevant KPIs. Track profiling-to-hardware accuracy by continuously comparing profiling/simulation results with actual hardware deployments, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, KV-cache management, or adjacent systems to the owning teams. Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements. Core qualifications Strong hands-on experience with AI models and inference systems , especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models. Strong Python and PyTorch skills and the ability to understand and modify model code. Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization. Strong problem-solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries. Proficiency in using AI-powered developer tools (e.g., Claude Code, Cursor). Strong pluses Experience with LLM serving and deployment stacks such as vLLM, SGLang, or similar systems. Experience working directly with customers or external technical partners . Experience bringing models up on new accelerators or non-standard hardware , including performance debugging across framework/runtime/hardware boundaries. Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes . Experience navigating and contributing to Rust codebases. Tensordyne Values Think big. Pursue ambitious technical and business goals. Aim for excellence. Quality matters in everything we build and deliver. Own it and get it done. Take responsibility and drive results. Operate with integrity. Be direct, transparent, and respectful. Win as a team. Make the people around you more effective. Value different perspectives. The strongest teams challenge assumptions and bring diverse experience to difficult problems. Tensordyne is an equal opportunity employer. We believe diverse teams are better equipped to solve complex problems and build exceptional technology. All qualified applicants will receive consideration for employment without regard to age, color, gender identity or expression, marital status, national origin, disability, protected veteran status, race, religion, pregnancy, sexual orientation, or any other characteristic protected by applicable laws, regulations, and ordinances. A note to recruitment agencies: Please do not contact Tensordyne employees or leaders regarding this role. We do not accept unsolicited agency resumes and are not responsible for fees associated with unsolicited submissions. Find Jobs in Germany on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-10-05. ApplySarthi collects openings and links to application pages; the role is advertised by Tensordyne, not by us.