ApplySarthi Match jobs to your CV

Staff Software Engineer - Inference Backends

Hume-Ai

Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

132 open inference roles across 34 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).

What inference roles keep asking for: LLMs (49%), Python (49%), Machine learning (36%), PyTorch (26%), System design (26%), Kubernetes (24%), AWS (23%), Observability (20%) — counted across their open postings here.

Remote Software Engineer jobs · C++ jobs · CI/CD jobs · Go jobs · Linux jobs

Hume-Ai has 9 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Hume-Ai.

Questions you are likely to be asked

  1. Why do you want to join Hume-Ai?
  2. What is your experience with Observability? Tell me one thing you learned the hard way.
  3. Describe a time a deadline forced a trade-off in quality. What did you choose and why?
  4. How would you design an API for a feature you have worked on?
  5. What do you do when a production issue happens on your code?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Staff Software Engineer - Inference Backends at Hume-Ai interview free →

Hume AI is looking for a systems-oriented engineer to own the path from trained model to production inference. Join us in the heart of New York City and contribute to our endeavor to ensure that AI is guided by human values, the most pivotal challenge—and opportunity—of the 21st century. About Us Hume AI is a Series B startup dedicated to building artificial intelligence that is directly optimized for human well-being. As the first company to release speech language models, we’re focused on expanding our research to encompass audio understanding models and evaluation platforms for enterprises. Our goal is to enable a future in which technology draws on an understanding of human emotional expression to better serve human goals. As part of our mission, we also conduct groundbreaking scientific research, publish in leading scientific journals like Nature , and support a non-profit, The Hume Initiative, that has released the first concrete ethical guidelines for empathic AI ( www.thehumeinitiative.org ). You can learn more about us on our website ( https://hume.ai/ ) and read about us in WIRED , Forbes , and Venturebeat . About the Role As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference: graph export, engine compilation, runtime integration, serving contracts, client libraries, artifact verification, and the performance and correctness of what runs in production. You will work closely with research scientists, machine learning engineers, backend engineers, and the Data Plane team to bring new models and inference capabilities into production. This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy, make sound architectural decisions, and own systems throughout their lifecycle—from initial design and implementation through deployment, observability, optimization, and production support. About the Role As the first engineer dedicated full-time to inference backends at Hume, you will own the systems that take trained models from checkpoints to production inference. This includes graph export, engine compilation, runtime integration, serving contracts, artifact verification, and production performance and correctness. You will also own Hume’s internal inference platform, which serves many of our in-house models across products. This includes deployment, routing, load balancing, health checking, observability, capacity management, and safe model rollout. You will work closely with research scientists, machine learning engineers, backend engineers, and product teams to turn new model capabilities into reliable production systems. This is a systems-oriented role focused on performance, numerical correctness, reliability, and efficient use of accelerators. You will operate with a high degree of autonomy and own systems from design through production. What You’ll Do Own the path from trained checkpoint to served request, including graph export, engine compilation, runtime integration, and serving configuration. Build and evolve Hume’s internal inference platform for serving multiple models across products and workloads. Design and operate serving infrastructure, including routing, load balancing, health checking, autoscaling, capacity management, and failure handling. Build reproducible, versioned inference artifacts and tooling for validation, deployment, promotion, and rollback. Build verification gates that catch numerical, behavioral, and performance regressions before production. Design and maintain internal client libraries and standardized serving contracts. Profile and optimize latency, throughput, memory usage, batching, scheduling, and accelerator utilization. Diagnose production issues across application, runtime, container, networking, driver, and hardware boundaries. Improve observability, resilience, testability, and operational safety across the inference stack. Write clear technical documentation for the systems and APIs you build. What You’ll Bring Significant professional experience building server-side, infrastructure, distributed, or systems software. Strong Linux fundamentals and hands-on experience troubleshooting and profiling production systems. Professional experience with at least one systems-oriented language such as Rust, Go, C, or C++. Experience with distributed systems concepts such as load balancing, health checking, failure recovery, observability, and capacity management. A practical understanding of neural-network execution, including computation graphs, tensor shapes, data types, and accelerator execution. Experience profiling and optimizing production systems. Comfort working across languages and tooling, including Python for model export, validation, and integration workflows. Strong ownership, independent technical judgment, and clear written and verbal communication. The ability to use AI-assisted coding tools effectively while retaining the ability to explain, validate, debug, and modify the result independently. Bonus Points Experience building or operating shared model-serving or inference platforms. Experience with GPU inference, CUDA, ONNX, TensorRT, PyTorch, Triton, vLLM, or similar systems. Experience diagnosing numerical correctness issues such as precision loss, numerical drift, or nondeterminism. Experience building high-performance client libraries, SDKs, or networked systems. Experience with containers, CI/CD, or deploying software into on-premises or customer-managed environments. Contributions to systems, inference, distributed-systems, or machine-learning infrastructure open-source projects.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on ashby · posted 2026-09-17. ApplySarthi collects openings and links to application pages; the role is advertised by Hume-Ai, not by us.