Senior Inference Engineer, AGI
Amazon.com Services LLC
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
132 open inference roles across 34 companies are on ApplySarthi right now, most of them in Bengaluru (3), Delhi NCR (2).
- Member of Technical Staff – AI Inference platform, featuresLyceum
- Senior Machine Learning Engineer, LLM Inference OptimizationNebius
- Senior Machine Learning Engineer, LLM Inference OptimizationJobgether
- AI Engineer 5 (FM Hosting, LLM Inference)Capitalone
- Software Engineer II - AI/ML, Neuron InferenceAnnapurna Labs (U.S.) Inc.
What inference roles keep asking for: LLMs (49%), Python (49%), Machine learning (36%), PyTorch (26%), System design (26%), Kubernetes (24%), AWS (23%), Observability (20%) — counted across their open postings here.
AWS jobs · C++ jobs · Deep learning jobs · Java jobs
Amazon.com Services LLC has 4,316 open roles listed here.
- Delivery Station Customer Service Associate
- Delivery Station Customer Service Associate, DSL
- Senior Business Intelligence Engineer, North America Supply Chain Execution Strategy
- Mechatronics & Robotics Tech
- Automation Engineer Manager
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for inference roles keep coming back to LLMs, Python, Machine learning, PyTorch. Practise those questions before you sit with Amazon.com Services LLC.
Questions you are likely to be asked
- Why do you want to join Amazon.com Services LLC?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- What do you do when a production issue happens on your code?
- Walk me through a system you built. How was it designed, and what would you change now?
- Tell me about a hard bug you tracked down. How did you find the cause?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Inference Engineer, AGI at Amazon.com Services LLC interview free →We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI. This is a full-stack inference role: you will work across the entire path a modeltakes from research Basic qualifications: - 3+ years of building machine learning models for business application experience - PhD, or Master's degree and 6+ years of applied research experience - Experience programming in Java, C++, Python or related language - Experience with neural deep learning methods and machine learning - 2+ years of hands-on experience optimizing inference for neural models — not just using inference frameworks, but profiling and improving them - Strong understanding of deep learning architectures (transformers, attention mechanisms, autoregressive decoding) and their application to speech/audio or other multimodal domains - Production track record delivering latency-constrained, real-time inference systems under concurrent load - Experience with GPU performance optimization — memory hierarchy, occupancy, KV-cache management, and the accelerator programming model - Demonstrated ownership of a technical area — driving execution for a workstream and collaborating effectively across scientists and engineers Preferred: - Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc. - Experience with production LLM/multimodal serving internals (e.g., vLLM, TensorRT-LLM): scheduler, batching, block manager, sampler customization - Hands-on experience building real-time or streaming AI systems — speech, audio, or video — with hard latency budgets - Experience authoring custom GPU kernels (CUTLASS, Triton, raw CUDA/PTX), fused attention (FlashAttention-style), or quantized GEMM - Familiarity with model-compression and efficiency techniques — quantization, pruning, distillation, speculative decoding, long-context optimization - Experience building offline inference or rollout/reward-serving infrastructure for reinforcement learning or large-scale evaluation - Experience with distributed training and post-training pipelines (SFT through RL) — parallelism strategies, training stability, and multi-accelerator communication (NCCL, NVLink) - Familiarity with multiple hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how architecture choices affect inference latency, memory, and cost - Background in speech-to-speech or audio generative models (codec models, autoregressive audio generation), speech recognition, or speech synthesis - Experience shipping research to production at scale — models serving real users, not just benchmark results - Contributions to open-source inference/kernel projects (vLLM, CUTLASS, FlashAttention, TensorRT-LLM, Triton, or similar) Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records. Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner. The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits . USA, CA, Sunnyvale - 192,200.00 - 260,000.00 USD annually USA, MA, Boston - 167,100.00 - 226,100.00 USD annually USA, WA, Seattle - 167,100.00 - 226,100.00 USD annually
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Automation Engineer ManagerAmazon.com Services LLC
- Sr. HR Business Partner, NACFAmazon.com Services LLC
- Flight Monitor, Amazon - Prime AirAmazon.com Services LLC
- Flight Monitor, Amazon - Prime AirAmazon.com Services LLC
- Business Intelligence Engineer, Amazon Prime Video Product AnalyticsAmazon.com Services LLC
- Senior Program Manager, Global Tax PMOAmazon.com Services LLC
- Export Compliance Manager, Global Trade ServicesAmazon.com Services LLC
- Senior Site EHS ManagerAmazon.com Services LLC
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on amazon · posted 2026-08-27. ApplySarthi collects openings and links to application pages; the role is advertised by Amazon.com Services LLC, not by us.