AI Research Engineer (Kernel & Inference Optimization)
Jobgether
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
274 open optimization roles across 74 companies are on ApplySarthi right now, most of them in Bengaluru (7), Hyderabad (3), Delhi NCR (2).
- 2027 Applied Science Internship - Reinforcement Learning & Optimization (Machine Learning) - United States, PhD Student Science RecruitingAmazon
- Senior Applied Scientist, Efficient LLM Inference & Model OptimizationNebius
- Senior Software Engineer, Mission OptimizationPlanet
- Internship in Cell Culture Process Development / Media Optimization (m/f/d)Roche
- Vice President — Card Acquisition Strategy, Offers & Channel OptimizationJPMorgan
What optimization roles keep asking for: Machine learning (25%), Python (22%) — counted across their open postings here.
Remote AI Research Engineer jobs · Machine learning jobs · NLP jobs
Jobgether has 4,572 open roles listed here.
- Account Director
- Account Executive (Multi-Product)
- Account Manager (Email Marketing)
- Account Manager (Email Marketing)
- Account Manager (Email Marketing)
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for optimization roles keep coming back to Machine learning, Python. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with Machine learning? Tell me one thing you learned the hard way.
- Tell me about a time the data was messy or wrong. What did you do?
- How would you explain your model's result to someone who is not technical?
- What would you check first if a model's accuracy dropped after going live?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the AI Research Engineer (Kernel & Inference Optimization) at Jobgether interview free →Accountabilities: Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization. Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms. Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability. Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates. Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions. Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations. Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL). Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding. Design and optimize distributed inference systems using approaches such as tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU workloads. Work with cross-functional engineering and research teams to integrate optimized inference frameworks into production and edge-device applications. Define evaluation methodologies, document experimental results, compare performance against established benchmarks, and continuously refine optimization strategies. Monitor production performance and use empirical research to identify opportunities for further improvements in scalability, efficiency, and reliability. Requirements: Degree in Computer Science or a related technical field; a PhD in NLP, Machine Learning, or a related discipline is highly relevant, particularly with a strong AI research track record and publications at leading conferences. Proven expertise in Metal Shading Language (MSL), including the ability to write custom compute shaders from scratch. Demonstrated experience with low-level kernel optimization and inference optimization on mobile or other resource-constrained devices. Track record of delivering measurable improvements in inference latency, throughput, and memory footprint for domain-specific applications. Deep understanding of modern model-serving architectures, inference engines, and optimization techniques for high-performance AI deployment. Strong experience writing GPU kernels for mobile devices such as smartphones. Practical experience developing and deploying end-to-end inference pipelines, from model optimization through production integration on constrained hardware. Strong ability to apply empirical research and systematic experimentation to solve latency, computational, and memory challenges. Experience designing robust evaluation and benchmarking frameworks for inference systems. Knowledge of distributed inference techniques, including tensor parallelism, pipeline parallelism, and expert parallelism for large-scale GPU clusters. Deep understanding of the mathematical foundations and architecture of diffusion models and Vision Transformers. Familiarity with modern inference optimization techniques including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding such as EAGLE. Strong analytical and problem-solving abilities, with an ability to investigate complex system bottlenecks and turn research findings into practical engineering solutions. Excellent English communication skills and the ability to collaborate effectively with distributed, cross-functional technical teams. Benefits: Opportunity to work on advanced AI systems spanning model serving, inference optimization, mobile computing, edge deployment, and large-scale distributed inference. Remote-first working environment with an international team. Exposure to cutting-edge AI research and practical systems engineering challenges. Opportunity to contribute to performance-critical infrastructure where improvements can have a measurable impact on real-world AI applications. Collaborative environment combining research-driven experimentation with hands-on engineering. Opportunity to work with advanced model architectures including diffusion models, Vision Transformers, and multimodal systems. Access to challenging technical problems involving GPU kernels, inference engines, memory optimization, and distributed computing.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- (Senior) Product Engineer (Backend)Jobgether
- (Senior) Product Engineer (Backend)Jobgether
- (Senior) Product Engineer (Backend)Jobgether
- (Senior) Product Engineer (Backend)Jobgether
- (Senior) Product Engineer (Backend)Jobgether
- (Senior) Product Engineer (Backend)Jobgether
- (Senior) Product Engineer (Backend) (m/f/d)Jobgether
- (Senior) Product Engineer (Backend) (m/f/d)Jobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-30. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.