ApplySarthi

AI Engineer (GPU & LLM Optimisation) – Canada / UAE / Remote

SuperQ Quantum Computing

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

25 open optimisation roles across 16 companies are on ApplySarthi right now, most of them in Bengaluru (2).

What optimisation roles keep asking for: Python (20%), Supply chain (20%), SQL (16%), LLMs (12%), Power BI (12%), SEO (12%), Stakeholder management (12%) — counted across their open postings here.

AI Engineer jobs in the United States · Remote AI Engineer jobs · C++ jobs · Docker jobs · Hugging Face jobs · Kubernetes jobs

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for optimisation roles keep coming back to Python, Supply chain, SQL, LLMs. Practise those questions before you sit with SuperQ Quantum Computing.

Questions you are likely to be asked

  1. Why do you want to join SuperQ Quantum Computing?
  2. What is your experience with LLMs? Tell me one thing you learned the hard way.
  3. How would you explain your model's result to someone who is not technical?
  4. What would you check first if a model's accuracy dropped after going live?
  5. When would you not use machine learning for a problem?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the AI Engineer (GPU & LLM Optimisation) – Canada / UAE / Remote at SuperQ Quantum Computing interview free →

Location: Canada preferred, also possible in UAE or fully Remote Experience: 4–6 years Company: SuperQ Quantum Computing Inc. **About the Role** SuperQ is looking for an AI Engineer specializing in GPU and LLM optimization to maximize the performance, scale, and efficiency of the models powering our multi-agent architecture. You will be the driving force behind accelerating inference, reducing latency, and optimizing hardware utilization as we scale our hybrid compute platform across industries like healthcare, manufacturing, and finance. **Responsibilities** * Accelerate Inference: Optimize LLM inference and serving pipelines using high-performance frameworks (e.g., vLLM, TensorRT-LLM, TGI, Triton Inference Server). * Model Compression: Implement state-of-the-art model compression techniques, including quantization (e.g., GPTQ, AWQ, FP8/INT8), pruning, and knowledge distillation to reduce memory footprints. * Hardware Profiling: Profile GPU performance to identify and eliminate bottlenecks, optimizing memory management techniques such as KV caching and PagedAttention. * Custom Compute: Develop and optimize custom CUDA or OpenAI Triton kernels for specialized, computationally heavy operations where standard libraries fall short. * Cross-functional Deployment: Collaborate with backend and platform engineering teams to deploy highly scalable, low-latency AI endpoints within our Kubernetes-based infrastructure. * System Monitoring: Ensure robust monitoring of GPU health, utilization metrics, latency, and throughput in production environments. **Requirements** * Experience: 2–4 years of experience in ML engineering, AI infrastructure, or high-performance computing (HPC) with a strong focus on Large Language Models. * Core Languages: Strong programming skills in Python; proficiency in C++ and/or CUDA is highly preferred. * Frameworks: Hands-on experience with LLM serving architectures, distributed computing, and optimization libraries (e.g., DeepSpeed, Hugging Face Accelerate, Ray). * Hardware Knowledge: Deep understanding of GPU architectures (NVIDIA), memory hierarchies, and parallel computing paradigms (e.g., NCCL). * Infrastructure: Solid understanding of containerization and orchestration (Docker, Kubernetes) tailored for GPU-accelerated workloads. * Bonus: Experience with hardware-aware neural architecture search (NAS), ML Ops, or working at the intersection of classical AI hardware and quantum compute interfaces. **Benefits & Work Culture Benefits:** * Competitive salary plus bonus based on module delivery and platform impact. * Stock options / equity – participate in the growth of a foundational tech platform. * Global remote-friendly roles: Canada, UAE or anywhere remote; with occasional team meetups and hack-weeks. * Learning & development stipend: conferences, AI/LLM workshops, certifications (e.g., in ML Ops, prompt engineering). * “Innovation time” built into schedule: work on passion projects, internal hackathons, share learnings with the team. * Cross-domain exposure: work across industries, build modules for different verticals – variety and career growth built-in. **Work Culture:** * Startup-scale agility within a mission-driven organisation: you’ll have impact and shape the direction of the platform. * Multi-disciplinary collaboration: you’ll work with quantum engineers, UI developers, domain experts – bridging tech and business. * Strong focus on autonomy, responsibility and ownership: you’ll own modules end-to-end, from conception through production and maintenance. * Culture of continuous learning: we encourage curiosity in new models, LLM architectures, AI trends, and provide time to experiment and publish/share. * Balanced remote-first mindset: flexible working hours, asynchronous collaboration but also moments of in-person/virtual team building. **Equal Opportunity Statement** SuperQ is dedicated to building a workplace where everyone can thrive. We are an equal opportunity employer and we make employment decisions based on merit, qualifications and business needs-without regard to race, colour, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other characteristic protected by law. We value diversity, equity and inclusion and strive to create a culture of belonging for all employees. [Interested in this role? Click here to apply.](https://zurl.to/gidB?source=CareerSite)

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on wellfound · posted 2026-08-24. ApplySarthi collects openings and links to application pages; the role is advertised by SuperQ Quantum Computing, not by us.