Agentic Compiler Engineer
Kog
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
71 open compiler roles across 12 companies are on ApplySarthi right now, most of them in Bengaluru (9), Pune (1), Hyderabad (1).
- Deep Learning Compiler Intern - 2027Nvidia
- Senior Compiler Engineer - Systems TeamRiverlane
- Senior ML Compiler Engineer, NeuronAnnapurna Labs
- Senior Compiler EngineerIntel
- Compiler Engineer, Neuron Automated Reasoning GroupAmazon
What compiler roles keep asking for: C++ (25%), Python (14%) — counted across their open postings here.
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for compiler roles keep coming back to C++, Python. Practise those questions before you sit with Kog.
Questions you are likely to be asked
- Why do you want to join Kog?
- What is your experience with LLMs? Tell me one thing you learned the hard way.
- What do you do when a production issue happens on your code?
- Walk me through a system you built. How was it designed, and what would you change now?
- Tell me about a hard bug you tracked down. How did you find the cause?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Agentic Compiler Engineer at Kog interview free →ABOUT KOG Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels. On the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming. On the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs. Kog generates 3,500 tokens/s per request on 8 AMD MI300X GPUs and 2,100 tokens/s per request on 8 NVIDIA H200 GPUs, in FP16 at batch size 1, with quantization and speculative decoding disabled. Our next major project is AGCO, our agentic compiler. AGCO is designed to optimize LLMs across different GPUs and optimization targets, including very fast inference. The team has 10 people, including 9 engineers and researchers and 4 PhDs. Test it at playground.kog.ai. Read the technical details on the Kog Labs blog. WHAT YOU WILL WORK ON You will work directly on AGCO. The goal is to build a system that can explore ways to optimize LLM execution, generate changes, compile them, check correctness, run them on real hardware, measure the results, and use this feedback to guide the next optimization. You will contribute to areas such as: Compiler and IR design for representing and transforming LLM computations. Optimization passes, lowering, and code generation. Search methods for exploring different implementations and execution strategies. Verification and correctness checks for generated changes. GPU execution, profiling, and performance optimization. LLM inference across operators, memory, parallelism, and communication. Optimization loops that connect generated changes to measurements on real GPUs. One direction we are exploring combines an IR, a verifier, a compiler, and a search optimizer. We plan to start with focused problems, build working prototypes, and extend the system from what we learn. Your main area will depend on your experience, skills, and interests. You may focus more on compilers, GPU systems, or LLM inference while working closely with people across the full stack. WHAT WE LOOK FOR We look for engineers with deep technical expertise and original work in at least one area relevant to AGCO. Relevant experience includes: Compiler engineering, including optimization passes, IRs, lowering, code generation, LLVM, or MLIR. GPU programming with CUDA, HIP, Metal, Vulkan, or similar technologies. GPU performance work involving kernels, memory, synchronization, profiling, or hardware behavior. LLM inference engines and performance optimization. Attention, MoE, parallelism, communication, or other systems-level parts of LLM execution. Formal verification, equivalence checking, SAT/SMT, or related methods. Systems that generate, search, test, benchmark, or optimize code automatically. We care about what you personally built and the technical decisions behind it. Strong candidates can explain the problem, their approach, the alternatives they explored, and how they measured the result. We review technical work during the process. This can be public code, an upstream contribution, a paper, a thesis, a technical project, or a detailed write-up based on work you can share. WHAT WE OFFER You will join a small team building AGCO as a core part of Kog's technology. Work at the intersection of compilers, GPU systems, and LLM inference. Direct access to engineers working across the full inference stack. A fast loop from an optimization idea to compilation, execution, verification, and measurement on real GPUs. The opportunity to go deep in your strongest technical area while expanding into the other parts of the stack. High ownership over technical decisions and systems that will shape how Kog optimizes LLM inference. This role is based in Paris, and we are looking for candidates who can relocate to Paris and work closely with the team. Find Jobs in France on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-09-29. ApplySarthi collects openings and links to application pages; the role is advertised by Kog, not by us.