ApplySarthi

Member of Technical Staff, Performance Engineering

Psi

Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.

Got this interview? Our apps help you get the job.

Skills named in this job

Read from the description itself, not inferred.

This role on the market

641 open performance roles across 166 companies are on ApplySarthi right now, most of them in Bengaluru (46), Hyderabad (26), Pune (8).

What performance roles keep asking for: C++ (16%), Python (15%) — counted across their open postings here.

Member of Technical Staff jobs in the United States · Remote Member of Technical Staff jobs

Psi has 19 open roles listed here.

Counted across 14 company job boards, updated as roles open and close.

Preparing for this interview

Interviews for performance roles keep coming back to C++, Python. Practise those questions before you sit with Psi.

Questions you are likely to be asked

  1. Why do you want to join Psi?
  2. What is your experience with C++? Tell me one thing you learned the hard way.
  3. What do you do when a production issue happens on your code?
  4. Walk me through a system you built. How was it designed, and what would you change now?
  5. Tell me about a hard bug you tracked down. How did you find the cause?

Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.

Practise the Member of Technical Staff, Performance Engineering at Psi interview free →

Overview Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking a performance engineer who owns a workload's time end to end, from the application code and the data path through the scheduler to the GPU kernels, across AI and HPC, to make one cluster produce more physics. Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit. The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era. We have one product: new physics, at scale. Role and Responsibilities Own a workload's time end to end. Your unit of work is a workload, not a box or a layer: a reinforcement-learning training loop, a multi-node physics simulation, an open-weight model we serve, a numerical campaign on an open physics problem, a coupled multiphysics engine. Its time is spent in application code, data loading and preprocessing, memory, storage and checkpoint I/O, container and queue overhead, service round trips, network collectives, and GPU kernels. All of it is yours. When it runs slow you find which layer and you fix it there. Where a researcher owns the code you land the fix in their repository and leave a benchmark behind. Write the kernels. Physics campaigns, simulation engines and learned-model work need custom GPU code: exact and extended-precision numerics, sparse and dense linear algebra, adaptive meshes and stencils, attention and cache paths for serving. You write them in CUDA C++ or Triton, profile them, and turn the good ones into a library other people call. Make simulations run at scale. Fluid, thermal, electrical, molecular and coupled multiphysics solvers, both established codes and the ones physicists write themselves, run on the cluster as multi-node MPI jobs: domain decomposition and strong scaling, GPU ports of solver kernels where the measurements say they pay, I/O and checkpointing, and continuous data generation for learned models. You decide with numbers when a solver belongs on GPUs and when it does not. Make training fast on our nodes. Distributed and reinforcement-learning training on the cluster: sampler and trainer placement, collective communication, checkpoint cost, resume from failure. You know where step time goes, what MFU means for an RL loop as against a dense pretraining run, and how to close the gap. Stand up and tune serving on our own capacity. Deploy open-weight models on our GPUs, choose the serving engine and the quantization, and make the endpoint cheaper per token than the frontier API it replaces. You own the throughput and latency numbers and the decision of which requests move in-house. Measure the cluster instead of guessing. Per-workload utilization and goodput, taken on real jobs, so that queue policy, preemption classes and the next capacity decision rest on numbers somebody observed. You supply the measured GPU-hours per project that the quarterly rent-or-expand call is made on. What We're Looking For Whole-stack performance ownership. You have taken a workload from slow to fast across more than one layer, with a trace that shows where the time went before and after, and at least one of those layers was not the GPU: a data loader, a filesystem, a scheduler, a container start, a service call, a Python hot loop. You work from measurements, instrument systems you did not write, and verify a win on a real workload before you claim it. Kernels you wrote yourself. Five or more years making GPU code fast, including CUDA C++ or Triton kernels that other people's systems depend on. You read Nsight Compute output the way a physicist reads a plot, and you can say where a kernel sits on the roofline and why. Scale on at least one side, and the appetite for the other. Either real multi-node AI training experience (distributed training throughput, collective-communication behavior, InfiniBand topology, where step time goes and why MFU disappoints) or HPC simulation at scale (a CFD, molecular dynamics, multiphysics or comparable code taken across nodes with MPI, scaling measured, solver kernels ported or tuned on GPUs, FP64 and extended-precision numerics). Half the demand on this cluster is simulation and physics numerics and half is AI; you bring one and learn the other on our hardware. You can explain a performance result to a researcher who will never open a profiler, and your numbers are usable by people who make decisions with them. Nice to Have Both sides of the scale requirement: multi-node AI training and HPC MPI simulation. Inference serving for large open-weight models: vLLM, SGLang or comparable, multi-GPU placement of mixture-of-experts models, KV-cache and quantization trade-offs. Source-level work in OpenFOAM, LAMMPS, GROMACS, Nek5000, or comparable codes; AMR or stencil frameworks; PETSc, Trilinos, MAGMA or comparable libraries. CUTLASS, CuTile, or comparable kernel frameworks; compiler-level work in Triton or MLIR. Kueue or Slurm queue policy on a shared cluster: priority classes, preemption, bin-packing, and the before-and-after that justified a policy change. You have owned a capacity decision with money attached and been held to the number. Experience turning research code into a maintained library that a team of scientists depends on. How We Work We hold a high technical bar and give people full ownership of their work, from spec to ship to on-call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI-native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage. Location and Compensation This role is based in Boston. We will consider remote candidates on a case-by-case basis. We offer competitive compensation including salary, benefits, and meaningful early-stage equity. We evaluate on whole-stack performance ownership, GPU depth across AI and HPC, measurement discipline, and shipping velocity. We are an equal opportunity employer and value diverse perspectives in building platforms for AI-driven discovery.

Match this job to your CV

ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.

Check my match →

Similar open roles

Need answers during your interview? Try Live Sarthi.

Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.

Try Live Sarthi free →

A Windows app, from the same team as ApplySarthi.

Listed on ashby · posted 2026-08-15. ApplySarthi collects openings and links to application pages; the role is advertised by Psi, not by us.