Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.
We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.
What You’ll Work On
You’ll work with GPU and accelerator kernel tasks involving:
•
Kernel implementation and debugging
•
CUDA and Triton optimization
•
Translation between kernel frameworks
•
Hardware migration
•
Operator fusion
•
Performance profiling and benchmarking
•
Numerical correctness verification
•
Compilation and runtime debugging
•
Memory hierarchy optimization
•
Kernel-level AI workload performance
You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.
What We’re Looking For
•
3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
•
Strong experience with at least two of the following:
•
CUDA
•
Triton
•
NKI / AWS Neuron
•
Pallas / JAX
•
Strong understanding of GPU performance optimization
•
Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
•
Understanding of:
•
Memory bandwidth
•
Compute throughput
•
GPU occupancy
•
Shared memory
•
Register pressure
•
Memory coalescing
•
Bank conflicts
•
Strong understanding of floating-point numerical correctness and tolerance thresholds
•
Experience debugging kernel compilation and runtime issues
•
Ability to distinguish software defects, environment problems, and genuine optimization challenges
Relevant Experience
Candidates should have experience with several of the following types of work:
•
Writing kernels from technical specifications
•
Translating kernels between CUDA, Triton, or other frameworks
•
Migrating kernels across hardware platforms
•
Debugging incorrect kernel implementations
•
Optimizing kernel performance
•
Fusing multiple operations into optimized kernels
Nice to Have
•
Experience across both NVIDIA GPU and custom accelerator ecosystems
•
Experience with AWS Trainium, TPU, JAX, or other accelerators
•
Compiler engineering experience
•
Familiarity with MLIR, XLA, or intermediate representation lowering
•
Contributions to GPU or ML kernel libraries
•
Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
•
Experience with AI model evaluation, RLHF, or technical benchmark development
What You’ll Be Responsible For
•
Reviewing GPU and accelerator kernel implementations for correctness
•
Comparing outputs against reference implementations
•
Evaluating numerical tolerance thresholds
•
Reviewing kernel benchmarks and determining whether comparisons are fair
•
Identifying performance bottlenecks and optimization opportunities
•
Assessing whether performance targets are realistic given hardware limits
•
Reviewing kernel translations and hardware migrations
•
Identifying compilation, driver, memory, shape, and runtime issues
•
Determining whether technical tasks are genuinely difficult or incorrectly configured
•
Providing clear, actionable technical feedback
Engagement
Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation
This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.