← All jobs

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Salary
$65/hour
Location
Argentina - Fully Remote
Work type
Remote
Posted
today
Verified live
today

Apply on company site (opens in new tab)

GPUTritonRLHFAWS

Filed underInference / ServingFine-tuning

This range sits 45% below the $246k median for Inference / Serving roles on this board that publish pay (228 of 486).

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging
  • CUDA and Triton optimization
  • Translation between kernel frameworks
  • Hardware migration
  • Operator fusion
  • Performance profiling and benchmarking
  • Numerical correctness verification
  • Compilation and runtime debugging
  • Memory hierarchy optimization
  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
  • Strong experience with at least two of the following:
  • CUDA
  • Triton
  • NKI / AWS Neuron
  • Pallas / JAX
  • Strong understanding of GPU performance optimization
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
  • Understanding of:
  • Memory bandwidth
  • Compute throughput
  • GPU occupancy
  • Shared memory
  • Register pressure
  • Memory coalescing
  • Bank conflicts
  • Strong understanding of floating-point numerical correctness and tolerance thresholds
  • Experience debugging kernel compilation and runtime issues
  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications
  • Translating kernels between CUDA, Triton, or other frameworks
  • Migrating kernels across hardware platforms
  • Debugging incorrect kernel implementations
  • Optimizing kernel performance
  • Fusing multiple operations into optimized kernels

Nice to Have

  • Experience across both NVIDIA GPU and custom accelerator ecosystems
  • Experience with AWS Trainium, TPU, JAX, or other accelerators
  • Compiler engineering experience
  • Familiarity with MLIR, XLA, or intermediate representation lowering
  • Contributions to GPU or ML kernel libraries
  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls
  • Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For

  • Reviewing GPU and accelerator kernel implementations for correctness
  • Comparing outputs against reference implementations
  • Evaluating numerical tolerance thresholds
  • Reviewing kernel benchmarks and determining whether comparisons are fair
  • Identifying performance bottlenecks and optimization opportunities
  • Assessing whether performance targets are realistic given hardware limits
  • Reviewing kernel translations and hardware migrations
  • Identifying compilation, driver, memory, shape, and runtime issues
  • Determining whether technical tasks are genuinely difficult or incorrectly configured
  • Providing clear, actionable technical feedback

Engagement

Work Type: Remote

Engagement: Part-time, project-based consulting

Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Apply on company site (opens in new tab)