GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Remote in Portugal·Part-time·Added 14 days ago

Anyone AI

5 open roles

Overview

Job details

  • US$65 an hour

    Gross, as stated in the ad.

  • Fixed-term contract

    Part-time hours.

  • Fully remote

    Per the ad.

Requirements

  • Have 3+ years of experience

    Mid-level role.

The role

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We're looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What you'll do

You'll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging

  • CUDA and Triton optimization

  • Translation between kernel frameworks

  • Hardware migration

  • Operator fusion

  • Performance profiling and benchmarking

  • Numerical correctness verification

  • Compilation and runtime debugging

  • Memory hierarchy optimization

  • Kernel-level AI workload performance

You'll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

  • Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations

  • Evaluating numerical tolerance thresholds

  • Reviewing kernel benchmarks and determining whether comparisons are fair

  • Identifying performance bottlenecks and optimization opportunities

  • Assessing whether performance targets are realistic given hardware limits

  • Reviewing kernel translations and hardware migrations

  • Identifying compilation, driver, memory, shape, and runtime issues

  • Determining whether technical tasks are genuinely difficult or incorrectly configured

  • Providing clear, actionable technical feedback

How you'll work

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Requirements

What we're looking for

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

  • Strong experience with at least two of the following:

    • CUDA

    • Triton

    • NKI / AWS Neuron

    • Pallas / JAX

  • Strong understanding of GPU performance optimization

  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

  • Understanding of:

    • Memory bandwidth

    • Compute throughput

    • GPU occupancy

    • Shared memory

    • Register pressure

    • Memory coalescing

    • Bank conflicts

  • Strong understanding of floating-point numerical correctness and tolerance thresholds

  • Experience debugging kernel compilation and runtime issues

  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Nice to have

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications

  • Translating kernels between CUDA, Triton, or other frameworks

  • Migrating kernels across hardware platforms

  • Debugging incorrect kernel implementations

  • Optimizing kernel performance

  • Fusing multiple operations into optimized kernels

  • Experience across both NVIDIA GPU and custom accelerator ecosystems

  • Experience with AWS Trainium, TPU, JAX, or other accelerators

  • Compiler engineering experience

  • Familiarity with MLIR, XLA, or intermediate representation lowering

  • Contributions to GPU or ML kernel libraries

  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

  • Experience with AI model evaluation, RLHF, or technical benchmark development

About Anyone AI

Anyone AI is an EdTech company that develops AI-powered educational tools and platforms. Based in San Francisco, the company focuses on leveraging artificial intelligence to enhance learning experiences, offering products that range from AI tutoring systems to adaptive learning technologies. While the company is relatively young, it has established a presence in the AI education space, working with experts across various domains to train and refine its AI models.

Anyone AI has a significant remote-first presence in Spain, hiring experts in fields like mathematics, physics, biology, and software engineering to contribute to AI training and development. The company offers fully remote positions within Spain, making it an attractive option for international professionals seeking flexible work arrangements. With a focus on AI model training and development, Anyone AI provides opportunities for experts to apply their knowledge in a cutting-edge field, though its brand recognition is limited outside of niche AI and EdTech circles.

Industry
EdTech
Employees
11–50
Headquarters
San Francisco, USA
All open roles at Anyone AI

Tags

WhatsApp