Role overview

Software Engineer- GPU Kernels

Requirements and responsibilities

Readable role content extracted into sections for faster review.

Details

  • Baseten Embeddings Inference: The fastest embeddings solution available
  • The Baseten Inference Stack
  • Driving model performance optimization
  • Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing
  • Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques
  • Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap
  • Implement cutting-edge features like quantization (FP8/FP4), sparsity, and compute/communication overlap
  • Identify and resolve performance bottlenecks using tools like Nsight Systems, Nsight Compute, and Torch Profiler
  • Collaborate with research teams to productionize theoretical advancements
  • Contribute to internal and open-source GPU libraries
  • Present technical contributions at industry conferences (e.g., NVIDIA GTC, AWS re:Invent)
  • Strong understanding of GPU architecture and programming paradigms:Memory hierarchy (global, shared, registers, L1/L2 cache)Thread/block/grid organizationSynchronization techniques and race condition mitigation
  • Memory hierarchy (global, shared, registers, L1/L2 cache)
  • Thread/block/grid organization
  • Synchronization techniques and race condition mitigation
  • Proficient in C++ and GPU performance profiling tools
  • Knowledge of:CUDA C++ APIMemory access patterns and bandwidth optimizationNumerical precision and quantization strategiesModern GPU features (e.g., tensor cores, async operations)
  • CUDA C++ API
  • Memory access patterns and bandwidth optimization
  • Numerical precision and quantization strategies
  • Modern GPU features (e.g., tensor cores, async operations)
  • Memory hierarchy (global, shared, registers, L1/L2 cache)
  • Thread/block/grid organization
  • Synchronization techniques and race condition mitigation
  • CUDA C++ API
  • Memory access patterns and bandwidth optimization
  • Numerical precision and quantization strategies
  • Modern GPU features (e.g., tensor cores, async operations)
  • Experience with Transformer models and attention optimization (e.g., Flash Attention)
  • Familiarity with GPU kernel libraries: Cutlass, Triton, Thrust, CUB
  • Background in GEMM tuning and distributed/multi-GPU compute
  • Contributions to open-source GPU projects
  • Research publications or conference presentations on GPU performance
  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Similar roles

Keep a backup shortlist.

Browse stack
FocusKernelsRole area
Seniority signalOpen levelCandidate level
StackAWS, SparkPrimary skills
Location2 accepted countriesEligibility

Stack

Use these tags to compare similar remote roles.

Location eligibility

Candidates should apply only when their profile country is listed here.

Your profileCountry not setSign in to check your country against this role.

Hiring flow

WithMira shows the role, then sends candidates to the company application.

1Check role fit, stack, and location eligibility in WithMira.
2Open the company application page from the tracked apply link.
3Save the role or subscribe for similar opportunities before leaving.