Baseten
Software Engineer — GPU Networking & Distributed Systems
Remote Model Performance role with clear candidate location fit.
PostedFeb 23, 2026
Eligible countries2 accepted countries
Seniority signalOpen level
Work settingRemote
Accepted candidate locations
CanadaUSA
Role overview
Software Engineer — GPU Networking & Distributed Systems
Requirements and responsibilities
Readable role content extracted into sections for faster review.
Details
- Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack, helping us move beyond TCP/IP to unlock order-of-magnitude improvements in bandwidth and latency.
- Optimize Distributed Inference: You will implement and tune the networking layers necessary for efficient Disaggregated KV Cache Offload and WideEP, ensuring seamless communication across NVLink and InfiniBand for our MoE models.
- Enable Serverless-Grade Startup Speeds for LLMs: You will work deeply with checkpointing and storage mechanisms to enable sub-10-second startup for trillion-parameter models.
- Deep-Dive into Hardware: You will characterize and validate networking performance on bleeding-edge clusters (H100/H200, B200/B300, GB200/300 NVL72), writing the acceptance tests that ensure our hardware delivers peak achievable throughput and minimal latency.
- Build Observability: You will design the tools that let us visualize packet flow, congestion, and effective bandwidth across the GPU interconnects, helping us diagnose complex distributed system behaviors.
- Optimize Kernels: You will work with communication libraries (NCCL, NVSHMEM) and potentially write custom communication kernels to overlap compute and data transfer.
- You have deep experience with high-performance networking protocols (InfiniBand, RoCE v2) and understand the physics of data movement.
- You are fluent in C++ or Python, with the ability to bridge the gap between high-level logic and hardware. You have a deep understanding of the memory hierarchy in modern NVIDIA architectures (H100/Blackwell) and know how to optimize for it.
- You like going deep. You aren't afraid to dive into TensorRT-LLM source code, write custom C++ / Python bindings, or debug NVLink topology issues.
- You know when to use an off-the-shelf solution and when we need to build a custom solution because the upstream tools (like standard Kubernetes networking) are too slow for our needs.
- Deep knowledge of NCCL, NVSHMEM, and UCX.
- Experience with GPUDirect Storage (GDS) or high-performance filesystems like Weka or 3FS.
- Familiarity with TensorRT-LLM, vLLM, or Sglang.
- Experience running low-level benchmarks to "qualify" new hardware clusters.
- Bleeding Edge Hardware: We are preparing to bring Blackwell (B200/B300) and then Rubin architectures online. You will be one of the first engineers in the industry optimizing networking for NVL72/GB300 racks.
- We go deep: We operate at every depth. Whether it’s tuning hardware interconnects, writing custom communication kernels, or designing distributed inference strategies, we work across the entire stack to deliver performance that goes far and beyond.
- High Impact: The networking optimizations you build will directly enable features that no one else in the industry has fully mastered yet, like seamless multi-node WideEP and instant model hydration.
- Competitive compensation, including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
Similar roles
Keep a backup shortlist.
Python, Spark USA
Senior Data EngineerTop Us Wealth Management FirmView role Kubernetes, Python 1 accepted country
Lead Platform EngineerAppianView role Kubernetes, Python 13 accepted countries
Senior Backend Engineer (AdTech)Leap ToolsView role Kubernetes, Python 13 accepted countries
Senior Backend EngineerLeap ToolsView role Stack
Use these tags to compare similar remote roles.
Location eligibility
Candidates should apply only when their profile country is listed here.
Hiring flow
WithMira shows the role, then sends candidates to the company application.
1Check role fit, stack, and location eligibility in WithMira.
2Open the company application page from the tracked apply link.
3Save the role or subscribe for similar opportunities before leaving.