Key Takeaways

  • On ParallelKernelBench (87 problems), the best frontier model solved 28 zero-shot, 22 faster than the PyTorch + NCCL reference; sampling raised correctness to 36 but left the "fast" rate near 31%.
  • Arora attributes the gap to reasoning, not syntax — models compile fine but fail on collective ordering, data partitioning, intra- vs inter-SM scheduling, and transfer-mechanism choice.
  • From A100 (2020) to B200 (2024), BF16 tensor core speed rose 7.2x while intra-node communication rose 3x and inter-node just 2x.
  • A mini-SWE-agent harness with Gemini 3 Pro took one model from 24 to 35 of 87 problems, 26 above 1x speedup — then plateaued as time scaled.

Frontier models can write CUDA that compiles. They cannot yet reason about the network. That is the short version of Simran Arora’s AI Engineer talk: on ParallelKernelBench, her team’s 87-problem benchmark for multi-GPU kernel generation, the best frontier model they tried solved 28 problems zero-shot, and only 22 of those beat a plain PyTorch-plus-NCCL baseline. Scaling test-time compute pushes correctness to 36 — but the share of kernels both correct and faster than the reference plateaus around 31%.

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI Can LLMs Write Fast Multi-GPU Kernels?Simran Arora, Together AI · AI Engineer · Watch on YouTube

The bottleneck moved, and most tooling hasn’t followed

Arora is a principal scientist at Together AI, where she leads the frontier performance research team; she did her PhD in the Hazy Research lab with Chris Ré at Stanford and is an incoming professor at Caltech. Her framing is that the field solved the wrong-shaped problem very well. GPU utilization used to be capped by poor intra-GPU memory access and weak single-GPU kernels — then came FlashAttention, memory-efficient architectures like DeepSeek’s, sparse attention, Mamba, kernel DSLs from Triton to ThunderKittens, and megakernels. In her words, “we’ve sort of shifted the bottleneck to multi-GPU communication.”

The hardware trend makes that shift permanent. Comparing NVIDIA A100s in 2020 to B200s in 2024, Arora said BF16 tensor core speeds improved by 7.2x while intra-node communication improved by just 3x and inter-node by 2x. On production distributed training and inference workloads, she said, communication is increasingly consuming the majority of runtime and yields low model FLOP utilization at scale.

Networking is also fragmenting in a way memory hierarchies never did: AMD uses XGMI point-to-point links, TPUs a 3D torus with optical wraparound links, NVIDIA the NVLink/NVSwitch pair — up to 900 GB/s unidirectional between two GPUs on the generation she cited, with multicast and reductions runnable inside the switch itself. Scale-up domains keep growing, from 72 GPUs in coming chips to a single 576-GPU system NVIDIA plans for 2027 — the capital-intensive hardware arms race, seen from the kernel writer’s desk.

Why the default stack leaves performance on the floor

Arora’s team first measured what standard tools achieve. NCCL is the library most teams lean on, and she credited the engineering investment behind it — but it is tuned for bulk transfers of large contiguous chunks. “The design really breaks down when you care about peak performance, fine grain communication, and sort of non-trivial collectives that you want to fuse together,” she said. Across their benchmark, a naive baseline representative of PyTorch plus NCCL falls below 50% of its communication-aware roofline bound on the majority of problems, and frameworks layered on NCCL — Megatron-LM, FlexFlow, NanoFlow — inherit the pattern.

The alternatives each have their own failure mode. Compilers and DSLs struggle to track how fast networking hardware changes; Arora said Triton-distributed, originally tuned around 8 H800 GPUs, fails to adapt efficiently to architectures like H100s. Hand-tuning operators one by one does reach peak performance, but she noted some of those methods were designed for one precision and take five or six months to scale to another.

The small set of trade-offs a kernel writer faces

Arora’s research question was whether a small set of principles governs multi-GPU kernel writing. Her team did the work by hand first: “we think it’s important to build our own fundamental understanding and to manually do the work to understand it rather than just throwing say an LLM at the problem.” That produced ParallelKittens, a set of primitives she said usually adds roughly a dozen lines of code on top of a single-GPU kernel, now used in production at Together AI and at partners including Cursor.

The space has two axes. The first is the transfer mechanism:

  • Copy engine — host-initiated, strong for large messages, and it burns neither registers nor SMs, leaving both free for compute.
  • Device-initiated transfers via TMA — saturates NVLink bandwidth with small messages using very few registers and processors, the tool for fine-grained overlap. Its limitation: it cannot use NVSwitch’s in-network computation.
  • Register-level instructions (the multimem load/store/reduce family in PTX) — the mechanism that can exploit NVSwitch’s in-network reductions.

The second axis is scheduling. Intra-SM specializes warps inside one SM for compute and communication concurrently, which works when the two patterns align on the same input data. When they do not, inter-SM specializes whole SMs for compute, communication, and memory. Her example: GEMM plus reduce-scatter favors intra-SM, while GEMM plus all-reduce favors inter-SM, because that path leverages NVSwitch’s in-network reductions. That is the entire curriculum — exactly the kind of bounded reasoning problem you would expect a frontier model to handle.

ParallelKernelBench, and what the scores say

Each task hands the model an unoptimized PyTorch reference using torch.distributed NCCL operations, plus a system topology specifying rank count and intra-node hardware configuration. The model must rewrite it into a performant CUDA kernel using unified virtual addressing. The 87 problems were drawn from GitHub repositories, optimized library implementations, and DSL kernels, then organized by a taxonomy — a transformer layer can be parallelized across data, sequence, tensor, context, layer, pipeline, and expert dimensions, each composition inducing a different communication pattern. Arora stressed that solutions should be net-new production kernels, not artificial ones.

Two metrics: pass@k (correct after k attempts) and fast_1@k (correct and faster than the PyTorch + NCCL baseline). The headline results:

  • Zero-shot, the best frontier model tried solved 28 of 87, with 22 beating the baseline.
  • Parallel sampling raised correctness to 36, but the fast metric plateaued at roughly 31%, with little room for gains from more generations.
  • On the speedup-threshold curve, the best model benchmarked, GPT-5.5, drops off very quickly as the required speedup rises; DeepSeek V4 Pro sat at the bottom.

Where speedups did appear, the mechanism was consistent: once correctness is established, gains come from eliminating NCCL staging in favor of direct NVLink loads and stores.

The failure is reasoning, not CUDA

The most useful result is diagnostic. “We found that there’s deeper issues than CUDA syntax,” Arora said. Multi-sampling or letting the model fix its own errors gets kernels to compile. What models cannot do is reason through the trade-offs above — collective ordering, data partitioning, intra- versus inter-SM scheduling, transfer mechanisms. In practice, she said, they often simply do not reach for the register-level transfer instructions or TMA at all.

The successes cluster where you would expect memorization to carry a model: collective primitives, tensor-parallel GEMMs, and Ulysses-style context parallelism. Those are, in her framing, patterns “heavily represented on the internet rather than necessarily patterns that the model has used its reasoning abilities to think through.” Recall dressed as reasoning is the transferable lesson, and anyone building evaluations for their own team should note that the second metric, not the first, exposed the ceiling.

Her team also tested the obvious next move: an agent loop. Using the mini-SWE-agent multi-turn harness with Gemini 3 Pro and a local bash environment — deliberately mimicking a standard Claude Code setup — the model went from 24 problems to 35 of 87, with 26 above 1x speedup. But as they scaled the time budget, performance plateaued again, and Arora said additional techniques would be needed to keep scaling. Iteration buys correctness, not insight — a sharper constraint than most self-improvement narratives acknowledge.

What happens next

Arora kept both halves of her conclusion. The optimistic half: there are not many patterns involved in writing effective multi-GPU kernels, and her team encapsulated them in a small set of primitives. The deflating half — models “do not currently understand how to reason through these trade-offs even when we provide them in context.” Note the qualifier: this is not a retrieval failure a better prompt would fix. The fundamentals were in the context window and the models still did not apply them.

There are signs of life. Arora pointed to net-new kernels emerging where nobody had hand-tuned before: a NeMo vocab-parallel filtering kernel, a Hyena-architecture context parallelism kernel, and an IOU suppression kernel for the SAM 3 video segmentation model. Her invitation was for others to attack ParallelKernelBench and design architectures that grow with where networking is heading — larger scale-up domains, a shift away from scale-out, massive on-chip memory.

For engineers the practical read is narrow. If your workload is communication-bound, an LLM will get you a distributed kernel that compiles and may hand you the easy win of removing NCCL staging. It will not make the decisions that separate a 1x kernel from a 3x one. That still needs someone who has read the fundamentals — the same division of labor teams are drawing as they fold coding agents into engineering workflows.

Quick poll

Where is your biggest performance bottleneck right now?

Arora said that from A100 to B200, tensor core speed rose 7.2x while intra-node communication rose only 3x.

FAQ

What is ParallelKernelBench? An 87-problem benchmark from Together AI. Each task gives a model an unoptimized PyTorch + torch.distributed reference and a system topology, and asks for a performant CUDA rewrite using unified virtual addressing.

How well did frontier models do? Zero-shot, the best model tried solved 28 of 87, with 22 faster than the baseline. Parallel sampling pushed correctness to 36 but the correct-and-faster rate plateaued around 31%. With an agent harness, Gemini 3 Pro went from 24 to 35 solved, 26 above 1x.

Why do the models fail if they can write CUDA? Arora’s diagnosis is that the problem is not syntax — kernels compile after retries. Models struggle with collective ordering, data partitioning, intra- versus inter-SM scheduling, and transfer mechanisms, and often skip register-level multimem instructions and TMA entirely.

What is ParallelKittens? A minimal set of primitives for multi-GPU kernels, adding about a dozen lines over a single-GPU kernel. Arora said it is used in production at Together AI and at partners including Cursor.