An open compiler harness for coding agents that write GPU kernels, from the MLC community behind MLC-LLM and XGrammar. Its authors found that agents optimizing kernels spent much of their effort on work around the optimization itself: predicting how code would be lowered to hardware behavior, diagnosing timing-dependent failures, and coping with measurements distorted by other GPU activity. They define a compiler harness as the agent-facing environment around a compiler foundation, and TIRx Harness has four parts. TIRx is a deliberately minimal, hardware-close intermediate representation with a thin PTX-level programming surface, written through the TIRx-lite DSL, so that the PTX ISA serves as the main source of truth for instruction semantics. Compiler analyses check synchronization, data races and numerical behavior, backed by profilers (NCU and IKET) and by numerical simulation when no GPU is available. A knowledge base pairs a kernel zoo of more than 60 TIRx kernels with PTX documentation, keeping concrete implementations rather than agent-written summaries, which tended to overfit the last debugged case. And the KCoral benchmark server owns the GPUs and schedules every correctness test, benchmark and profile centrally, so concurrent agents cannot distort each other's timings.

On NVIDIA Blackwell GPUs, with web access disabled during optimization, agents using the harness produced kernels with family-level geometric-mean speedups of 1.33× to 6.84× over each family's optimized reference across Kimi Delta Attention (KDA), MiniMax Sparse Attention, Multi-head Latent Attention and Video Sparse Attention; KDA kernels reached 2.94× over FlashKDA in the forward pass and 6.84× over Flash Linear Attention in the backward pass. Example traces show an agent finding tcgen05.mma output-lane masking in the PTX ISA, which raised the observed GPU frequency for a 2–3% gain; adapting FlashAttention-4's mixed native and software exponential for about 2.9%; and replacing a serial 32×32 recurrence in KDA forward with Gated DeltaNet's hierarchical inverse decomposition for about 17%. In another, synchronization analysis caught a latent mbarrier bug in a candidate that had passed the correctness tests. The harness was released on 29 September 2026 as a PyPI package (tirx-harness) with a companion TIRx-kernels repository and an online book, Agentic GPU Programming for MLSys, and had about 70 GitHub stars a day later; the repository declares no license. Five of its seven GitHub contributors list Carnegie Mellon affiliations, and the post thanks "the NVIDIA CAKE team, Kernel Design Agent (KDA) team and SOL-ExecBench team".

Library

Language Python, Rust
Install pip install tirx-harness
agentsagent-harnessharnessgpu-kernelscompilersinfrastructureefficiencyopen-source

Related