Roth describes Jalapeño as an application-specific chip built for large-language-model inference rather than a general-purpose GPU. He highlights early SemiAnalysis testing that reportedly showed strong performance per watt and high interactive token throughput, while noting that the reviewers had not completed their preferred full benchmark suite.
The more consequential claim concerns software. OpenAI reportedly lacked an existing multi-head latent attention kernel for DeepSeek R1, so Codex produced functional, efficient low-level code for the new ASIC. Roth explains that these kernels determine whether capable hardware can actually run a model efficiently.
Nvidia's CUDA moat rests on nearly two decades of libraries, tools and scarce engineering expertise. Roth argues that AI-generated kernels could let a new chip bypass much of the need for a large human-friendly software ecosystem, especially when a compact low-level language such as Gluon is paired with a capable coding model.
Roth frames the chip as an early recursive-improvement loop: models help design hardware, write its supporting software and then run on the resulting infrastructure. Small data-center volumes are expected first, so the long-term competitive impact on Nvidia remains uncertain despite the promising initial results.
Watch on YouTube


