OpenAI presented its first benchmark results for Jalapeño, its new inference processor developed in close collaboration with Broadcom, at the Hot Chips conference on Tuesday. Tested against SemiAnalysis’s InferenceX benchmark, Jalapeño outperformed current state-of-the-art systems, including an Nvidia Blackwell configuration, on both tokens served per user and throughput per kilowatt. Richard Ho, OpenAI’s head of hardware, said the chip represented a significant performance advance, capable of serving more AI workload per unit of power while also reducing response latency. First announced last October, Jalapeño is intended as a multigenerational platform integrating models, chips, and memory in concert, with small-volume deployment expected by the end of 2026 and broader rollout in 2027. OpenAI said the system was designed to minimize data movement and communication delays during the prefill and inference phases that typically create processing bottlenecks.
The hardware announcement arrived alongside a broader look at OpenAI’s push into agentic software through ChatGPT Work, a $20-per-month product built on its Codex coding tool and aimed at extending agent-based automation beyond software engineers to white-collar workflows. Andrew Ambrosino, lead engineer for OpenAI’s desktop app, described granting the tool extensive access to his own inbox, Slack, and other applications as necessary for testing its capabilities, while acknowledging inherent privacy tradeoffs. Thibault Sottiaux, who leads OpenAI’s core product work, framed the effort as central to the company’s mission of expanding AI’s usefulness beyond simple question-answering.
Engineers described early friction in adapting agentic tools built originally for coding contexts into general-purpose use, noting that non-engineering employees initially found the tools poorly suited to their workflows before adjustments were made. OpenAI acknowledged that ChatGPT Work remains far behind ChatGPT’s overall user base, with adoption still concentrated internally. Joe Gershenson, engineering lead for OpenAI’s harness team, said the company prioritizes minimal, precise tooling around its models rather than heavily engineered interfaces, betting that continued model improvements will reduce the need for elaborate scaffolding. Executives characterized the parallel hardware and software advances as reinforcing efforts to expand AI’s economic reach beyond coding into broader knowledge work.