OpenAI Jalapeño Custom AI Chip Launched for Faster LLM Performance
OpenAI has shared initial performance data for Jalapeño, its first custom inference chip. According to an OpenAI post, benchmark results on InferenceX show the silicon delivers superior throughput and lower latency per watt across major language models, paving the way for deployment in its infrastructure.
OpenAI has released early performance data for Jalapeño, its first custom inference chip, demonstrating significant advancements in both processing speed and energy efficiency for large language models. Tested extensively across prominent model architectures, the specialized silicon is designed to eliminate common hardware bottlenecks by integrating compute, memory, and networking into a single unified architecture.
As per a post by OpenAI, the custom chip delivers higher throughput and lower latency simultaneously, overcoming traditional trade-offs seen in standard hardware systems. Evaluated on SemiAnalysis's public InferenceX benchmark, Jalapeño achieved substantial gains in performance per watt and end-to-end response times compared to leading commercially available platforms. OpenAI Previews GPT-5.6 Sol Ultrafast Mode Powered by Cerebras for 14x Speed.
Performance Benchmarks and Efficiency for OpenAI Jalapeño
The initial benchmark figures highlight strong results across multiple model families, including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Across these evaluations, the architecture yielded up to 1.9 times more AI work per watt at peak throughput and reduced end-to-end latency significantly.
For highly interactive workloads, which rely heavily on sequential step processing for autonomous agents, the chip recorded up to 4.1 times higher performance. By keeping model states and the KV cache local within a connected domain, the hardware minimizes data movement delays that typically slow down complex generation tasks.
Hardware Integration and Deployment for OpenAI Jalapeño
Developed in collaboration with partners like Broadcom, Jalapeño represents a full-stack design approach where models, software, and silicon are optimized together. Artificial intelligence tools were also utilized during the development cycle to accelerate chip design and optimize arithmetic circuits. Anthropic Claude AI Shared Memory Feature Launched for Free Pro and Max Users.
OpenAI plans to begin deploying Jalapeño within its compute infrastructure. The processor marks the initial phase of a multigenerational roadmap, with subsequent chip generations already in active development to support future scaling demands.
(The above story first appeared on LatestLY on Aug 25, 2026 11:49 PM IST. For more news and updates on politics, world, sports, entertainment and lifestyle, log on to our website latestly.com).