New-ZZZ
RU / EN
Chips and Hardware 25 August 2026

OpenAI Details Its Full-Stack Compute Strategy

N
New-ZZZ desk
OpenAI Blog · 3 days ago

OpenAI says its AI strategy depends on improving the whole technology stack together, from data centers and chips to models, software, products, and AI devices. The company reported its first performance measurements for Jalapeño, a custom chip built to run trained models efficiently. On the public InferenceX benchmark, the chip reportedly delivered higher peak throughput per kilowatt and lower token latency than the commercial systems tested, with strong results across GPT‑OSS 120B, DeepSeek R1, and Kimi K2. OpenAI argues that designing chips, memory, networking, and serving software together can reduce cost and energy use while improving speed. It will keep using infrastructure from Microsoft, NVIDIA, AWS, AMD, and other partners, choosing different systems for different jobs. The broader goal is to maximize useful AI output per dollar, making demanding applications such as contract review, personalized analysis, and long-running agents affordable at greater scale.

Why it matters

  • OpenAI now has working custom silicon, giving it more control over the speed, energy use, and cost of running AI models.
  • A broader supplier portfolio could reduce dependence on a single chip or cloud provider and strengthen OpenAI's negotiating position.
  • More efficient infrastructure could make advanced AI services and long-running agents cheaper for businesses and consumers.

Key facts

  • Jalapeño is OpenAI's first custom inference chip, designed to run already-trained AI models.
  • OpenAI says Jalapeño achieved higher peak throughput per kilowatt and lower token latency than the compared commercial systems on InferenceX.
  • The chip was tested with GPT‑OSS 120B, DeepSeek R1, and Kimi K2, suggesting its benefits are not limited to one model family.
  • OpenAI plans to combine its own hardware with systems supplied by Microsoft, NVIDIA, AWS, AMD, and other partners.
  • The company says GPT‑5.6 Sol set a new high on a coding-agent index while producing 54% fewer output tokens than another leading model.
Read the original

The full text is in the original source. Here we provide a brief summary and key facts.

/ related