OpenAI took Jalapeño, its first custom chip for AI inference, from project start to tape-out in nine months, using its own artificial intelligence models to accelerate development, optimization, and verification stages. The achievement was highlighted by CFO Sarah Friar during the Goldman Sachs Communacopia + Technology conference.
Tape-out marks the completion of a chip's design before it is sent to manufacturing. According to OpenAI, its models helped explore different implementations and reduce design, measurement, and circuit verification cycles.
Jalapeño was co-developed with Broadcom, which handled part of the silicon implementation and network infrastructure. Celestica is involved in integrating boards, racks, and systems.
Chip will be used in OpenAI's infrastructure
The processor was developed specifically for inference of large language models, the stage in which already-trained models process requests and produce responses. The company plans to begin deploying Jalapeño in its infrastructure by the end of 2026.
Tests released by the company indicated between 1.5 and 1.9 times more processing per watt and end-to-end latency between 1.7 and 3.6 times lower than the commercial systems used in the comparison, depending on the model.
The tests included GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
OpenAI also competes on cost with Chinese models
Friar also said that the 80% price cut on Luna, OpenAI's lowest-cost model, helped boost its use by roughly tenfold.
According to the executive, running Luna can be cheaper than operating some Chinese open-weight models through cloud providers. She cited GLM 5.3, from Z.aias an example.
The comparison reinforces a broader OpenAI strategy of reducing the cost of running its models both in software and in the infrastructure used to run them.



