OpenAI Announces Its First Self-Developed Inference Chip Jalape: There Are Actual Test Results, But This Does Not Mean It Can Completely Replace GPU
CoinMeta
2h ago
Ai Focus
On August 25th, OpenAI disclosed for the first time the measurement results of its self-developed inference chip Jalape. The company stated that, using a public InferenceX benchmark with a capacity of 120B GPT-OSS, this chip achieved a higher peak throughput per kilowatt and lower token latency compared to the commercial systems involved in the comparison; it also performed strongly on DeepSeek R1 and Kimi K2. OpenAI indicated that they currently possess a working first-party chip, and subsequent generations are also under development.
Helpful
No.Help

OpenAI在8月25日首次披露自研推理芯片Jalapeño的测量结果。公司称,在使用GPT-OSS 120B的公开InferenceX基准中,这款芯片相较参与比较的商用系统实现了更高的峰值每千瓦吞吐量和更低的token延迟;在DeepSeek R1与Kimi K2上也表现强劲。OpenAI表示,当前已经拥有可工作的第一方芯片,后续代际也在开发中。

这次发布的准确边界是“首批测量结果”,不是Jalapeño已经大规模铺进所有数据中心,更不是训练和推理都迁移到自研硬件。公开文章没有给出量产规模、制程、单卡功耗、内存容量、系统售价或完整部署时间表,也没有说明比较系统的全部配置。可以确认的是OpenAI已经把模型、服务软件、芯片、内存与网络协同设计从规划推进到可测硅片阶段;不能确认的是它在真实业务总拥有成本上已经普遍领先。

每千瓦吞吐与token延迟要同时看,峰值不能代替持续负载

推理芯片竞争常被简化为每秒生成多少token,但大规模服务还受到首token时间、token间延迟、并发用户、批处理大小、模型精度、上下文长度、内存带宽和功耗约束。OpenAI特别强调在既定token间延迟下扩大吞吐,以及用每千瓦吞吐衡量能效,说明Jalapeño瞄准的是持续运行的推理经济性,而非只追求单一峰值算力。

使用GPT-OSS 120B作为公开基准有利于外部比较,因为测试对象不是公司内部不可获得的私有模型;DeepSeek R1和Kimi K2的结果则用于证明优化并非只绑定一个模型家族。不过,新闻稿没有同步公开足够细的原始数据、软件版本、精度模式、批量参数和复现实验脚本。不同系统若在量化精度、延迟目标或机架功率定义上不一致,排名可能发生变化。因此现阶段更适合把结果视为第一方性能声明,等待InferenceX条目和第三方复测补充。

芯片与模型共同优化能减少通用硬件为兼容众多工作负载保留的冗余,但也带来生命周期风险。模型架构变化速度通常快于芯片设计和制造周期,今天为注意力、混合专家或特定内存访问方式优化的硬件,未来可能面对不同计算图。OpenAI要证明的不只是首代芯片能跑得快,还要证明编译器、运行时与软件栈可以随模型变化保持利用率。

能效也不只等于芯片功耗。机架网络、内存、冷却、电源转换和低利用率待机都会影响数据中心实际每token耗电。峰值每千瓦吞吐是重要指标,却不能直接推导全年平均成本。真正决定商业价值的是在真实请求波动、长上下文、工具调用和多模型路由下,系统还能否保持同样优势。

自研芯片增加的是供应选择与议价能力,不是结束合作

OpenAI把Jalapeño定位为合作伙伴加速器之外的一条可信第一方路径。公司明确列出Microsoft与NVIDIA对其增长的基础作用,并表示计算组合还包括AWS、AMD、Broadcom、Cerebras、CoreWeave、Oracle、SB Energy和SoftBank。不同训练、批量推理、低延迟推理和常驻智能体任务需要不同硬件,企业会把负载放到性能与成本最匹配的系统,而不是只押注一种芯片。

对OpenAI而言,自研芯片的战略价值有三层。第一,针对自身高频推理形态优化,降低边际服务成本;第二,在供应紧张时增加可控产能路线;第三,通过可替代选择维持采购议价能力。即使Jalapeño占比不高,只要能承接足够多稳定工作负载,也可能影响外部加速器价格和部署节奏。

这种控制力仍依赖制造、封装、HBM、网络与数据中心建设伙伴。所谓“自研”通常指架构和系统设计由公司主导,不代表所有物理环节都由OpenAI完成。供应链良率、软件迁移成本和新代际迭代都会决定它能否从测量芯片走到稳定生产。文章没有公布量产日期,因此不应把“已经有工作硅片”写成“已全面商用”。

对云服务和API用户而言,Jalapeño短期内未必改变产品接口。更高能效可能为更低延迟、更高配额或更低单位成本创造空间,但是否传导到价格由容量、需求和产品策略共同决定。用户更应关注服务等级、模型一致性和隐私边界,而不是假设自研芯片会自动带来降价。

这次披露真正改变的是OpenAI在计算栈中的位置:它不再只是大型加速器买家,而开始拥有可测量的第一方推理硬件。下一阶段需要验证的不是宣传中的“领先”,而是公开基准能否复现、真实集群能否长期稳定运行、部署规模如何扩大,以及公司能否在继续使用多家伙伴的同时,把自研硬件变成持久而非一次性的经济杠杆。

来源:OpenAI,The full stack behind abundant intelligence,2026年8月25日,https://openai.com/index/the-full-stack-behind-abundant-intelligence/

Tip
$0
Like
0
Save
0
Views 14
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
LayerZero Launches the ATLAS Market Backend: 200,000 TPS was the initial plan, not the actual transaction figure in the live network
LayerZero released ATLAS on August 25th, which is Aggregated Trading Liquidity and Settlement. It is positioned as a general-purpose backend for the trading market. It does not have its own consumer frontend; users will access it through independent trading venues. The trading venues are responsible for the interface, distribution, and customer relations. ATLAS provides the underlying capabilities for matching, liquidity, and settlement. Zero records the ultimate ownership in the network, while ZRO ensures security, Gas, governance, and cost efficiency.
币界网
·2026-08-26 09:57:01
45
EU's goods trade turned into a deficit of 21.8 billion euros in the second quarter: Widening energy gap is the main reason
Data released by the European Statistical Office on August 25 showed that in the second quarter of 2026, the EU's imports of goods from non-EU countries amounted to 701.8 billion euros, while exports were 680 billion euros, resulting in a trade deficit of 21.8 billion euros. This is the first quarterly deficit since the second quarter of 2023; previous deficits mainly occurred during the period of sharp increases in energy costs from late 2021 to mid-2023.
币百科
·2026-08-26 09:56:48
18
Jetson Orin Nano 2 Entry-Level Robots: The Conditions for Implementation Behind 78 TOPS and 15 Watt Energy Efficiency
NVIDIA released on August 25th the Jetson Orin Nano 2 robot computer, aimed at entry-level edge AI, robots, delivery and inspection drones, as well as visual AI systems. The official specifications include 78 trillion operations per second, 8GB of memory, and an 8-core Arm CPU; compared to Jetson Orin Nano Super, the inference performance has been doubled, while the physical dimensions remain unchanged. In 15-watt mode, it operates with 40% less power consumption while maintaining the same performance.
CoinMeta
·2026-08-26 09:56:15
17
Robot AI and startup Generalist valued at $3 billion
Generalist reportedly completed nearly $200 million in additional financing, raising its valuation to $3 billion, reflecting continued capital betting on the AI model of general-purpose robots.
TechCrunch
·2026-08-26 08:48:54
20
View More