摘要摘要 1.超大规模超大规模AI厂商资本开支厂商资本开支:预计将从2026年的8000亿美元(亿美元($800B))增长至增长至2028年的年的1.4万亿美元万亿美元(($1.4T)),谷歌和亚马逊合计将占据总支出的50%-60%。2.AI业务盈利表现业务盈利表现:专有基础设施上的模型API业务盈利能力最强;GPU租赁(IaaS)的投资回报率超过30%((>30% ROIC)),增量EBIT利润率约70%((~70%))。3.全球算力总容量全球算力总容量:预计到2028年将达到115-120GW,亚马逊将保持整体算力规模第一,谷歌预计将在2026/2027年超越微软。4.业务负载结构业务负载结构:推理工作负载占总算力的比例从50%-60%持续提升,但训练仍是前沿模型保持竞争力的核心;像GB300这类芯片可以同时适配两种任务,通用性极强。5.AWS边际利润边际利润:AWS约**50%(~50%)**的增量边际利润来自于裁员、H100产品涨价以及高盈利的Bedrock会计核算业务,长期利润率预计将稳定在40%左右。6.内存成本趋势内存成本趋势:Rubin新一代服务器架构中,内存成本将占到机架物料成本的25%,但高投资回报率框架意味着超大规模厂商愿意承担这一成本上涨。7.开源模型影响开源模型影响:开源模型会降低每GW算力的营收,但算力规模增长以及定制芯片(TPU/Trainium)上更高的令牌吞吐量将抵消利润率的压缩。 Q&A Q1: What is the current outlook on hyperscaler CapEx, compute capacity, and the potentialROIC for Generative AI investments? The ecosystemis currentlycompute-constrained, which is driving continued capital expenditure growth. For the four largesthyperscalers, total CapEx is projected to reach$1.2 trillion in 2027and$1.4 trillion in 2028, with an upward bias to thesefigures. However, 2027 is expected to be the final year of significant second-derivative acceleration in CapEx for this cycle. This investment will lead to afourfold increase in compute capacity between 2025 and 2028. The ROIC is analyzed throughthree emerging business models: 1.Hyperscaler GPU rental business: Could achieve after-tax ROICs of over 30% with incremental EBIT margins in the70% range.2.Model-enabled APIs running on proprietary data centers: Such as Google's Gemini, represent the most profitablemodel with the highest incremental margins and ROIC.3.Third-party capacity rental models: Can still generate strong returns, with incremental EBIT margins of 30% or higher andROICs exceeding 25%. Q2: What are the specific CapEx and compute capacity projections for the major hyperscalersthrough 2028? For 2026, total CapEx across Google, Amazon, Microsoft, Meta, and SpaceX's terrestrial compute is estimated at nearly$800billion. This is projected to grow approximately 60% to between$1.2 and $1.3 trillion in 2027, and further to$1.4 trillion in2028. Google and Amazon are expected to account for 50-60% of this spending in each year. In terms of compute capacity: 1.These five companies are forecasted to double their capacity from10 gigawatts at the end of 2025 to around 20 gigawatts bythe end of 2026.2.They are projected to add another 30 gigawatts in 2027 and 35 gigawatts in 2028, bringing the total incremental capacity toapproximately 85 gigawatts by year-end 2028.3.This will result in an aggregate capacity of 115 to 120 gigawatts. Google is expected to add the most incremental capacity, slightly ahead of Amazon. However, due to its existing lead, Amazon isprojected to maintain the largest overall compute capacity through 2028, followed by Google, which is expected to surpassMicrosoft Azure's capacity during 2026 or early 2027. 问题问题1:生成式:生成式AI三大三大ROIC框架的假设与关键变量框架的假设与关键变量 本分析基于NVIDIA GB300硬件开展。 1.基础设施即服务(基础设施即服务(IaaS)框架)框架:预计每吉瓦可实现230亿美元年收入亿美元年收入,该测算基于每小时8.50美元的GPU租赁价格与GPU运行时长假设。成本端方面,每吉瓦的资本开支(CapEx)为390亿美元亿美元,其中230亿美元为使用寿命5年的IT设备开支,160亿美元为使用寿命15年的非IT基础设施开支,外加每瓦约2美元的能源与其他运营支出(Opex),最终可实现30%的的ROIC。2.自有基础设施的模型自有基础设施的模型API框架框架:该模式可获取溢价API收入,无需承担高昂的GPU租赁成本。该框架的关关键变量键变量包括:算力在创收推理任务与非创收训练任务之间的分配比例,以及令牌吞吐量与令牌定价之间的权衡关系。3.第三方基础设施的模型第三方基础设施的模型API框架框架:该模式预计每吉瓦的收入更高,因为其租用算力的GPU运行时长为100%,而自有基础设施的GPU运行时长仅为75%。该框架的核心成本驱动因素核心成本驱动因素为每小时的GPU租赁价格。 问题问题2:开源模型渗透率提升对超大规模云服务商:开源模型渗透率提升对超大规模云服务商ROIC的影响的影响 开源模型仍然需要算力支持,而AWS这类超大规模云服务商正是算力的提供方。超大规模服务商不太可能达成单位经济效益不佳的合作项目。尽管开源模型的每吉瓦收入可能更低,但整体的底线收益预计不会为负。哪怕利润率从30%下滑至20%区间,也完全可以通过算力规模的大幅增长抵消利润损失。此外,开源模型通常体量更小,加载所需的参数更少,因此可以实现更高的令牌吞吐量。这种低定价与高吞吐量的权衡低定价与高吞吐量的权衡是核心的市场动态逻辑。超大规模服务商还可以利用定制化芯片(比如TPU)进一步优化这类小型开源模型的吞吐量。 问题问题3::GPU租赁费率的演化趋势与租赁费率的演化趋势与Meta潜在潜在Neocloud业务前景业务前景 尽管算力市场在一段时间内仍将处于供给紧张的状态,但近期出现的每瓦每瓦50美元的极高租赁费率美元的极高租赁费率并不具备可持续性,随着更多算力产能上线,租赁费率预计将出现下滑。供需失衡仍将维持算力价格的坚挺,但整体费率将回归合理水平。 一、一、Meta的新云业务战略定位的新云业务战略定位 exorbitant rates are considered temporary.The ROIC models do not depend on these peak rates to be viable. RegardingMeta's neocloud opportunity, it is viewed as a "Plan B." The company's primary focus for its capacity is on its core products andnew initiatives like messaging and the Meta AI API. While Meta is open to a neocloud offering, likely utilizing older chips like H100srather than Blackwell, it is not the primary strategy. If pursued, the neocloud could establish a higher floor for earnings and the stockprice, butthe key driver for Meta's long-termvaluation will be the successful launch and scaling of its own AI-drivenproducts. 二、训练与推理负载的长期结构展望,以及芯片跨任务复用性二、训练与推理负载的长期结构展望,以及芯片跨任务复用性 What is the long-termoutlook for the mix between training and inference workloads, and howfungible are the chipsbetween these tasks? The proportion of compute capacity dedicated to inference is expected to continue increasing fromthe current level of around 50-60%. However, it is unlikely to reach 100% in the near future, as labs will continue to allocate significant resources to training tomaintain model compet