Nvidia: Q2 Preview
Deep-dive into what this quarter means for the stock
Brief intro
Nvidia isn’t just a GPU seller; it’s a systems provider. It sells hardware (GPU, CPU, high-speed interconnects like NVLink/InfiniBand), software (CUDA platform) and AI-specific solution (DGX Cloud, AI Enterprise, Omniverse)…
Nvidia has a triple moat.
They offer the best-performing hardware products in the industry, backed by solid execution and the fastest design cycles.
Adding to this, they have a multi-year first-mover advantage on the software side, which is their biggest moat. Most of today’s AI code run on CUDA optimizations which only Nvidia GPUs support (in the future ROCm by AMD will be able to translate from the CUDA architecture).
Their ability to provide completely integrated platforms with their best in-class hardware, software and networking…
The lead-up to earnings has been pretty crazy, with the yen carry trade, recession fears, and export controls shaking things up. On top of that, a lot of investor anxiety around ROI on AI capex, delays in Nvidia’s Blackwell offerings made it a very volatile stock dropping by -30% and up 30% from its low within a month and a half !
Recap of last quarter
Last quarter was very strong for Nvidia, with solid results, they beat the quarter by $1.5B and guided $1B above consensus… Alleviated air pocket fears and the stock was up +10% the next day. Large CSPs represented mid-40% of data center revenues.
They also relieved some anxieties over ROI on capex investment by claiming that for every $1 spent on NVIDIA AI infrastructure, cloud providers have an opportunity to earn $5 in GPU instant hosting revenue over four years.
They gave out some numbers around the sovereign AI opportunity: Japan spending would be $740m this year and they expect sovereign AI to be around high single digit billion dollars this year.
They quantified their inference opportunity, which was much bigger than investors expected at ~40% of data center revenues.
Finally, they talked up their upcoming Blackwell product line, which has 25x lower TCO relative to Hopper and expectations that we would see “a lot” of Blackwell revenue this year… more on this in the next section.
Debate on the stock
Q: Impact of Blackwell delays on CY24 numbers?
This seems to be the key focus for the quarter. Every quarter there’s been strong concerns around Nvidia’s numbers… 1Q24 was risks of an air pocket, 4Q23 the issue was sustainability of China revenues… this quarter it’s how does the Blackwell delay impact CY24 numbers
Let’s dive in this new problem. First, what is it all about?
Nvidia is struggling to ramp up high-volume manufacturing (HVM) of the Blackwell product line. They were expected to ship in Q3/Q4 this year and it seems like this has been delayed by a couple months.
The reason for this delay is related to Nvidia’s back-end design. This is not a front-end design issue, which would take much longer to solve. In essence, Hopper architecture uses CoWos-S while Blackwell will start using CoWos-L, which is a new technology for TSMC that can handle more chips and higher memory stacks… CoWoS-S ramp up was much smoother because TSMC had been using CoWos-S for years with Xilinx.
The back-end design issue relates to bridge dies, which are connectors between different chips within a package and enable chip to chip communication. The bridge dies were not positioned correctly on the Blackwell chip which led to errors during the manufacturing process at TSMC and has led Nvidia to a re-design of the bridge dies.
To solve this issue, Nvidia has introduced a new de-tuned chip called the B200A which is based on a B102 monolithic die and allows the chip to be packaged using CoWos-S instead of CoWos-L. They’ve also cancelled the B100 chip.
Chips are important but the focus for Nvidia is systems; more specifically their GB200 NVL 36/72 racks. While you can still get the HGX 8-GPU systems, those are generally better suited for smaller models that don’t need as much memory. The MGX system is likely to be the most popular choice because of its flexibility; meaning CSPs can pick which components (ex: NIC cards) they want to put in their system. The choice of system has an impact for Nvidia numbers as the average GPU content comes closer to $40k for MGX vs ~$50k for DGX… (This is still higher than the $30-$35k you'd pay for a single B200 chip).
I'll break down the differences between these platforms below.
Brief overview of systems:
The DGX system is the reference architecture for Nvidia combining 8 GPUs, NVLink, high-end x86 CPUs and some other high-speed networking. You can have both any kind of Ampere, Hopper or Blackwell based DGX systems. And it is also scalable with the DGX SuperPOD for example, which is connecting 36 Grace CPUs and 72 Blackwell GPUs into one system by connecting all the different racks using InfiniBand. The DGX is fixed so it is put together by a company like Foxconn and then sold directly to a Microsoft.
The HGX system, usually suits companies that need more customization than the DGX platform. It offers similar performance than DGX but customers can modify the CPUs, RAM, storage, and networking configuration. It’s the original system which is 4 or 8 GPUs per rack.. these could be H100/H200 but also the newer B200A on CoWoS-S were also built to be used on the HGX platform… The HGX platform can either be sold to hyperscalers (Google, Azure, AWS…) via one of the major ODMs (Wiwynn, Inventec or Quanta) or sold to the more enterprise clients (Stability.ai, Tesla..) via Supermicro.
The MGX system provides a more modular architecture means it is highly customizable and each ODM has their own client. HonHai/Foxconn is mainly partnered with AWS and MSFT, while SMCI is mostly partnered with Tesla/Coreweave…
This leads to the dual debate on the stock:
In CY24, what’s the impact of Blackwell delays and can Hopper fill in the gap?
In CY25, What will the product mix between DGX and MGX?
1 / How does the Blackwell delay impact numbers?
For the moment, it seems unlikely that this supply issue changes CY25 numbers, but it could change CY24 numbers, depending on how long the delay is. Keeping in mind Blackwell would have mostly been Q4 revenues, with no material revenues in Q3. So, this quarter, the focus won’t just be so much on the numbers; but on commentary around potential delays and how that might affect the year end.
Let’s take a look at some of the commentary from ODM/OEMs to get an idea as to the impact.
No pushouts
Asia Vital Components (sells server chassis & cold plate modules) sees no delay for the GB200 superchip. They were told about the B100/B200 cancellations and B200A replacement.
KYEC (provides testing equipment to Nvidia): Nvidia asked KYEC to build its new tester on time… aka no delays
Quanta sees minimal impact from Blackwell delay..
Pushouts
Super Micro talked about pushouts from Blackwell.. said it didn’t expect any volume in Q3 and they expect volume in Q4 but will be very small.
Hon Hai said GB200 rack shipments would commence in small volume in 4Q24, which is 2-3 months later vs prior expectations, with a stronger ramp in 2025.
The ODM/OEMs are suggesting pushouts from Blackwell vs individual component suppliers or equipment vendors are not seeing any change… which suggests delay is real and they need ramp up on B200A or H200 to fill in the gap in revenues.
Can Hopper be enough to offset Blackwell delay?
Some sell-side analysts were forecasting Nvidia would ship around 500k B100/B200 chips this year. With the delays, this number should come closer to ~100k in volume; That's a $14B gap that Hopper needs to fill.
We are hearing from supply chain Hopper is ramping up faster than expected with 400k capacity per month in Q2 (vs end of Q2 expected) and 600k units per month by end of Q3 (vs 500k previously expected). That’s 300k extra capacity priced in at $25k means $7.5B of incremental revenue, which would still leave us with ~$6.5B hole to fill.
I don’t think it would make sense for TSMC to go back to CoWos-S to make more Hopper chips as most of 2025 will use CoWos-L for Blackwell. Instead, what they could do, is outsource more of the H20 packaging to Amkor, which would free up additional H200 capacity so that they can fill this gap.
This is some back of the envelope maths right there, but it gives us an idea how much Hopper they need to sell. No matter how you slice it, Blackwell's delay puts a lid on how much Nvidia can beat each quarter (at least for this year) and the element of surprise is getting smaller and smaller each quarter, I think we should be fine this quarter but the setup in Q3 will be quite tricky.
Q: Past the Blackwell delays, how sustainable is the capex spend? What’s the ROI on investments?
It’s fair to say one major topic of interest regarding Nvidia and the broader AI space has been whether or not hyperscalers are getting their money’s worth from all this spending. They might have deep pockets, but they’re spending close to 30% of their cloud revenues on capex, which is not nothing - and can’t continue forever. Investors want clarity on the revenue generated from all this spend and are rewarding only companies that can generate revenue from it..A great example of this is how the market is now rewarding Meta for increasing its capex this quarter because it’s seeing AI drive engagement vs earlier this year stock when the stock got severely punished for increasing its spend on AI.
To start off, some good bearish papers on this topic have been AI’s $600B Question by Sequoia and GenAI: Too much spend, too little benefit by Goldman Sachs both highlighting the massive spending vs the very little tangible benefit we’ve seen so far.
Although I agree with the premise that the amount of capex currently spent on AI is astronomical, I do think investors are setting their expectations too high. Companies like Microsoft are already generating substantial revenue from AI (This quarter AI contributed 8 points to its 30% growth).
You could also compare AI revenue generation by Microsoft vs Cloud when they first entered the business. A good representation of this is in the GenAI paper by Goldman, which shows it took ~7 years for Microsoft to generate $3.8B of cloud revenues vs only 1 year for AI.
I think that’s the best example of ROI on AI.
If we look at META it’s a bit of a different story. They invest their capex in both their core advertising business and in GenAI. Within their core business, they’ve introduced AI features like Advantage+ which is driving 22%+ higher ROAS for US advertisers.. their new video recommendation system/ranking has already increased engagement on Facebook and Instagram reels helping drive their topline last quarter…
They’ve also integrated Meta AI (powered by Llama 3) into the Instagram/Facebook search bar, as well as in WhatsApp and FB messenger.. and that has already been used for billions of queries. The ROI here will take longer to trickle into their P&L as META is playing a different game. They’re trying to make the production of content as inexpensive as possible using open-source LLMs which would help drive engagement. They’re essentially providing cheap supply for content creators to use their platforms - their love for open-source models isn’t the only reason they’re spending billions on infrastructure.
Some datapoints on AI adoption:
Amazon: “we shipped 79% of the AI code reviews… this saved us the equivalent of 4500 developer-years of work”
Walmart: “we’ve used LLMs to improve our catalog… this work would have required nearly 100x the current headcount.”
Another way to look at it is Doug O’laughlin’s framing of AI capex as a Dollar Auction.
Whoever is first to the model that creates a killer app so good that it makes information processing (Microsoft), information search (Google), or sharing things online (I guess Meta) obsolete, then the AI Capex problem is best explained as a dollar auction. And your risk of not investing enough is much worse than investing a bit too much. Whoever wins will get the entire dollar.
This echoes the comments heard during Google’s earnings call recently but under-investing poses a greater risk than over-investing… and the winner of this race gets everything.
I’ll finish this part with a reminder that we are still in the very early stages of AI… The US census bureau published a paper earlier this year on AI’s use within enterprise and the results are… quite surprising. Only 5.4% in Feb24 of companies are currently using AI (from 3.7% in Sep23) and is expected to end the year at around a ~7% penetration rate.
If you compare that to the internet, 6% penetration rate would put you before the 2000s. That’s where we are. The levels of monetization we are seeing only reflect 1) the first year of AI infrastructure buildout 2) only represents an environment where ~5% of enterprise are using AI.
Q: Is Nvidia in a bubble? Comparisons with dot-com..
Nvidia is one of the best quality companies to own. They do 70% gross margins, 55% free cash flow margins at $100B+ revenues, growing 50% next year and they’re trading at a low 30 P/E multiple on a CY25 basis (40x CY24) compared to a Cisco which traded at 132x times at its peak.
It has a huge technological advantage, focuses its R&D on GPUs while its closest competitor (AMD) is using a smaller amount of R&D and shares that spend across CPUs, GPUs and FPGAs… Custom silicon is also a competitor but as we’ll see in a later section, there is a large gap in performance between custom silicon and Nvidia’s latest products, and this competitive advantage is only getting larger with Nvidia’s 1 year design cycle.
Q: Competition from custom silicon / AMD?
That’s more of a longer-term question; we’ve been hearing that the latest Gemini LLM was trained on TPUs so could it be possible that hyperscalers start using their own solutions vs buying off-the shelf processors from Nvidia?
To answer this question, let’s take a look at the different AI chips, their pricing and performance.
As you can see the performance to cost ratio for Nvidia’s latest Blackwell chip is 9x, which is much higher than any other chip in the market, and by a wide margin. Even the older generation Hopper chips remains one of the best AI chips on the market and Blackwell is seeing 25x better total cost of ownership (TCO) and energy consumption. And hardware isn’t even their main competitive advantage - software is. Programmers have been coding on CUDA for the past 10 years, building the libraries and there is still no software that can translate all this code, even though AMD will succeed in that one day.
Now, that’s just one metric and we shouldn’t directly extrapolate on products based off 1 metric but one thing is clear, the B200 chip has the best ROI out of all AI chips in the market. Additionally, any additional space or energy constraints that are facing Nvidia’s data center customers is actually improving demand for Nvidia’s products. Given the limited rack space in existing data centers, customers now need to be selective as to which chip they should put in their data centers. They will only put the most power-efficient, ROI chip that they can find, which are Nvidia chips.
Outside of hyperscalers, any AI startups or research center will tell you they are only buying TPUs or MI300s because they can’t get enough Nvidia chips.
Q: What’s going on with gross margins?
Nvidia’s gross margin in Q1 came in at 78.4%, a 240bp increase from the previous quarter. Next quarter, and for the remaining of the year, gross margins are expected to go back to the ~75% levels, representing a 340bp dilution from last quarter.
They guided this in their Q4 call in 2023, claiming this was due to higher component costs.. The question is, if their products are so good and they have a quasi-monopoly, why are they not able to pass-on costs?
The answer to this, is they are able to but they’re making the voluntary decision to price aggressively to make it a lot tougher for competition and to keep their market share. The only problem with that is that’s affecting their margins, as well as ramping up a new chip on a new technology which has low yields leads to a higher cost per chip and hence lower gross margins.. which is what we’re seeing right now.
The real problem is, they are currently overearning. Their customers are very price sensitive, there’s increased competition from custom silicon & AMD, they have higher component costs (HBM memory in Blackwell costs 3x the amount in Hopper), TSMC is raising their wafer prices…
During the Hopper era there was virtually no competition and H100s would be sold for 2x the price on eBay.. this period temporarily inflated their DC gross margins… and at the same time data center was taking a larger percentage of the company’s revenue mix (which is the margin accretive business) which explains why you would see 400bp gross margin expansions QoQ. Now, mix won’t change margins that much and aggressive pricing/increased competition means every new product will have lower gross margins… based on that I think it’s unlikely the stock will re-rerate and numbers will have to do the heavy lifting for the share price to work.
Demand-side of the equation
All this intense scrutiny on supply is important but what really matters is, how’s the demand environment?
Let’s start with a chart of Nvidia’s main customers:
Starting with the US hyperscalers, here is the latest commentary on their capex outlook:
Capex is expected to continue to grow for these companies, as shown this quarter. We can even see from the Tesla call how difficult it still is for companies to get Nvidia GPUs - and this forces them to put more effort in their in-house Dojo solution.
Comments from the latest earnings calls from Alibaba and Tencent have also been extremely positive for AI/Nvidia.
I just want to point out one part that really captures the demand environment well:
“as soon as we get a server up, that server is essentially instantly running at full capacity” - Alibaba CEO
That’s similar commentary that we’ve also been hearing from the tier 2 cloud service providers (Coreweave, Lambda…).
SMCI also guided for the next 12 months, with revenues $4-$5B above consensus..
To finish off, we can take a look at some AI peers that have already reported.
TSMC raised guidance on stronger smartphone and AI environment, AMD raised its MI300 sales projections for the year and Hynix reported higher capex given the strong demand it has for its HBM product line.
All checks here are very positive for this quarter. Customers are still spending, Nvidia is still supply constrained. As long as demand holds up, I expect Nvidia to continue to do well, even if supply poses a risk to CY24 numbers.
Longer term discussions around AI
I wanted to touch on the more longer-term perspective on AI. It’s easy to get all worked up on Blackwell delays or slower economic growth but it’s often necessary to take a step back and remind ourselves of the big picture.
I think that most jobs will see some kind of AI automation, whether this will work as a productivity enhancement tool or replacement. We’ve already seen this trend start. Earlier this year we’ve heard from Klarna, the private buy now pay later company about its use of OpenAI’s Chatbot . The data on this is very interesting:
In 1 month the AI assistant has had 2.3 million conversations, that’s 66% of all their customer chats.
It’s doing the work of 700 full time call center agents and are on par with human agents in terms of satisfaction score.
It is more accurate in resolving customer inquiries than humans.
This is only a small example but it shows - call center jobs are going extinct. You can train a model on millions of call center conversations and the chatbot will be able to solve your issue 5x faster than a human (average chatbot conversations = 2minutes vs 11 minutes with humans).
Once latency becomes a non-issue and we can talk to models, most office support, legal jobs will also go extinct, or at least be largely automated.
And we’re only in the early innings, what happens when we have AI agents or AI researchers that are able to do research by themselves?
This brings us to one of the most fascinating and at the same time unsettling papers I've come across on AI, which is the "Situational Awareness" by Leopold Aschenbrenner discussing AGI and Superintelligence. He speaks about a world where you would have AI agents in the workplace, that would join your zoom calls, research things online, message/email people, use your apps, code and so on… And that by 2027, you would have Artificial General Intelligence (AGI) which means these AI systems could automate all cognitive jobs or any job that can be done remotely.
Once we reach AGI, AI agents will be able to automate research and learn independently. Imagine having 100 million automated research agents, each operating at 100x the speed of a human. This makes it very likely that we could achieve superintelligence just a year after reaching AGI, meaning AI would possess intelligence that surpasses that of humans.
In order to achieve this, we need infinitely more infrastructure. The following is Leopold’s prediction as to how many H100 equivalent achieving AGI/ASI would require.
100m of AI chips by 2028… versus 4M that Nvidia is currently planning on selling this year. That’s the bull case for Nvidia.
Financials
To the fun part, what does all of this means for Nvidia, in terms of actual numbers?
A large part of Nvidia’s revenues next year will be systems, notably the NVL36/72. ODMs expect to build between 40-60k of these racks in CY25, with prices ranging from $1.5M to $3.7M per rack. The range in ASPs described in the model reflect the mix between MGX (lower content) vs DGX.
This is what I get in a best case/ worst case scenario: ~$143B Data Center revenues for CY25 bear case and $206B bull case.
Looking at different supply chains from upstream/downstream, CoWoS capacity… leads us to these scenarios. Obviously, there is a lot of volatility around these numbers: a year ago, Nvidia was expected to make $50B of revenues in 2025 vs >$150B as of now. I understand that the > $200B in revenues for data centers in 2025 looks elevated, but this is where supply points at, in an environment where Nvidia does not need to be more aggressive with its pricing.
I find comfort in the fact I am not the only one with these kind of expectations:
Mainstream sell-side analysts seem to assume only 10-20% year-over-year growth in Nvidia revenue from CY24 to CY25, maybe $120B-$130B in CY25 (or at least did until very recently). Insane! It’s been pretty obvious for a while that Nvidia is going to do over $200B of revenue in CY25 - Leopold Aschenbrenner in Situational Awareness
Forecast gross margins per product and adjusting for R&D… gives me a worst case scenario of ~$3.4 EPS in CY25 vs best case scenario of ~$5.11
At the ~$125 stock price level:
Bear case, 25x P/E multiple leads to a $87 stock price, which gives us 30% downside.
Best case, 35x P/E multiple leads to a ~$179 stock price, which gives us $45% upside.
The risk/reward looks fairly compelling at these levels. It is less interesting than when Nvidia was under $100 a couple of weeks ago, but I would still be comfortable owning the stock into earnings.











