I joined my friends Brian Schechter and Gaby Lorenzi, investors at Primary, on The Compute 100 podcast. It was a great chat and I wanted to share it with you. You can also watch it on X:
Things we cover:
Why the CUDA moat shrank once LLM inference became the biggest market
Intel Foundry’s culture change, and EMIB as an alternative when CoWoS capacity is short
Advanced packaging, from OSATs to 3D stacking
Prefill and decode disaggregation, and why it becomes a rack and data center design problem
Edge compute for physical AI, including Rivian’s RAP1
AI-native EDA startups vs. Cadence, Synopsys, and Siemens
Rapid fire on TSMC, Anthropic, OpenAI, Hock Tan, and quantum
Also, Primary is hiring:
This podcast is lightly edited for clarity.
Introducing Austin
Gaby: Today’s guest is Austin Lyons. He is the writer behind Chipstrat. When I was getting up to speed on the compute market, one of the first places I looked for information was Substacks. All these independent analysts were writing incredible pieces at the intersection of technology and business, and Chipstrat was one of the earliest ones that I found. So getting to talk with him and hear his journey, not only building the media engine around Chipstrat, but also being in conversation with some of the most senior leaders in compute, has been really fun. What did you think?
Brian: One of the things about Austin, one, he’s just a really nice guy, but he has this set of experiences, from his MBA to startups to working in industry, that has now given him some credibility to predict what are the most important themes coming, whether it’s emerging ASICs or the reemergence of Intel as a really exciting force in the semiconductor industry. Just an exciting conversation to hear what he’s thinking about these days of what’s coming.
The Compute 100 is a podcast about the most consequential market in AI.
Gaby: These are the conversations with builders and investors about the technology and the capital powering the greatest infrastructure buildout ever.
Brian: I’m Brian Schechter.
Gaby: I’m Gaby Lorenzi.
Brian: We’re investors at Primary.
Gaby: Let’s dive in to The Compute 100. This podcast is for informational purposes only and does not constitute investment advice or an offer to buy or sell any securities.
From Intel Engineer to Chipstrat
Gaby: Austin Lyons, thank you so much for joining us on The Compute 100 podcast. We’re super excited for the conversation.
Austin: Awesome. Yes, thank you for having me.
Gaby: You sit at one of probably the most interesting places in the content-meets-semiconductor world over the past five years. I’m sure your career has gone through a complete evolution as this market has caught up to the world that you’re really interested in. I’d love to maybe start 10 years back, even your first career post-grad, and hear your story up to what you’re working on today at Chipstrat.
Austin: Sure, totally. I’m older than I look, so it might be longer than people will think. I got trained as an electrical engineer in undergrad and in grad school. Did my master’s at the University of Illinois. And interestingly, when I was at the University of Illinois, I was actually in the PhD program, but I started a startup with some friends, some other grad students. This was back in the Kickstarter days.
We launched something on Kickstarter, had a lot of fun with that, and it really opened my aperture for what’s possible. So I was like, oh man, I’ve got to go get into industry and maybe get into startups. I actually went straight to Intel out of grad school, down in Austin, Texas, and was working on a startup at night, not in semiconductors. This was more of the IoT, iPhone apps, B2B SaaS early days era.
So I was doing hardware engineering by day, I was doing software engineering by night. And that set me down a path of thinking beyond just engineering, but more about who is the customer? What do they really want? Is this useful? Ultimately, we raised some money, we built a product, and it turned out customers didn’t really want it. So you learn early on the hard way. Fast forward, I ended up eventually moving into product management and getting an MBA.
I was really impacted by those entrepreneurial experiences. So I sit at this nexus of, I understand the deep technology, I can think like a product manager and ask the questions about roadmap, value prop, pricing, that kind of thing. And then I have a bit of an entrepreneurial spirit myself. So of course, I always love to study startups as well.
Gaby: And so all of this was phenomenally well-timed with what is today the compute revolution. Take us through the last five years of deciding to go out on your own with Chipstrat and really build around the ecosystem that is education and technical deep dives around the compute ecosystem.
Austin: Yes. So I was working as a product manager by day and of course the world was changing. ChatGPT was launched and all of a sudden semiconductors, which was my original first love, became super interesting. Early on, of course, it was just from the perspective of logic, but now it’s so much more. It’s logic, memory, optics, power, foundry, all the things. I was spending a lot of time consuming media.
I was training for marathons, so consuming podcasts, reading people’s things, trying to pay attention. And I just saw a gap in the market from an education content perspective. There’s either very technical things or very non-technical, thin, FinTwit-type things. And I just thought, we’re missing a Ben Thompson for the semiconductor industry here.
I felt like my background gave me a particular take on certain things. So I just decided to start writing, because it was as accessible as ever. And then it snowballed from there. You start writing and essentially you start putting your public thoughts out there in the court of public opinion. It’s a little intimidating at first, but it turns out most people are really nice and really interested, and you find the other nerds like you who want to talk about this stuff.
And then, as you guys know how it goes, it just snowballs from there.
Early Takes: AI Accelerator Startups and AMD
Brian: So what is your particular take? I know that there’s a lot in that, but walk us through some of it. Ben Thompson has some core ideas that animate his writing. Curious to hear what are some of your core ideas right now. I’m sure they’re constantly evolving, but as some of them become solidified, what are they?
Austin: Sure. So some of my early takes. Very early on, there’s an AI accelerator startup, Etched, and I started looking at them. Early on, it was just the Nvidia gravy train. Everyone realized, holy cow, GPUs are the future of AI, and Nvidia is obviously dominating in training, and we weren’t even into the inference era yet. And I saw that Etched was out there and I started asking questions like, could there be companies in this space other than Nvidia?
So Etched was one of them that got me into thinking about, what would it take for a startup to compete with Nvidia? Would the model labs trust and buy from a startup? Would the hyperscalers trust and buy from a startup? And so looking at first principles of, what’s wrong with the GPU? And what’s right with the GPU? I started pulling on threads and realizing, hey, GPUs are super flexible.
That’s great for training, that’s great for inference. But at the same time, once you start to fast forward, if transformer-based inference is the defining workload of our time, which has turned out to be true, then it would make a lot of sense for hyperscalers who actually know their workloads to invest in something that might be a little bit less flexible, but as a trade-off, it’s hyper-designed for the workload at hand, for attention and LLMs and KV cache.
So early on I started saying, hey, I think there’s something here. I think these AI accelerator companies can compete. And then similarly, early on, I was actually pretty bullish AMD. I think that was one of my first posts, about a year in, that started to get a lot of traction. Everyone at the time was talking about this CUDA moat.
Part of the CUDA moat early on was the fact that Nvidia had written software so that anyone in almost any domain could accelerate their workload, from scientific computing to oil and gas to quantitative finance and so on and so forth. So anyone who’s trying to compete, say AMD was trying to compete, people would say, oh, well, you have to write libraries in all of these different domains and it’s going to take you hundreds of thousands of man-years.
But once it became that the TAM for LLM inference was going to be way bigger than any of these other high-performance computing things, all of a sudden that CUDA moat kind of shrank. To compete for all these dollars, maybe AMD doesn’t have to even worry about all those other markets. Maybe if they just focus on making their Instinct chip and their ROCm software competitive for LLM inference and start there, they could actually compete.
So those are actually two very similar takes, which is, I think startups can compete in this market and I think that AMD can compete. I’ll just pause there, but it was very focused on logic early on.
Intel: Fix CPUs and Foundry First
Brian: I also enjoyed some of your articulation around Intel and where you see their place in the market right now. Would love to hear you go on that. Maybe we could set the table of Austin’s takes on various dimensions of the compute world and then we’ll double click wherever seems most interesting.
Gaby: I think, Austin, we’ve had probably the most ex-Intel people on this podcast. So we’re happy to have another one and just excited to get your take as someone who was there.
Austin: Yes. Well, I’m wearing my Intel blue today. So, Intel. Early on, once Nvidia started selling all the GPUs and making a ton of money, everyone was saying, where’s Intel? Where’s Intel’s GPU? At the time, I said, hold the phone, Intel is a CPU company. And frankly, at the time, this was back in the Pat Gelsinger days, when they were even trying to figure out, are we going to have an external foundry or not? What is our strategy? And Intel was behind.
Basically, the advantage that Intel always had was they could do their own manufacturing. For many, many years, that gave them a performance benefit compared to everyone else, because they could get to the next node from a manufacturing process faster, and therefore they could reap that and bring more performant, more power-performant products to the market. Again, mostly CPUs, both for client and for server.
Then there was a period of time around 10 nanometer, 7 nanometer, where Intel actually fell behind for various reasons. One being not getting on EUV fast enough. Basically you can think of Intel as manufacturing and as product. They’re an integrated device manufacturer. They do both. They don’t use TSMC. Manufacturing used to really benefit the product team. Once manufacturing fell behind, all of a sudden that really started to drag down the product team and expose, if you have the same technology as everyone else or worse transistor technology, are your designs actually as performant and do they compete? So manufacturing started to fall down and then the product team started to fall down.
Meanwhile, everyone’s starting to say, okay, Intel’s struggling, but oh, they just need to make GPUs to catch up. And my question was, should they make GPUs? Or should they not make GPUs, and not make all these other networking chips and stuff, and actually just focus on fixing both of those things? Make better CPUs like they used to. Of course, AMD was starting to eat their lunch, Arm was starting to eat their lunch. And figure out the manufacturing thing.
Not only get back to being competitive, but then of course the problem with manufacturing. Everyone used to manufacture their own chips, because there’s actually a benefit if you can co-design and design for manufacturing. If you can do that together, not just design the chip but also design the manufacturing, you can optimize up and down the stack. But with every generation of foundry, it becomes more and more expensive, to the point where most companies didn’t have the volume to fully utilize and amortize all of those huge CapEx investments. So people kept falling off, but Intel had enough volume. Even they were starting to get to the limits. So it started to beg the question, if their manufacturing could catch back up and if they could take on external companies, could they get back into the game and compete with TSMC, both on performance but also financially? Could it actually be competitive from a cost perspective?
So early on, my Intel hot take was simply, they shouldn’t focus on GPUs. They should actually just right the ship, make sure their CPUs are performant, and double down on foundry. And now, I will say, that has paid off for Intel. CPUs have exploded with agentic AI. I didn’t necessarily expect that early on, but it has worked out for them in that their CPUs have been in high demand and can help continue to fund the business as they’re starting to bring foundry up.
And of course, within the past six months, we’ve started to hear a lot about Intel Foundry and external customers. So they are on stable ground with CPUs and with manufacturing. It still begs the question, what is their AI strategy? What is their GPU strategy? But I think now Lip-Bu has put enough sturdy ground underneath their business that they have earned the right to expand into GPUs.
Gaby: Was your belief that investing in manufacturing is what will enable Intel to be a better partner to external customers, hence using that spare capacity to deliver for whoever they might be selling to in the next decade?
Austin: Yes. So the question for Intel is, why didn’t people just raise their hand when everyone wanted to make GPUs and TSMC said, we are supply limited, and it depends on where you are in line, and you get allocated so many wafers? And Intel said, hey, we’re going to get into the game. Why didn’t people rush over to Intel?
Obviously, there’s a big cost in saying, now we have to redesign everything for a different foundry. But ultimately it was cultural. When you’re an IDM, an integrated device manufacturer, your foundry ultimately serves your product team, and you do have this internal groupthink of, this is the way we do things and this is the best. That is not a customer-centric mindset. Your customer is the product team, but you’re all the same company. So when Intel said that they were going to become an external foundry and they started working even with small customers, it did give them an opportunity to start trying to build that muscle.
And probably the best thing that has happened at Intel is they’ve brought in external hires who’ve worked at other foundries, whether it was GlobalFoundries or even Micron or Samsung, memory companies, who already have that correct orientation of, I work in a service business. Because that’s what manufacturing is.
So yes, I think Intel has had to undergo a lot of cultural change, but they seem to have worked through this. And ultimately it will end up benefiting Intel products, because the Intel manufacturing team now has to make the decision, who should get the capacity? Who should get the wafers? Should it be our internal products team or should it be our external customers? And that’s actually going to be great for both of them, because now the products team can’t be sloppy and just be like, ah, we screwed up here, do a hot run and run a couple more wafers for us. It’s like, no, no, no, you have to behave like an external customer. You need to plan in advance, and there might be times where you have to go to the end of the line and wait your turn.
Gaby: Yeah, it’s interesting. It’s almost similar to what’s playing out in some of the hyperscalers that are designing their own chips now, and the way that they have to think about resource allocation between their own internal teams using their systems for the development of AI versus selling those chips, which are ultimately becoming really valuable product lines. So it is going to be interesting to see how that ultimately affects Intel and what their business looks like in the future. Very, very curious to see.
Austin: Yes, totally. It’s a brand new world. And of course, again, none of us could have predicted all the geopolitical stuff that’s going on, which also indirectly benefits Intel Foundry as customers are incentivized to say, I’ll give you a chance.
Advanced Packaging
Gaby: Totally. So one area that you spoke to that Intel has done a lot of really great innovation in is packaging. Something that we’ve been tracking from a startup perspective is how much value is going to be accruing to the fabs themselves. As we start to see systems like power become more integrated into the package, cooling becoming more integrated, so much value is going to be going to packaging and who’s doing that. So I’d love to hear a little bit more about what you saw and those learnings from your deep dives into the Intel space, and then do a little bit of a segue into what you’re tracking in packaging these days.
Austin: So packaging used to be a fairly boring space, low value, low margin. If you go way back in time, packaging is literally taking discrete transistors and wiring them together. And then it was taking discrete chips on a board, printed circuit boards, and just big copper traces in between them.
But nowadays, we’ve moved into advanced packaging, where we’re taking chiplets or dies and trying to connect them in three dimensions. And that’s way more sophisticated from the perspective of how do you route everything electrically. It’s complicated from both a design and a manufacturing perspective. From a design perspective, you’re thinking about crosstalk and noise and electrically routing all these signals to the right places.
Brian: Austin, maybe if I could break it down even more simply and just check for understanding, because the 3D thing was confusing for me for a while. One way that I think about advanced packaging right now is, we used to have city streets where every building had at least alleys between them, if not roads.
And we now have begun building the buildings to be touching each other and have shared plumbing and electrical connectivity between the two of them underneath. And so that is now the 3D element, where it’s no longer just everything laid out on a board assuming that it’s its own piece. Is that what we mean by 3D?
Austin: Yeah, that’s a good analogy. Historically it was just, you had a chip and that was it. You’ve got a house and it’s in the country, whatever. We’re not even dealing with roads. And then eventually there’s this 2.5D, where it was, yes, let’s start to put buildings near each other, and now we have to put roads and alleys and connections between them.
And then as we’re headed into 3D, for example, high bandwidth memory is 3D-stacked DRAM chips. You’re literally stacking them up and you have to think about the internal plumbing. If you’re talking about a high-rise building, how do I get heating and cooling to all of these, get the waste out, electricity?
And we will continue. In the future, there are actually companies who are working on stacking logic chips. You could either have logic chips with memory on top or logic chips stacked on logic chips. So yes, now we do have to think about the infrastructure in three dimensions.
It used to be OSATs that did this packaging, and that stands for outsourced semiconductor assembly and test. It says outsourced. Right there you can tell it’s low value. But now it’s obviously getting higher and higher value and getting more and more complicated. At the end of the day, it’s very similar to what foundries do when they’re fabricating logic chips or memory chips. It’s that level of deposition and etch and patterning.
Now the OSATs are trying to move into that space, because they want to accrue the value at the OSAT layer and they want to say, no, no, we can do CoWoS now, so if you can’t get it from TSMC, we can do that too. But of course it’s also an opportunity for TSMC and Intel to expand from just making logic chips to doing more of the full-on logic, dicing, packaging, everything.
Gaby: And so when you think about it, maybe selfishly from a startup perspective, thinking about where there’s new technologies that are interesting to invest in, as so much more value is going to be accruing to people who are doing packaging or TSMC as a fab, how do you think about the startup landscape evolving? Is it largely going to be startups that are building around new processes or startups that are enabling new types of packaging techniques?
We’ve seen companies that want to be a next-gen Amkor or a next-gen OSAT more broadly. But would love to just hear how you think this affects the newcomers and how they build something that’s really valuable.
Austin: Yeah, this is tough. One of the hard parts when you first look at the semiconductor industry is it feels like everywhere in the value chain, there’s two or three or four companies and they’re huge. So you have to ask, okay, if I am a PhD student and I’m studying some really interesting new material or something, how do I take on Applied Materials or someone like that?
I don’t have scale, but what could I do? Obviously, one thing that you’re thinking about right away is, could I build some interesting technology that they would want to acquire? But if you want to be a full-on company, let’s talk advanced packaging. As this is getting smaller and smaller and the power consumed by these dies continues to increase, there’s really interesting questions around how do you deliver power there?
Power used to come in from way off to the side. Are there ways to integrate it much closer on the package, or even as a die, as part of this 3D stack? So you can think about, are there power delivery opportunities? Are there power dissipation opportunities? Interestingly, I think what probably doesn’t get appreciated enough is from an EDA or a design tool perspective.
There used to be people that would design and think about the electrical characteristics, and maybe different mechanical engineer type people thinking about the thermal characteristics. And now we live in a world where you have to always be doing both and thinking about both, because it’s just such small places with such hot, hot dies that are interconnected in three dimensions.
So I think there’s opportunities to ask, are there ways that we could help solve this or create a new user experience from a design perspective, where the electrical design and the thermal design is taken into consideration? There’s definitely opportunities. It’s just about that seam of how do you take your innovation and bring it to market and get in here.
EUV, EMIB, and CoWoS
Brian: Can we actually spend another minute on that? Austin, take us further into that mistake of not investing in EUV, and then also the technical bets that they did make that people are now starting to say look like they’re starting to be really promising.
Austin: Both of those, because both are important. From the logic perspective, one thing that people don’t appreciate about Intel is they actually do a lot of the transistor research and development that enables future generations that the whole industry uses, TSMC included.
Intel’s always been really good at that early R&D. They have innovations from the early days like strained silicon and using different materials as oxides and things like this. So they’ve always really pushed forward on stuff. But there was a period where they took some bets on some materials that didn’t work, and they ended up delaying jumping to EUV. I can’t remember off the top of my head exactly why, if it was a cost concern.
Brian: I think it was a cost concern. I think the cost was too big.
Austin: That actually aligns with the idea that they had some more financial-oriented folks who were driving the ship during those times and not engineering-centric folks. So Intel, yeah, they slowed down there. Basically they were stuck at a particular node, 10, 7, while TSMC was able to catch up and keep going. And then of course everyone, AMD for example, is using TSMC and they’re getting on smaller and smaller transistors, and Intel was essentially stuck.
At the same time, Intel has always been innovating on packaging, and EMIB is one of their packaging processes that they use. The idea is, once you start to make chips that are really big and you hit the reticle limit, you might want to make two of them and tie them together.
TSMC, their approach is CoWoS, chip on wafer on substrate. And right there you can hear it. There’s three levels there, chip, wafer, substrate. So there’s this interposer where the connections happen. Intel has historically had EMIB, where basically just in the substrate itself, they embed these little bridges. So they just put the connections in between the chips right where you need them. The rest of it is a substrate.
And it hasn’t gotten much attention at all until very recently. But one of the benefits of being an integrated device manufacturer is, when your technology works and you have the scale of Intel’s product team making CPUs for tons of laptops and every company worldwide, you have the opportunity to take your packaging and put it in production and run it at scale. So now Intel Foundry is talking to potential customers who either just can’t get enough CoWoS capacity, ultimately they might be gated by the amount of wafers they can produce at TSMC, or, there’s different flavors of CoWoS, but some of the early flavors that were adopted have particular trade-offs.
And even as TSMC has come up with something called CoWoS-L, which is local silicon interconnect, and it’s fairly similar to EMIB, there’s still the three layers versus two. But there are some scaling problems with the way that CoWoS is implemented. And Intel Foundry can essentially stand on the sideline and say, hey guys, we’ve got this thing called EMIB, it scales really well. We are already making it in volume.
We’re servicing all of our Intel products team and all of their products that they’re shipping. So we know that it works and we have spare capacity. So yes, as we’ve moved to these AI accelerators that want to just be as big as freaking possible and have all this HBM that needs to be interconnected, it’s the perfect timing for EMIB and Intel Foundry.
Information Sources and the Master Plan
Brian: Where are you getting your information from? We listen to various thought leaders in the compute space and they all have their information edge. Where do you get your information? And also, to the extent that you’re open to it, what are your goals? What’s the Austin Lyons master plan?
Austin: Yeah, these are awesome questions. So where do I get my information? I of course try to consume all the public information that I can, earnings reports, whatever. When I can get my hands on expert interviews, those are interesting. You always have to take those with a grain of salt. And then I have the opportunity in my day job with Creative Strategies as an analyst to talk to a lot of these merchant silicon companies and talk with their executives.
Sometimes I’ll bring them on and publicly interview them, or other times I’ll get to talk to them behind closed doors and just ask them what’s top of mind. So that is an awesome privilege. But I also am always trying to talk to startups, see what people are saying on X. Again, a lot of it is public conversation. If there’s one thing I could do more of, it’s continue to try to talk to people who are in the trenches as much as possible.
Going to OCP or conferences is a really good place to meet people and have the opportunity to just ask them. I try to balance, okay, what is the market narrative around something? What is the FinTwit narrative? What are the thought leaders saying? What is Gavin Baker saying? But what am I hearing from companies? What are they saying?
And then thinking about questions where it’s like, well, the company said this, but did they actually mean this? Or they specifically didn’t talk about this. So trying to find those gaps and noodle on some of those. And then I will say, one of my advantages, I think, is that I don’t live in Silicon Valley. I’m outside of the thought bubble of, everyone’s thinking this or saying this. I live in the Midwest, and so we have a particular perspective on how the world works. And as you guys are experienced with as well, I’ve worked in industries outside of this.
That lived experience also gives you an ability to think critically and differently about certain things. And then to Brian’s question of the master plan. When I was a product manager, what was awesome was, ah, I like to think about strategy, I like to think about technology, this is so fun. I get to make a plan. What’s our one-year plan? What’s our three-year plan?
And then you think about it for a little bit and then you just spend the rest of your life implementing it. So it’s a little bit of strategy and a lot of execution. What I like about my current role, working as an analyst, writing Chipstrat, and then I also host a podcast called Semi Doped with Vik Sekar, is I just like to be on the forefront as new information comes out and think strategically. Oh, what does this mean?
China’s going to shut down access to indium phosphide. Oh no, is that real? What does it mean? So a lot of just thinking strategically. And then fortunately I’ve been able to build up relationships with executives and people at these companies to start to ask, I think this means this for your company. What is your answer to that?
Brian: Yeah, it’s really interesting, because what it points to, for the hardware designer, whether it’s a startup or an incumbent or a hyperscaler, is having a point of view around the TAM associated with different types of value that AI can create. Because ultimately the TAM needs to be there, and the margin profile will need to be there as well.
Disaggregated Compute
Gaby: I’d love to spend more time on disaggregated compute. One of the areas that people are spending time thinking about is, in response to disaggregation, a lot of new areas emerge, or maybe not new, but the downfalls of disaggregation emerge, like networking or how close your power supply is. So walk through your framework. If we move to this disaggregated world where a single vendor has all of these different compute offerings for different parts of workloads, how do you actually bring that together in a way that you can still enable efficient serving of inference, or maybe it ends up being for training, but largely for inference, when you have to deal more with networking, power, memory, et cetera?
Austin: Yes. So when it comes to disaggregation, we’ll start with prefill-decode disaggregation, because everyone knows that. It actually makes a lot of sense. There’s these two subtasks and they have two different memory constraints or profiles, how much memory capacity and memory bandwidth they need. Prefill doesn’t need much and decode needs a lot.
So right there you’re like, wait a minute, we could have two different hardware SKUs, and then we could run part of the workload on one and part of the workload on the other. But then to your point, Gaby, now that means you do add some complexity. You need to orchestrate across them, you need to send information across them. So right away that starts to point to, we’re not talking about chip design, we’re talking about rack design and ultimately data center design. You start to ask questions about, what is the most power-efficient, most cost-effective, fastest way to send information back and forth?
You could go so far as to ask, how should we even architect the data center? Do you have racks of prefill chips, racks of decode chips, or do you have a rack that has prefill and decode in it so they can talk back and forth? And that even goes to co-design with the model companies. How are they taking their mixture of experts and splitting it across? We’ve got these layers on this GPU and these layers on that GPU, this expert on this GPU, this expert on that GPU.
I do think that’s why it pushes Nvidia to just say, oh, I want to try to get 576 GPUs in the same rack, even if the power’s crazy, if the cooling’s crazy, if the manufacturing’s crazy. You’re essentially trying to keep as much of the compute as close as possible to make it easier to send everything over copper if you can. But it does start to become a problem, and it opens up opportunities again to rethink. Oh yeah, but what about long context size and KV caching, and a cache miss versus a cache hit?
Should we get clever and interesting with having a rack right next door that’s just a ton of storage, so you can store the KV caches as close as possible? So we’re obviously disaggregating things and saying, how can we have the proper SKU for the proper workload need? But we’re also realizing, oh crap, there’s trade-offs, and how can we also aggregate it as much as possible and get them as close as possible?
Time will tell where other interesting innovations happen, from networking to memory to cooling, to continue to let us get more done faster but not waste as much power sending bits all over the place.
Physical AI and Edge Compute
Brian: As we’re talking about this, a lot of the reason for all of this comes back to the size of the models and the size of the context windows that they benefit from in terms of developing more and more useful forms of intelligence. This is a change of topic some, but what are you most excited about in terms of AI and what capabilities are coming online? And what are you most hungry to see the semiconductor industry do to help enable those types of possibilities?
Austin: I’m most excited about physical AI, about embodied AI. If you think about it, yes, in the last 30 years, an increasing amount of our world’s GDP is impacted by digital things, Netflix and Facebook and the internet and Riverside.fm and what have you. But so much of the world is still physical. Everything that you buy, everything that’s manufactured, Apple, Nvidia, it’s all supply chains, it’s all physical hardware, it’s all getting moved, it’s all getting shipped.
Of course, I live in places where people are growing corn and soybeans and hogs and things. Anyone who eats anything, that had to get physically moved. And most of that stuff is still not touched by AI. It’s barely touched by digital. There’s some logistics and things that go on. But literally, I was at a large grocery retailer and we were using mainframes still, and this was not a decade ago.
And they will continue to. There’s places where people are still using pen and paper. So it’s crazy to think, obviously just getting digital is one thing, but man, how can physical AI just unlock even more productivity? And it’s not just making what you’re doing more productive, like saying, here’s the way we work today, how can we be more productive? One of the cool things with anything digital that’s AI-enabled is it enables some solopreneur to do so much more than they ever could before.
Oh, hey, I’m a domain expert. I’m a dentist and I’ve always thought our software sucks, but now with AI, I can just write better dental software, because I’m the domain expert and I can just do it, send Fable off for a week. My oldest is in middle school and he’s been spending this summer coding an indie game and it’s super cool. He knows how to code but not that good, but Fable can just go off and do it.
And then he can spend a lot of time thinking about the story and the graphics and the art and stuff like this. Now imagine if we could take that and put it in the physical world. What if I could realize that I hate this particular part of my car? Of course I could 3D print it and stuff, but what if I could actually manufacture and assemble it and be like, hey, I’m going to make this bespoke add-on piece to cars and start selling it Etsy-style or something?
But it’s actually something physical. In Iowa where I live, lots of small towns are supported by a local manufacturer. It might be snowmobiles or basketball hoops or whatever. But they suffer from brain drain like everyone else does, which is everyone going off to the big city, and they don’t want to come back and live there and work on it. What if you could use robotics and physical AI innovations somehow to make it so the same amount of people could get more physical work done, or they could graduate to managing fleets of humanoids instead of being the person doing the thing?
It’s obviously very fuzzy, but what I’m excited about is, man, it feels like there’s so much human productivity that’s not unlocked yet, and there’s so many interesting, capable people that have very interesting mechanical or physical ideas, but we just don’t have the tools for them to just build.
Brian: And so what hardware, compute-specific unlocks do you think will lead to breakthroughs in physical AI? Or do you think that’s part of the equation?
Austin: Yeah, good question. Historically, you’ve just got embedded compute, and it’s little CPUs and not much memory and it’s all pretty wimpy. Nvidia has had Orin and Thor. They have a small portfolio, and it has certain power requirements, and it fits perfectly for certain use cases. But if you go talk to people who are doing AI at the edge, you’ll find out there are customers who are like, oh yeah, we just have to take three Nvidia SOMs and bolt them together so that we have enough compute to do what we want to do.
But that requires a ton of power. There’s all these trade-offs today. So I just don’t think we’re yet to a point where we’ve got enough different compute options for the edge, whether it’s more FLOPs or more memory or less power or just more different configurations. We are in the early days and it is a little bit of, you get what you get. Think about it. With a humanoid, it’s got a brain, it’s got sensing, it’s got all these little sensors at the fingertips, and you could have distributed computing within it.
I don’t think we even know what is the right hardware architecture there. There’s CPUs, there’s GPUs, there’s FPGAs. But are we designing the systems we build today based on the hardware that’s given to us, or are we building the right hardware for where these workloads are going?
Gaby: We talk a good bit about the hardware lottery, and this idea that what is working and the research ideas that win are just what is suitable to what exists at the time. If you think about it, cloud was enabled by CPUs, AI enabled by GPUs, and there’s likely another iteration of this going to happen in the physical AI world, a little bit more than just what a Jetson offers, if we use that as the best edge compute that exists today. But it almost makes it hard to think about what are the use cases that a next-gen Jetson, or the next best piece of edge compute, could enable.
We’ve spent most of our time thinking about how edge will impact personal computers, or the way that we’re running more models locally, and how that eventually will trickle down into the economics of the AI economy. But of course the physical AI world is going to be, hopefully, an explosive customer of edge compute. In terms of the conversations you’re having with people around this market, where are the real pull and pain points of what exists today in the physical AI space? We hear a lot about power, a lot about weight, about latency. But I’d be curious, in the conversations that you’re in, where you feel like physical AI’s compute is not stacking up today.
Austin: I think what Rivian is doing with their RAP1, the Rivian Autonomy Processor, is pretty interesting, because they are designing their own platform. They said what they needed off the shelf wasn’t there. And it was very specific. For personal vehicle autonomy, we need a real-time operating system and some real-time safety that has hard latency budgets and everything. But we also have cameras, we have lidar, so all this sensor data input, lots of data, it’s moving around. They actually created their own communication protocol. They called it RivLink, which is like NVLink or something.
They made a lot of trade-offs for their particular portfolio of vehicles, and I think that they’re also extending it to Mind Robotics potentially, which is, hey, can we also have humanoids that help make our vehicles someday? I thought it was just an interesting example of needing certain amounts of compute, a certain amount of interconnect bandwidth to send all this information around, and yet still having particular power constraints. Now, of course, at the end of the day, they’re connected to a big battery in your car. They don’t want to drain your battery, but they have more power.
Whether it’s humanoids or, what’s interesting is what Meta’s doing with those Orion glasses. Those were going to be crazy expensive, of course for some of the optics reasons, but still at the end of the day, it’s battery, compute. And if you think about it, most of the edge compute that we’re using today was probably designed two or three or four or five years ago. What was the AI that they were designing for at the time? It was convolutional neural networks and more traditional vision AI and ML.
So I actually don’t know if anyone has designed edge compute around end-to-end transformers yet, other than, again, maybe Rivian, maybe Tesla. That would probably be a big opportunity. If you started from first principles and you said, I want to run a transformer end-to-end at the edge, what is the right compute to support that?
Where Startups Can Get a Foothold
Brian: What are you most excited about from a startup or early-stage company creation standpoint?
Austin: I always love hearing about the technology, but what I’m most interested in is, where can a startup actually gain a foothold against a very large incumbent? In networking, what always seems to happen over the last 20 years is a startup, it’s usually people with a lot of experience, spin out of a company and then they go start something and then they get acquired by Cisco or Marvell or Broadcom. And it’s amazing, it works. If you want to start a startup and get acquired, go to networking. But it’s a little bit less interesting, because you’re like, oh, I know the playbook here.
What’s more interesting is the people who are playing in the space of ASML or foundry, like we talked about, where everyone on the internet’s just like, oh, you can never do that. That’s the dumbest thing I’ve ever heard of. How could you ever compete with ASML? There’s some startups, Substrate, xLight, where they’re saying, hey, what if we use X-ray lithography, or what if we use electron beams? What if we just think really crazy from first principles? What could that unlock?
I’m excited, from a strategy and entrepreneur perspective, to see people with chutzpah who say, no, no, no, we actually think there’s a way to pull this off. And I just like to watch it. One, I want to understand their crazy. Tell me more. You think this is possible. Why is it possible? I want to learn from you. And then two, I just want to see how it plays out.
Brian: One of my passions as a VC is thinking about people and what they want to do and where they’re going, and then the questions associated, like, how do you get there? One thing that leads me to wonder about for you is, what is the question that you think people are asking that maybe they don’t even know how to fully frame, that comes up again and again, maybe in the context of consulting or maybe in the context of people who are engaging with your content that are hungry to come to you? What are people really trying to find out?
Gaby: Can I give a guess and then you can tell us if it’s right? My guess is that people want to know how to trade in the stock market and are trying to get some intel.
Austin: Yeah, totally. Ultimately that’s what it boils down to, and I was trying to abstract it in my head. The simplest thing people ask is essentially, is there an interesting insight that you have that I could profit off of that is not obvious to other people? If I abstract that, even when I work with tech companies, I think what the executives want to know is, what am I not taking seriously enough that I should be? It could be a new idea.
One of the things that I put out into the world before a lot of people saw it coming was, hey, I do think CPU demand is going to shoot up. I didn’t see it as early as I should have, but once you start to think it through, you’re like, oh, wait, I’m just going to have a bunch of interns doing my bidding. What are they going to do? They’re going to use Microsoft Excel or whatever, Google Drive, whatever. And then I had tech executives say, oh, we read your piece and we hadn’t really thought this through, but this could impact our supply chain decisions if we’re going to need more CPUs than we realized.
So abstracted a little bit, it’s, because I’m afforded the time and I have the experience, how can I traverse the frontier of knowledge and just think through interesting things and find one that has an implication that’s going to really matter? Could matter for institutional investors, could matter for tech executives.
And actually, I was in a PhD program. I just stopped after my master’s. But when I was doing research, it was actually very, very similar. I was doing research in graphene and carbon nanotubes, and what you would do is you would survey the landscape. I’d work with my professor and he’d be like, go read all these papers as a new student and then see what people are doing, and what haven’t they done that should be done?
It’s that same muscle. What is Nvidia doing? What is AMD doing? What are the problems that startups are working on? Because that can tip you off into problems that the big companies aren’t solving. And then, what’s an interesting space that people just haven’t really thought through much? Let me quick go think about it. Are there any interesting implications? And then I could publish it or I could sell it to clients or what have you.
That’s what also gets me up, because you don’t know what that’s going to be tomorrow. You’re like, well, I did this yesterday and I came up with something interesting, but what about next week? I should go put pen to paper and come up with, what are my theories, and try to name them and come back to them.
But for me, it does usually end up at, okay, this is the way the industry works. There’s various incentives for two or three players to exist. Yet there’s always technological revolutions, even on a cadence. If you look at optics and networking, it’s 400G, then it’s 800G, then it’s 1.6T. So there’s the status quo, and then there’s inflection points that are going to happen.
So there’s questions of, can incumbents take advantage of those inflection points, or could startups come in and take advantage? I guess some of my rules of thumb are, just because you’re the incumbent doesn’t mean you will continue to be the incumbent. And just because you’re a startup doesn’t mean you can’t compete here. Usually you just have to be clever about routes to market and whatnot.
Even with transistors, there’s a company, imec, that’s over in Belgium, and they research this stuff and they publish a roadmap out to 2040. So it’s not crazy for a grad student to just be like, well, let me go think about some of that 2040 stuff. Is there a company to build there?
AI-Native Chip Design
Gaby: It’s fascinating, because as they’re publishing roadmaps till 2040, we’re also living in this world of just accelerated, compressed development times, research periods, all of this. What Dario calls the compressed 21st century, where all of research, whether it be across pharma, semis, whatever, is just going to be compressed. I imagine you’re an avid user, and it sounds like your son is an avid user of AI. I’m curious how you’re seeing the strongest teams actually applying AI to this industry.
We’re investors in a company that’s doing AI for materials science with a focus on compute. We see that as a really valuable way to be identifying areas where there’s new opportunities for materials to drive innovation. But obviously chip design has been a big area, verification, manufacturing. Would love to hear about where you’re actually seeing AI applied to this industry from a software perspective.
Austin: I think this is going to be the Harvard case studies of our time, when we look at this period where you’ve got companies that weren’t built AI-native, and then there’s going to be startups that are built AI-native. So even AI-native chip design, let’s talk it through. Right now there’s three big EDA companies, Cadence, Synopsys, Siemens.
They have all the customers. They have been qualified with TSMC with their PDKs, process design kits. So there’s all of this momentum, institutional knowledge, and everything that’s built up. On the other hand, you’ve got startups who have the freedom to think from first principles and build AI-native chip design. So the question is going to be, who’s going to win?
Is it the big guys, who have actually solved a lot of really hard non-technical problems, but have to overcome their old way of thinking and figure out, do they bolt AI into EDA? Or the startups, who are going to build EDA for AI from the start, but then they have to go convince TSMC to work with them and they have to go convince customers to use them?
I think that’s what’s going to be fascinating to watch play out. As far as companies using AI, I’m surprised when I go talk to even some of my friends in Iowa. They’ll be working at power companies, traditional power companies. And even they are saying, Austin, all we talk about is AI, and are we doing enough and are we thinking enough? And hey, we’re doing on-premises. We’re buying huge racks of GPUs.
It’s totally new to us, but we have to either sink or swim, and we’ve got to learn how to swim quickly.
Gaby: Yeah, it’s interesting. On the AI EDA side, we saw, not Chipstrat, we saw ChipStack get acquired by Cadence. And I think a lot of people are interested to see how Synopsys is going to evolve. But like you said, a lot of these duopolies, or areas where there’s two to three major players, whether it be in EDA or other areas, are thinking about, how can we integrate AI into our stack?
And a lot of the question, like you mentioned, is going to be, do we bolt it on or does someone new come in entirely? Part of the question is going to be how relationship-driven and vendor-connected these industries are. Beyond EDA, and the power thing you mentioned around everybody thinking about AI at the forefront, are there other areas in the semiconductor world or compute landscape more broadly where you feel like AI hasn’t been applied enough? One area I’ve been learning a little bit more about has been the defect detection world.
If you think about defect detection, maybe a company like KLA or Lam Research is building technology that can use computer vision and really powerful AI models to detect whether a chip has been produced correctly. That’s one area I’m spending time in and learning more about, but I’d be curious if there’s areas you feel like are underserved.
Austin: Interesting. Yeah, that one’s cool. AMD just bought a company, I think they’re called MEXT, and they were doing something really cool, which is applying AI, it is more traditional AI, but to memory. Basically trying to predict cache hits and cache misses and trying to predict access patterns. I thought that was a pretty cool use case. Even something as simple as what you study in computer science 201 or something, when do you write something to a cache, and when there’s a cache miss, there’s a write-through, some of this very simple stuff.
It’s probably actually still some very simple heuristics or algorithms. What if you applied AI to that? So I do think there’s probably an opportunity to literally go up and down the stack, at the operating system level, down to how we access memory, and there’s probably a lot at the networking and data sending layer. Could you predict how you send information even more intelligently than we do?
I think there’s a lot of interesting opportunity to explore where to apply AI next.
What’s Next at Chipstrat
Gaby: Perfect. So Austin, tell us about what’s on the docket. What are you working on? What’s coming out soon in the Chipstrat world?
Austin: What I’ve been starting to do lately is zoom out from individual companies and look at the markets as a whole, because there’s a lot of new dynamics happening because of the huge demand-supply imbalance. So memory, zooming out and saying, oh, there’s all these LTAs, these long-term agreements. How should we think about the dynamics of memory markets, which were traditionally more of a commodity, when they have an LTA in place? They’ve got agreed commitments from people to buy capacity for five years.
How does that impact pricing? How does it impact whether they should spin up more capacity or not? There’s a lot of interesting game theory things where it’s like, oh, the rules just changed a little bit here. Trying to look at that across optics and networking. And then of course, I think foundry is super interesting again. For all intents and purposes, TSMC was an earned monopolist, which means if they wanted to, they could charge whatever prices they want, but they hadn’t historically. Then they started leaning into selling their value.
But now all of a sudden, as Intel Foundry and Samsung Foundry come into play, what does that mean for how quickly they raise their prices, and also how quickly they expand capacity or don’t? Because expanding capacity is super expensive and risky for someone like TSMC. Again, there’s these dynamics I’m trying to dive into, which is, how are they calculating how much they could spend and how much capacity they are leaving on the table?
And this is also so interesting because if you don’t build it today, that impacts you in four years. And if you don’t build it today, you know that it’s going to go to Intel or Samsung potentially. So are you lifting them up? Or is TSMC believing, no, we don’t think they’ll be able to fulfill it? There’s a lot of gamesmanship. It’s fun to look across memory, look across foundry, look across all these markets and figure out, how are people thinking about this world where demand is so crazy?
And then we still don’t know. Companies are ramping up capacity for 2028, 2029. And even just predicting, is the rate of growth going to slow or is it going to stay? Because of course we see free cash flows going towards zero and there’s all these financing questions. So I’m just trying to also project a couple years and think three, four years down the road.
Gaby: Awesome. Very excited to be reading whatever market you’re digging into next. Brian, what are your rapid fire questions for today?
Rapid Fire
Brian: So this is a new experiment for us on The Compute 100 podcast. Gaby, we’ll ping pong. And Austin, you just need to answer quickly. All right. You have to put your entire net worth into one company in the semiconductor industry. Which one would it be?
Austin: Oh man. TSMC.
Gaby: What pending IPO are you most excited for?
Austin: Anthropic.
Brian: If one marquee name in the AI revolution, public or private, were to take a true nosedive within the next six months, which would you guess it would be and why?
Austin: Off the cuff, my brain goes straight to OpenAI. They’re awesome, but it’s just their path dependency of how they got to where they were, essentially at the core being researchers and then thinking, oh yeah, we will go to consumer and we’ll make a bunch of money. And then it’s like, dude, people only want to pay $20 a month. Meanwhile, Anthropic’s like, I know where it’s at. It’s charging enterprises hundreds of thousands of dollars for tokens.
Now, of course, to Sam Altman’s credit and the OpenAI team, they have done great with Codex and they’re chasing that, so I think actually they’re fine. But if they fell off, it could be them. Obviously there’s distractions around hardware and lots of other vectors it could be tempting to go into. So if they’re trying to prevent a nosedive, I would just say, dude, focus on enterprises and big customers.
Gaby: Who is your favorite CEO on earnings calls?
Austin: Favorite CEO on earnings calls. Oh, Hock Tan. Hock Tan from Broadcom. I love him because he is very wise, very experienced, very old, also has a lot of power in the industry, and sometimes he doesn’t really give a you-know-what. And I love how that comes through.
Brian: Shoot, I had one for you. That was a great one, Gaby. Who would you like to hear us interview on a podcast who is not someone you typically hear?
Austin: Interview Richard Ho from OpenAI. He’s their head of hardware. Of course, they’re building their own hardware, which we’ve heard about recently with the Jalapeño chip and Broadcom stuff. But they’re also building the AI that helps them to build the hardware. So what a cool perspective that he has, building the AI, building the hardware. Amazing.
Gaby: If you weren’t doing this, what would you be doing?
Austin: Oh man. If I wasn’t doing this, I would be trying to create the best lifestyle I could, and using my brain. How I make career decisions is just, am I going to get to use my brain more or less? Of course, you’re balancing the trade-offs there too. When I was in grad school, you have total intellectual freedom and you also don’t make any money. So at some point you’re like, well, I have to feed kids.
Gaby: The thing that is exciting to your brain is also the biggest market ever, so you’ve coincided in a good way.
Brian: All right, let’s keep going. Again, I had one, but this is fun. Gaby, your questions are good and I want to get Austin’s answers.
Gaby: What is your favorite book in the semiconductor world?
Austin: Focus: The ASML Way is really good. It’s about ASML. I thought that one was cool. I thought Tae Kim’s book on Nvidia was good too when it came out. And of course everyone’s going to say Chip War, Chris Miller. That’s a great one too. There’s actually lots of good ones out there.
Brian: All right, maybe one more. What’s a question for us that you’d be curious to hear about before we wrap up?
Austin: What are you guys most excited about, looking three or four years out, not three months out, into the future about compute?
Brian: That reminds me of one of my questions for you, which was actually, quantum relevant within the next five years, yes or no?
Austin: Quantum companies, please convince me yes, but I say no.
Brian: Cool, great. Most excited. Gaby and I spend a lot of time thinking about the rising generation of entrepreneurs and operators who are AI-native and want to build the future of compute for AI. Watching them do seemingly impossible things, and transforming the industry along the lines of what we’ve been talking about, is a thing of constant fascination.
Austin: Yes, I love that. It’s like, dude, AI is magic. It’s like Harry Potter before he went to Hogwarts and after he went to Hogwarts. Think about these kids that just have this magic wand. It’s amazing.
Gaby: It’s crazy. It is a different world spending time on college campuses. I feel like they are learning in their freshman year what I learned in all of college. And that’s not actually a curriculum thing. That is just the speed at which they can keep up with the world around them and get exposure to things, and a lot of that is obviously because of AI. It’s just so fascinating. So the new generation of talent here is just really exciting.
Austin: Love it. That’s exciting.
Gaby: Amazing. Well, Austin, thank you so much for joining us. We really appreciate the conversation, and thank you for coming on the podcast.






