This is an odd time to be starting a publication that covers the personal computing market, to say the least. Massive AI data-center build-outs are sucking up all of the chips, and things look bleak in the consumer space as a result. Prices for memory and storage have skyrocketed. PC makers are raising system prices, pulling back on system specs, and asking for ridiculous things like 4 GB graphics card configs. Nvidia appears to have torn up its roadmap for GeForce products to focus on those sweet, sweet high-margin data-center GPUs. And folks are talking about this crisis continuing for several years—unless the bubble bursts on data-center construction.
It’s easy to look at the situation today and be glum. But I have a few reasons to be optimistic in spite of it all.
For one thing, the present moment isn’t entirely unprecedented. We’ve had two prior waves of demand—driven by cryptocurrency mining—hoovering up all of the high-end consumer GPUs, and we got through that. I think Jensen Huang is right to cite a sort of Jevons paradox situation for GPUs: as the cost per flop of data-parallel computing power decreases, demand rises. That dynamic has held true essentially since the birth of the GPU, and it’s a great thing for the computing market overall. The other side of that coin is that supply will rise to meet demand, made possible by the rising prices that fund the build-out of manufacturing capacity—eventually. Also, the bottom could drop out of the AI data center bubble, and we could find ourselves laughing maniacally while bathing in stacks of cheap DRAM and flash chips in a couple of years.
Regardless of how it all goes down, I believe the PC market will come out of the current moment transformed.
AI inference as the impetus
If, like Jensen Huang and, uh, maybe the pope, you believe open-weights models are the key to a healthy AI ecosystem going forward, then surely the ability to run beefy models well on local systems will be part of that picture. For cost and privacy reasons alone, I think local AI inference will become a keystone workload for personal computers. More importantly, I think PC chipmakers believe that, too. Their internal roadmaps surely already reflect it.
PC system architectures will have to evolve in order to handle the demands of these workloads. For LLM inference in particular, that means we’ll need a balanced mix of data-parallel computing power, copious memory capacity, and ample memory bandwidth.
Even before the AI-hardware crunch really hit, we were in an intriguing moment for PC hardware. Qualcomm entered the market with some compelling new SoCs, and Microsoft’s Prism emulator made running x86 applications on Arm processors truly viable. (I bought a Snapdragon-based laptop for my wife—on purpose!) Rumors about an Nvidia-MediaTek collab circled for ages, and then we finally got the GB10 chip in the DGX Spark. New packaging tech has made “system-on-a-chip” solutions into something more like “system-on-a-package” setups with wild variety between the various offerings. Meanwhile, Apple has emerged as the apparent market leader with its M-series processors initially derived from its iPhone SoCs.
With the M2 and M3 Ultra processors, Apple’s platform team built what are arguably the best consumer chips for running generative-AI workloads like large language models (LLMs). I’m not sure what their targets were when they conceived the M2 Ultra. Local LLMs weren’t yet an obviously important workload, I don’t think. But the current two-week lead times and crazy eBay prices for Mac Studios have surely validated their architectural choices.
Nvidia and AMD have dipped their toes into the AI SoC waters with the GB10 (DGX/RTX Spark) and the Ryzen AI Max chips, but they haven’t yet fully committed. The table below tells the story.
In its M Ultra offerings, Apple built SoCs with roughly RTX 5070-class GPUs. They have similar peak compute flops, somewhat higher memory bandwidth, and—crucially—many times the peak memory capacity: 512 GB unified for an M3 Ultra Mac Studio vs. 12 GB of dedicated VRAM for the GeForce RTX 5070.
Meanwhile, memory bandwidth is the glaringly obvious bottleneck on both the GB10 and the Ryzen AI Max+ solutions. They have less than half the memory bandwidth of the RTX 5070, and that bottleneck tends to show up when systems based on these chips run relatively sizeable LLMs. Although they have the memory capacity to load large models, their token output rates are often uninspiring.
This situation is not likely to persist.
Hence my core argument, which is Dirk Meyer’s thesis from 20 years ago: the future is fusion. AMD has been teasing the benefits of CPU-GPU integration for decades while reserving its most capable large-scale implementations for the Xbox and PlayStation. The rise of generative AI will supply them—and their many competitors—with the courage needed to do the right thing.
PC SoC makers are going to build truly large and capable SoCs for AI, with big GPUs and well-matched unified memory subsystems. When they do, they will ship products that call into question the need for discrete GPUs for a great many folks, not just those seeking AI processing power.
You’re gonna like it
This development will be a very good thing for PC enthusiasts, gamers, content creators, and anybody else invested in serious desktop computing. Finally, we are going to get a powerful, complete PC gaming and content creation setup on a chip—or at least on a package.
The potential benefits of this sort of integration should be apparent to anyone who has watched the personal computer evolve over time. Once upon a time, one had to buy a separate FPU and install it in order to handle floating-point math properly. Memory controllers lived on core-logic chipsets, and getting a second CPU core meant populating another socket on the motherboard. We all benefited mightily as CPU makers integrated those capabilities into the main products—which now have an embarrassing number of CPU cores and many times the performance at a fraction of the cost of past systems.
Graphics card memory capacities have stagnated in recent years as GPU firms ran into cost and power challenges. Gamers have felt the pain on this front and made their frustrations known, seemingly to little effect. But SoCs tuned for AI inference will fix that problem in ways it’s hard to fully digest. Imagine a GTA VII game authored for systems where the GPU can access 128 GB of fast, unified memory. Future Los Santos is gonna look insane.
Big SoCs will be more efficient in various ways, too. Unified memory means reduced data movement. Not having to shuffle bits between system memory and VRAM can save time and energy. CPUs, GPUs, and other accelerators can work collaboratively on workloads as appropriate, assuming the hardware guys work out some potentially thorny issues like cache coherency and shared data formats.
If things go well, I expect future gaming workloads to involve a mix of CPU work, traditional graphics rendering, neural graphics, and generative-AI inference. This mix should be better served by a single, shared pool of high-bandwidth memory.
Another traditional benefit of integration is a reduced physical footprint. I feel weirdly scandalized by today’s high-end video cards being the size of a shoebox. I mean, I love big GPUs, and priorities are priorities, but holy crap. Do they make Ozempic for PCIe cards? Meanwhile, the DGX Spark and Ryzen AI Halo boxes are much more compact than a PC with just a mid-range video card. Not having a separate video card plugged into a PCIe slot with its own GPU, memory, power distribution, and cooling should enable some nicer choices—or, honestly, just better use of the space inside an ATX case.
There are cost benefits to the reduced size and complexity of big-SoC systems, too. That fact seems tough to wrap one’s head around right now, given the twin challenges we face in the AI hardware crunch and the slow death of Moore’s Law. After all, the AI-inference-focused development systems I’ve been talking about are currently selling for around $5K. But I still believe integration and innovation will make big-SoC systems affordable in the long run. We’ll likely have a spectrum of price and performance trade-offs from which to choose across a range of SoC sizes. The key thing is that well-balanced large- and medium-sized options should exist in the future—and should have compelling benefits over traditional PC configs.
So what happens, you may be asking, to discrete GPUs? Am I saying that PC enthusiasts will have to give up the awesome ability to pop in a video card of their choosing to upgrade their systems?
I guess maybe that day will come, but probably not for a very long time. Discrete GPUs will likely persist, at least as ultra-high-end options.
But I can also see this situation another way: at long last, the GPU—or, more broadly, the data-parallel computing portion of the PC—is becoming the heart of the system. Honestly, it kind of has been in gaming PCs for quite some time. It is for data-center systems now, too. Upgrading that part of the system will still be something you want to do, and in this new world, you’ll likely get some improved CPU cores in the process, since it’s all packaged together.
So I guess I’m saying that when the time comes, you probably won’t mind.





I have been rolling the same thoughts on the future of desktop PC and SoC evolution for some years. Looking at the console convergence on CPU and GPU integration and performance scaling already charted the path to a possible alternative. Then Apple just broke new ground with their M-series products. The SoC concept is no longer just a slow and skimpy solution for mobile and SOHO platforms. With the pricing crisis affecting the traditional PC component market, the small form factor with powerful SoC offerings will be tempting many (non-enthusiast) customers to ditch the age old ATX platform for a more discrete box. Big towers will still be around for workstation and boutique niche clients of course... with the appropriate price premium. I just hope the ThunderBolt interface becomes a standard feature for external GPU option in those cases when we need to scratch that benchmarking score ego. 😅
Hey Scott, I get where you're coming from and mostly agree.
The one thing that I think you might have missed, is I think AMD and Nvidia have intentionally gimped their little SoC boxes. If they perform too well, it might dissuade people from buying the big iron. After all, if two or four of those DGX Sparks with a theoretical GeForce RTX 5070's worth of bandwidth can run enormous models at acceptable speeds for a small company with 15-20 users, it could hurt the datacenter rollout. It feels like they were purpose-designed to be proof of concept machines you'd use to test your models on and then push them into the really big cloud environments where they can lease you GPU time forever and ever, amen.
I do agree that big SoC systems with actual passable performance could very well take over for many workloads. Nvidia seems to be giving a passing nod to their gaming roots with RTX Spark systems, even if it's just the same GB10 chip in a smaller thermal footprint so it works in laptops.
I kinda wouldn't mind a world where these big SoCs can work in tandem with PCIe add-in GPUs as well. I don't want these systems to be as small as possible and louder than PCs with equivalent power draw like they tend to be.