AI tools require physical hardware infrastructure in data centers, where GPUs handle heavy computation-intensive tasks (like the 'cooking' of model responses) while CPUs manage coordination, routing, and planning across thousands of simultaneous requests. The distinction between training (building AI knowledge) and inference (using AI) is crucial, as inference is what users experience daily. As AI evolves toward agentic systems that perform multi-step tasks, the CPU's role becomes increasingly important, shifting the hardware balance from the traditional 1:4-8 CPU-to-GPU ratio toward a 1:1 ratio or even more CPU emphasis. Understanding this infrastructure layer helps professionals evaluate AI platforms more effectively and distinguish between model problems and infrastructure problems when experiencing performance issues.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
What Actually Powers the AI Tools You Use Every Day?
Added:Here's something that almost never comes up in any AI conversations at work.
What's actually running it? Not the model, not the algorithm, the physical hardware. Every time you send a prompt to an AI tool, [music] a server somewhere in the world runs a computation and sends you a response back in seconds. Most people never think about that layer. Most of the time, you don't have to until a tool your company deployed feels slower than the consumer version you use at home. [music] Until you're in the meeting where someone's asking which AI platform to invest [music] in. Until you're expected to have an informed opinion on AI infrastructure and you realize you have no real framework for it. That's what this video is about. I'll walk you through what actually happens when you hit send, why the hardware layer matters more than most people realize, and what this means for you as a professional using or making decisions about AI at work. [music] I would not have been able to make this video without the sponsor of this video, AMD. More about them later. So, let me first start with what happens when you hit send. Let's start with the pipeline. When you type a prompt and hit send, your request leaves your device and travels to a data center. Think of a data center as a warehouse-sized building, sometimes multiple buildings connected together, filled floor-to-ceiling with industrial-grade servers.
These aren't laptops or desktop computers. They're machines designed to run continuously, 24 hours a day handling millions of requests at the same time. This is where your AI tool actually lives and thinks. Inside those servers, two types of chips do most of the heavy lifting, GPUs and CPUs. If those terms don't mean anything to you, here's the simplest way to think about them. A GPU was originally designed to render video game graphics, but it turned out to be pretty powerful for AI because it can run thousands of calculations simultaneously. Think of it as a team of highly specialized line cooks, each doing the same task over and over and over again at incredible speed.
A CPU is a general-purpose chip, kind you'd find in any laptop or desktop. It handles many different types of tasks and coordinates complex operations.
Think of it as the head chef running the whole kitchen, managing the flow, routing the orders, making sure every station is working together. Most people have heard that AI runs on GPUs. For one specific phase of AI, that is largely true.
But there are two very different phases, and they are not the same thing. And this is where I need to talk about inference and why you should care about it. Think of it this way. Training an AI model is like putting a new employee through an intensive onboarding program.
It takes months. It involves absorbing enormous amounts of information. The goal is to build up their knowledge and capability before they can even start the job. You never experience this directly. It happens once behind the scenes before the product ever reaches you. [music] Inference is that same employee actually doing their job every single day, every time someone asks them a question.
That's what happens every time you send a message to an AI tool. This distinction matters because training is where most of the headlines go.
>> [music] >> Billions of parameters, months of compute, research breakthroughs, that's the flashy stuff. But here's the thing.
You will never experience training or super unlikely. What you experience constantly is inference, >> [music] >> and inference is where speed, cost, reliability, and quality all live.
Here's a way to picture what actually happens during inference. [music] When you place an order at a restaurant, several things happen before the food arrives. Someone takes your order and routes it to the right station. A prep cook gets the ingredients ready. The line cook actually makes the dish, and someone plates up and sends it out. Now, imagine that kitchen is serving not 200 or 300 people, but 50,000 people simultaneously. GPUs are doing the cooking, the heavy computation-intensive part where the model generates your response. But, CPUs are handling everything around it, the routing, the data preparation, the coordination across thousands of simultaneous requests, running smaller, specialized models, managing the outputs before they reach you.
In enterprise environments where thousands of employees are sending AI requests at the same time, CPUs are doing a large amount of work. This is why when a response comes back in 2 seconds instead [music] of 15, it's not just the AI was being slow. The kitchen was understaffed. The wrong hardware was being asked to [music] do too much.
That's an infrastructure problem, and it has a very specific fix. And [music] here's the big update because it's 2026.
Up to now, most of AI you've used works like a vending machine. One question in, one answer out. Agentic AI is different.
[music] Instead of answering a single prompt, an AI agent takes a goal and breaks it into steps. It plans, calls upon tools, queries databases, checks permissions, runs other models, and loops back to check its own work before it's done. Here's the part where hardware matters so much. All of that planning and coordination, the deciding what to do next and routing between each step runs on the CPU. The GPU still does the heavy model math. But, the agent's actual thinking about what to do next lives on the CPU. So, the balance is shifting. Regular AI inference has typically run on roughly one CPU for every four to eight GPUs.
>> [music] >> In agentic setups, we are seeing that move towards something closer to a one-to-one ratio, and in some cases even more CPU than GPU. [music] In simple plain English, the more AI starts doing things for you instead of just answering you, the more the CPU matters. And that's a problem most infrastructure was never prepared for.
Think back to the kitchen analogy for a second, because not every kitchen is built for the same job. Cooking dinner for your kids at home, [music] you've got time, a couple of pans, and that's all you need. A good restaurant serving 50 people in one evening needs more hands on deck, more stations, and everyone fed at a reasonable pace. Now, a fast-food chain putting out thousands of meals a day, each one in five minutes or less, is a completely different operation. Different equipment, different staffing, all of it built for speed at scale. AI works roughly the same way. The right setup depends entirely on the job, how many people you're serving, and how fast they expect a result. Not every CPU or GPU is the same, and sizing the hardware to the actual workload is the whole game. I spent nearly a decade working inside large-scale data environments at a major financial institution, and the tools and the infrastructure that held up under real organizational pressure weren't the ones with the flashiest demos. They were the ones built for exactly the kind of work they were being asked to do.
AMD EPYC [music] processors are what a lot of serious enterprise AI deployments run on, and [music] as agentic AI pushes more of the work onto the CPU, that side of the equation matters more than it used to. This is the fifth generation of EPYC, and it holds 21 world records in data processing and database performance across the workloads that enterprise organizations actually run. Companies and governments around the world choose AMD EPYC for their most demanding operations, from corporate AI enablement to large-scale cloud infrastructure. Not because of marketing, because when performance under pressure actually matters, this is what serious operations have validated and chosen. If your organization is thinking about AI infrastructure, the link in the description takes you straight to AMD EPYC tools and resources. A huge thanks to AMD for sponsoring this video, and now let's talk about what this means for you at work. So, [music] why does any of this actually matter to you? If you work somewhere that has deployed AI tools internally, the performance you experience is a direct result of infrastructure decisions someone above you made.
When a tool feels slow, inconsistent, or frustratingly limited, it's worth asking whether that's a model problem or an infrastructure problem. These have very different fixes.
>> [music] >> Confusing one for the other leads to very expensive decisions. For anyone in a leadership role evaluating AI platforms and signing off on budgets, here's my honest take. You are making infrastructure bets whether you realize it or not. When you choose an AI platform, part of what you're buying is the hardware it runs on. Understanding what good server infrastructure looks like for AI workloads is a genuinely useful lens for evaluating vendors.
Look, if there's one thing you take away from this video, let it be this. You don't need to become a hardware engineer. That's not the point. The point is that AI tools don't run by magic. There's a physical pipeline behind every response you get, and the quality of that pipeline directly affects what you experience.
Understanding this makes you a more informed user and a more credible voice in AI conversations at work and better equipped to ask the right questions when something isn't performing the way it should. The trust gap a lot of professionals feel around AI, the uncertainty about when to rely on it and when to push back, often comes from not understanding the layers underneath.
This is one of those layers. And now, hopefully, you understand it or a little bit of it. If you found this useful, drop a comment and let me know which AI tools you're using most at work right now. And if you haven't subscribed yet, now's probably a good time. Thanks so much for watching, and I shall see you in the next one.
Related Videos

TOP 15 Data compression Interview Questions and Answers 2019 Part-2 | Data compression | Wisdom jobs
wisdomjobs
281 views•2019-06-28

CTS 158: 802.11w Management Frame Protection
ClearToSend
4K views•2019-02-04

NDSS 2019 Send Hardest Problems My Way: Probabilistic Path Prioritization for Hybrid Fuzzing
NDSSSymposium
496 views•2019-04-02

How realistic is Cities: Skylines?
CityBeautiful
159K views•2019-02-14

GUIs & TUIs: Choosing a User Interface for Your Python Project | Real Python Podcast
realpython
2K views•2025-04-04

The OSI Model - Explained by Example
hnasr
225K views•2019-05-12

Cloud Computing - Introduction
elithecomputerguy
98K views•2019-10-07

From Traveler's Dilemma to Dynamic Routing | Demystifying Networking
IITBombayJuly
5K views•2019-08-04
Trending

WOW! Judge TURNS THE TABLES on Trump in His OWN $10B LAWSUIT!!!
MeidasTouch
197K views•2026-07-23

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23