To embed our video on your website copy and paste the code below:
<iframe src="https://www.youtube.com/embed/bioFQN51WAo?modestbranding=1&rel=0" width="970" height="546" frameborder="0" scrolling="auto" allowfullscreen></iframe>
Prashant Agarwal, Intel (00:09):
It seems I have some sort of special talent to pick the last session before the lunch at these sort of events. So don't worry, I will keep it short so we can all have a timely lunch. So everybody is rightly excited about what AI can do, but before even talking about what AI can do, we also need to think about what sort of foundation infrastructure is required. Because infrastructure, compute, and connectivity, they are not anymore the side player in AI story. In fact, they are the story, because without the right foundation, AI is just a very expensive PowerPoint slide.
(00:57):
So before we start talking about the AI-native networks, it's worth stepping back and looking into the scale of the transformation underway. So if you look at this graph, within the last five years, AI has moved from a technology discussion to a business priority. So five years back, around 2020, AI represented a very small number of overall workloads in the data centres. But now at 2025, you can see it's almost 23% of all the workloads are AI. There is still 77% traditional workloads also. But it's not only what has happened till now. If you think about the future, trajectory is even more important.
(01:57):
By 2035, almost 50% of all the workloads are supposed to be AI workloads. Okay, 50% will still be foundation workloads, but the speed with which AI is evolving, is deploying, is much faster than the traditional workloads. And to be honest, it will continue because we will have more agentic AI, automation networks, digital twins, and this will all drive the AI growth. So the real question today is not whether AI is relevant. The question is: do we have the infrastructure ready which can be used by AI to deliver the true value? And to answer that question, we need to look beyond the headlines and try to understand what exactly is driving this AI demand.
(03:01):
Okay, so that brings me to my next slide. So the first slide was important to see how the AI workload is growing. But this one points even the more interesting things here. So when we talk about AI workloads, people think about training large language models on a massive GPU cluster. So training gets all the headlines, but yet it is only the part of the overall story. In fact, inference is where the AI delivers real value. And see, training builds a model, but inference put that models to work. Every time when a customer or application or AI agent is using the model, so inference is taking place.
(04:02):
And as AI grows, so does the demand for inference. In fact, if you see in this graph, currently we see, okay, it's not a very big amount inference, but if you see the scale, by 2035, 37% of the AI workloads are supposed to be the inference workloads, compared to 13% training, and of course, yeah, 50% will be still traditional workloads. So in fact, inference workload by 2035 will be almost three times than the training workloads. And that is very clear distinction at least for telcos, telecom operators, because most of the AI use cases in telco worlds are inference. The use cases, they require low latency, scalability, throughput, and energy efficiency. And as my friend told earlier, they don't necessarily always need a GPU.
(05:23):
So we need to really think about that the goal here is not just to deploy GPU, but we need to think about what sort of hardware is required to run what workloads. Most of these inference workloads can be run easily on the modern CPUs with the purpose-built accelerator also and the software optimisation. And as Warren mentioned, we work very closely to bring those software optimisation, working with the partners also. So now if we think about this, now if we agree that inference is the workload which is going to be required predominantly in future, now we need to think about the next question: How can we run this inference economically? Because we need to agree that AI infrastructure is not exclusively about GPUs. The success will come that how can we make the inference economics better rather than the number of GPUs deployed.
(06:30):
So that moves me to my next slide. So if we agree that inference workload is predominant, now we come: how and where we can run this inference workload economically and better way. So you see, traditionally, the workloads in telecom, they were very centralised. You run the workloads in big data centres, and that's where most of the training takes place. But again, as I mentioned, telecom environment is different. Now, most of the telecom use cases like optimisation and network slicing, those sort of use cases, they require decisions to be made in milliseconds. So what we are seeing, the inference is moving to the place where the data is being generated and the decisions are being made. And that's moving towards the edge.
(07:35):
So now here we are getting a distributed compute, and now we are moving the architecture where we have a core data centre, the regional data centre, and then we have open RAN edge and enterprise edge, so that our whole architecture is getting distributed. But again, the key point is, to have the distributed architecture, it doesn't mean then we duplicate the same hardware everywhere. Because each use case have a different requirement in terms of latency, throughput, energy efficiency. And that's the key part: we don't need to put GPU everywhere. GPUs are very good for training for the large language models and those. But as we see earlier, most of the use cases will be inferencing. And those use cases can very much be done by using CPUs and purpose-built accelerators.
(08:37):
So what's the future? What's the future of AI infrastructure? It's about heterogeneous compute, where you have a mixture of GPUs, CPUs, accelerators, and its success will not come just by the number of GPUs you deployed. It will come by matching the right workload on the right compute in the right place. And that's what we need to think about. So when we move to production AI, the question is not about how many AI models can we train. The real question is how can we inference at scale? And that's the key point. The success will not come by how much technology we are deploying, but by what value we are creating by deploying those technologies. Yeah, thank you very much. I will just leave it here.
It seems I have some sort of special talent to pick the last session before the lunch at these sort of events. So don't worry, I will keep it short so we can all have a timely lunch. So everybody is rightly excited about what AI can do, but before even talking about what AI can do, we also need to think about what sort of foundation infrastructure is required. Because infrastructure, compute, and connectivity, they are not anymore the side player in AI story. In fact, they are the story, because without the right foundation, AI is just a very expensive PowerPoint slide.
(00:57):
So before we start talking about the AI-native networks, it's worth stepping back and looking into the scale of the transformation underway. So if you look at this graph, within the last five years, AI has moved from a technology discussion to a business priority. So five years back, around 2020, AI represented a very small number of overall workloads in the data centres. But now at 2025, you can see it's almost 23% of all the workloads are AI. There is still 77% traditional workloads also. But it's not only what has happened till now. If you think about the future, trajectory is even more important.
(01:57):
By 2035, almost 50% of all the workloads are supposed to be AI workloads. Okay, 50% will still be foundation workloads, but the speed with which AI is evolving, is deploying, is much faster than the traditional workloads. And to be honest, it will continue because we will have more agentic AI, automation networks, digital twins, and this will all drive the AI growth. So the real question today is not whether AI is relevant. The question is: do we have the infrastructure ready which can be used by AI to deliver the true value? And to answer that question, we need to look beyond the headlines and try to understand what exactly is driving this AI demand.
(03:01):
Okay, so that brings me to my next slide. So the first slide was important to see how the AI workload is growing. But this one points even the more interesting things here. So when we talk about AI workloads, people think about training large language models on a massive GPU cluster. So training gets all the headlines, but yet it is only the part of the overall story. In fact, inference is where the AI delivers real value. And see, training builds a model, but inference put that models to work. Every time when a customer or application or AI agent is using the model, so inference is taking place.
(04:02):
And as AI grows, so does the demand for inference. In fact, if you see in this graph, currently we see, okay, it's not a very big amount inference, but if you see the scale, by 2035, 37% of the AI workloads are supposed to be the inference workloads, compared to 13% training, and of course, yeah, 50% will be still traditional workloads. So in fact, inference workload by 2035 will be almost three times than the training workloads. And that is very clear distinction at least for telcos, telecom operators, because most of the AI use cases in telco worlds are inference. The use cases, they require low latency, scalability, throughput, and energy efficiency. And as my friend told earlier, they don't necessarily always need a GPU.
(05:23):
So we need to really think about that the goal here is not just to deploy GPU, but we need to think about what sort of hardware is required to run what workloads. Most of these inference workloads can be run easily on the modern CPUs with the purpose-built accelerator also and the software optimisation. And as Warren mentioned, we work very closely to bring those software optimisation, working with the partners also. So now if we think about this, now if we agree that inference is the workload which is going to be required predominantly in future, now we need to think about the next question: How can we run this inference economically? Because we need to agree that AI infrastructure is not exclusively about GPUs. The success will come that how can we make the inference economics better rather than the number of GPUs deployed.
(06:30):
So that moves me to my next slide. So if we agree that inference workload is predominant, now we come: how and where we can run this inference workload economically and better way. So you see, traditionally, the workloads in telecom, they were very centralised. You run the workloads in big data centres, and that's where most of the training takes place. But again, as I mentioned, telecom environment is different. Now, most of the telecom use cases like optimisation and network slicing, those sort of use cases, they require decisions to be made in milliseconds. So what we are seeing, the inference is moving to the place where the data is being generated and the decisions are being made. And that's moving towards the edge.
(07:35):
So now here we are getting a distributed compute, and now we are moving the architecture where we have a core data centre, the regional data centre, and then we have open RAN edge and enterprise edge, so that our whole architecture is getting distributed. But again, the key point is, to have the distributed architecture, it doesn't mean then we duplicate the same hardware everywhere. Because each use case have a different requirement in terms of latency, throughput, energy efficiency. And that's the key part: we don't need to put GPU everywhere. GPUs are very good for training for the large language models and those. But as we see earlier, most of the use cases will be inferencing. And those use cases can very much be done by using CPUs and purpose-built accelerators.
(08:37):
So what's the future? What's the future of AI infrastructure? It's about heterogeneous compute, where you have a mixture of GPUs, CPUs, accelerators, and its success will not come just by the number of GPUs you deployed. It will come by matching the right workload on the right compute in the right place. And that's what we need to think about. So when we move to production AI, the question is not about how many AI models can we train. The real question is how can we inference at scale? And that's the key point. The success will not come by how much technology we are deploying, but by what value we are creating by deploying those technologies. Yeah, thank you very much. I will just leave it here.
Please note that video transcripts are provided for reference only – content may vary from the published video or contain inaccuracies.
Prashant Agarwal, Head of Business Development, Telco EMEA, Intel Corporation
At the AI-Native Telco Forum 2026, Prashant Agarwal, head of business development telco EMEA at Intel Corporation, discussed why infrastructure, compute and connectivity are the foundation of the AI story, the scale of the shift towards AI workloads, why inference rather than training is the predominant telco workload, why many inference use cases can run economically on CPUs and accelerators rather than GPUs, and the move to distributed, heterogeneous compute.
Broadcast live Sept 2026