To embed our video on your website copy and paste the code below:
<iframe src="https://www.youtube.com/embed/Gb7SnXnH9P0?modestbranding=1&rel=0" width="970" height="546" frameborder="0" scrolling="auto" allowfullscreen></iframe>
Shawn Zhang, Sanas (00:08):
Thank you everyone. Thank you Düsseldorf for having us here. It is awesome to be here. My name is Shawn and I'm the technical co-founder and the CTO of Sanas. And I'm so honoured to be joined up here by my friend and our partner, Nilesh, SVP of AI and strategic growth at Mavenir. And together, the two of us are so excited to tell you more about our vision and our mission. And that is to use real-time speech AI to break down communication barriers and build communication bridges. That's to elevate all voices worldwide with AI. And to begin with our presentation, I would actually love to tell you more about how passionate we are about speech. Number one, we believe that speech is the most emotional, most connective medium of communication because speech is more than just the words that we say. It is how we say it, where those words are coming from.
(01:10):
It's the difference that your mom feels when she receives a phone call, hearing your voice, telling her that you love her versus her seeing that same message over a text message. Second, we believe that speech is the most productive, most collaborative medium of communication. I mean, who here has been part of those Zoom, Microsoft Teams meeting where everybody's talking with their cameras off? And I think personally, I hate those meetings, but people are still talking, they're still discussing, work is getting done. And I say these things because from our perspective, speech is that largest surface area for telcos to provide value and impact to the people. And in the midst of one of the largest technological advances in our generation with AI, that opportunity is bigger than ever. But it's also an opportunity that requires action now because if you look at a few numbers, 42% of the population right now prefers WhatsApp over their actual mobile calls.
(02:16):
50% of B2B enterprise comms are moving away from their telco mobile network over to cloud hyperscalers. 89% of subscribers report call quality issues on their network in the last three months and $25 billion voice ARPU has declined year over year. Nilesh, what is Mavenir seeing?
Nilesh Parikh, Mavenir (02:37):
No, I think thanks, thanks Shawn. I think we are seeing the same shift from the operators. They're asking us, "What can you do for AI, for the voice networks?" Operators are concerned about security, spam, scam, call security. So we're seeing a shift where the operators are asking us what can we do with voice in our networks? And we are seeing a big shift and an opportunity for us to bring AI into voice and we'll talk more about that and we'll demonstrate that.
Shawn Zhang, Sanas (03:13):
Sounds good. And exactly, this is where we believe that Sanas is coming in. So just to give people more context about us, we were founded back in 2020. We have a mission for a more kinder, more understanding world. And how do we do that is that we build our own speech AI models at a speech AI lab. And why that matters to telcos is that we are deploying that directly into the core, into that network. So there's going to be no third-party cloud API, there's no data leaks, there's no skyrocketing token costs. This is what total sovereignty looks like. Just about Sanas, we are right now live with over 1.2 million active users relying on our technology. If you think about the top 20 call centers in the world, they have all been using Sanas. If you think about 20 of the Fortune 100 enterprises, they've been using us directly in their critical CX operations, including Verizon, including Comcast.
(04:07):
But what about telcos? And that's why we're so excited to be partnering with a thought leader and innovators like Mavenir.
Nilesh Parikh, Mavenir (04:15):
Yeah, thanks Shawn. We were looking for a partner who has a mature speech technology that we can bring to telcos. And I was so lucky to find Sanas who has been in the industry in the enterprise space, in the B2B space, who as Shawn explained, they have been deploying the solutions for tier-one enterprises and we wanted to have a mature speech technology that we can bring to telcos at scale. And that's what we are doing in this partnership. We are bringing a partnership to bring mature speech technologies to solve operators' problems, to bring monetisable use cases to the consumer networks as well as B2B networks. So with Sanas's help, we have been able to bring real-time language translations with 34 languages. They have deployed it in the enterprise. Now we're deploying it for the telcos. We're bringing the protected calling with the deepfake detections so we can identify the clone users.
(05:23):
We're bringing adaptive voice which allows us to bring HD voice quality with noise detections, noise cancellation rather, and a number of other features that comes with the AI HD voice that we will demonstrate. And then we have the in-call assistance that during. So with Sanas and with our Mavenir technologies, we're able to insert AI into the pre-call should the call be connected. So we are filtering the calls, screening the calls. Once the call is connected, we are monitoring the calls for any clone or any spam or any scam detections. Also, we're providing in-call assistance. So during the call, if you need to provide any assistance for any reason, so for providing any kind of services during the call, we can provide that. And then AI receptionist are solutions that we are trying to offer to the enterprises, AI receptionist, AI concierge services that we can provide using agentic model on our AVA platform.
(06:28):
So Mavenir is providing an AVA platform to bring all these capabilities to the telco operators on our sovereign AI. This is on-prem. This is not connected to the cloud. We offer both the Sanas on-prem models as well as we are working with the frontier models. If for some reason we need any capabilities that frontier models operator has a flexibility to deploy those frontier model capabilities with our AVA platform. So sovereign AI compliant to GDPR, so compliant by means if anybody's worried about, I don't want anybody using AI to eavesdrop my calls, we provide consent, we are providing compliance. The data stays with the operator's network on-prem and we do not store the data. So it's fairly compliant. It's AI native experience, it's carrier-grade obviously as it's Mavenir and it's monetisable. And by monetising, I think you heard Bijoy from Mavenir talk about tokenisation.
(07:31):
So we can tokenise the capabilities. We can build meter both on voice minutes, on the data and the tokens used. So it's fully monetisable. This solution is not coming in the roadmap. It's ready, it's here and we have a booth here to demonstrate all these capabilities. Awesome.
Shawn Zhang, Sanas (07:52):
No, thank you Nilesh. And onto the next slide. We can sit up here and tell you more about these different solutions. We have something called AI HD Voice Enhancement. On the next slide, I can show you about the different dimensions that we can do to touch up and enhance that voice. Even on the next slide, we can talk about language translation. How do we connect people worldwide globally to truly connect and understand one another? Even on the next slide, we can show us about AI fraud shield and how do we actually protect our subscribers from being vulnerable, being at risk? And I can be up on this stage to tell you more about these solutions, but what more can I do than to actually have you show it, experience and see what this looks like in action onto the actual live demo on stage. All right.
(08:41):
So what I have here is actually one of my coworkers, Luis. And so Luis, he's going to do a bit of a walkthrough through the different solutions. And so the first experience that I want to have you guys feel is going to be Sanas's AI HD voice enhancement. So Luis, why don't you take it away?
Luis Chajon, Sanas (08:58):
Sure. Thank you. Thank you so much, Shawn. So basically, AI HD, what it does is that removes all the noises from a call basically, let's say background noises and network connectivity issues. So for the demonstration, I will place some background noise so you can explain how it works. So I'm playing right now, very loud noise right now. I'm just going to enable right now AI HD so you can hear how it works. Okay? Okay. So I have enabled AI HD right now. I can still hear the loud noises here on my side, but I'm pretty sure that none of those noises are coming through on your side. I can disable it right now and you will still hear that noises. So I'm just going to pause it and you can take it back, Shawn.
Shawn Zhang, Sanas (09:49):
No, thank you. So that's awesome. The second demo that we would love to be able to show you is about Sanas's proprietary language translation. Something that we actually have partnered is actually in the customer experience side. Sanas has actually worked with America's E911. So that's America's 911 services to contact fire department, hospitals, police department. These are critical scenarios where it's where you need to be able to connect and speak and communicate and understand more than ever. Luis, why don't we show the team here more about Sanas's language translation?
Luis Chajon, Sanas (10:24):
Sure. No worries. Thank you so much Shawn again. So before I start, I would like to prove that I can speak Spanish. So I will give you a couple of lines in Spanish and then I will turn on the application so you can hear how it works. So that was just a quick introduction about how much I love to walk my cat and have some exercise as well. So I'm just going to enable Sanas right now and just to give you a little bit of context, what you will hear will be my native voice in Spanish and then you will hear the translator audio in English. I will have a quick conversation with Shawn, okay? So I'm just going to turn it on now. Well,
(11:12):
So as I was saying, I really like going out for walks with my cat and exercising, but I would like to know how the event is going, my friend.
Shawn Zhang, Sanas (11:21):
Yeah, I think the event is going really well. I'm having so much fun being to talk about what does it really look like to be this AI native telco? How's your cat doing? Tell me more about walking your cat.
Luis Chajon, Sanas (11:38):
Well,
(11:39):
The truth is it is a pretty fun action that I have, I really like going out for a walk on Tuesdays and Saturdays with my cat. It was a bit complicated teaching him to go out for a walk like a dog, but I think it is very easy for him to learn.
Shawn Zhang, Sanas (11:54):
All right. Awesome. Thank you, Luis, for showing the team here more about Sanas's language translation. And for the final experience that we would love to do on stage here is actually Sanas's accent translation. What does that technology particularly look like? Luis, if you could please take it away.
Luis Chajon, Sanas (12:11):
Sure, no worries. So as you could have here, I have an accent and I am from Guatemala. This is how I normally speak without any voice modulation or something. So what accent translation does is basically softens the accent so that you can be more understandable to other people. So just imagine you are in a call with someone that you're not able to understand the accent and it's been difficult to understand basically, but you have accent translation on your phone, so you turn on accent translation
(12:46):
And all of a sudden the accent becomes more easier to understand even from the person you're talking to and your accent reduces the noises as well and increases the quality of the sound. So I'll just give you a couple of lines as well so that you can hear a little bit more. So my name is Luis. I am from Guatemala. I live in Guatemala City. I have been walking my cat for around two years already, and this is how I sound with Sanas. So you can take it back, Shawn. Thank you.
Shawn Zhang, Sanas (13:21):
Thank you so much, Luis. Thank you so much for your time. Back onto the final presentation. We just want to conclude our presentation with what exactly does it look like to be a speech AI native telco? Why don't you take us on here, Nilesh?
Nilesh Parikh, Mavenir (13:34):
Yeah, I think Shawn, I think you mentioned about the E911 and I wanted to make a couple of points. Traditionally E911, it takes two minutes to find an interpreter for a distressed person who's calling on the line. And with voice AI, we can do that in seconds where we can provide the services immediately for a person who needs a translation or may need an accent or may have in a noisy environment, we can provide a true mission critical experience with voice AI. In addition to that, I think this voice is here, I think we are trying to now make it monetisable. We want to take voice to bank. So we have identified six or seven use cases right now that we are bringing it live, but we have a roadmap of so many activities that the customers are asking for. We talked about physical AI, 6G and everything, but right now I think if you're looking for something that you can fully monetise and utilize AI, voice is the solution that I think will make a big difference in terms of how AI can impact your top line, bottom line as well.
(14:46):
So I think I'm looking forward to this partnership and growing that. And for those of you, I'll let you close this.
Shawn Zhang, Sanas (14:54):
Sure. No, thank you. And so together with Sanas and together with Mavenir, we are creating the world's first AI native speech ecosystem directly into the telco. And exactly what that looks like is that we're here to drive those conversations with real-time speech understanding and speech intelligence like never before. So what our platform looks like is that we are empowering these conversations through frontier models like Sanas algorithms, like our AI HD, voice enhancement, language translation, accent intelligence, Sanas intelligence, but not only that, what we're also doing is that we're actually having to analyze those calls in real time to detect scam, fraud, abuse, and deepfakes to allow every enterprise, every telco, every user, everybody to go from being reactive to proactive now in real time. Ultimately, our mission is how do we have people speak confidently, to be confident that they're going to be understood, to be confident that they're able to understand, to be confident in the speaking scenario that they're in, no matter who they are, wherever they are and whatever they're using.
(16:02):
Thank you everybody.
Guy Daniels, TelecomTV (16:03):
Round of applause. Come on. Join us on stage. Come on. Great. What a great demo. As I said, it's the first live demo we've had at the forum, so I'm delighted it worked so well. Very, very impressive. Look, we've got a couple of minutes before our networking break. Tony, have we seen any hands in the air for any questions from our audience?
Tony Poulos, TelecomTV (16:26):
Everywhere. Oh
Guy Daniels, TelecomTV (16:27):
Right.
Tony Poulos, TelecomTV (16:28):
Hold on a second. I've got to get all the way over the other
Guy Daniels, TelecomTV (16:31):
side. Well, it's a law, isn't it, that the questioner is always the furthest away from you.
Yashesh Shah, Amdocs (16:39):
Hello, I am Yashesh from Amdocs. Fantastic demo and thank you for detailed explanation. We have been trying with few of the voice agents and wanted to understand how do you handle the real world challenges for real conversations, especially let's say in case of autonomous call handlers in call centre agents. So when you have a real time talk, it is possible that your agent's conversation is overridden by the human conversation or there are latencies, garbled voices, network issues where delayed information is reaching the LLM or whatever is the model behind the scenes. So how do you handle these challenges?
Shawn Zhang, Sanas (17:27):
Yeah, really great question. I think first thing I want to share about Sanas is that we have been a hardened, tried and trusted solution. So before us entering to telcos, we've been in call centres and these are our call centre agents that are sitting shoulder to shoulder, speaking with one another. So beyond just cancelling noises, you're having to cancel voices, actual human voices around you in extremely noisy scenarios, speaking to customers worldwide. And things that we're doing here at Sanas is that we've been building our speech AI models to reflect the real world reality. And so that includes the noises, the rate of speech, the accents, the languages. And I think the final thing about latency, and that's the structural advantage of being an AI native telco, especially because we are integrating these capabilities directly into the network, into the core. You're not having to have those two extra transcoding hops to go to servers and OpenAI and back.
(18:24):
You're actually doing this directly at transit, and that's exactly how you're going to reduce that latency, making sure you're able to provide a more natural experience to all your subscribers.
Nilesh Parikh, Mavenir (18:33):
Yeah. I think the reason we are putting this on-prem is to provide better latency, better accuracy, and better user experience. And obviously it's sovereign. You can experience this, come and see the demo. So when you're actually talking, you will see that you will be able to hear and understand even if you try to talk over sometimes. So there are challenges, but you will be able to see how we are addressing these challenges. So good question.
Guy Daniels, TelecomTV (19:03):
Great. Thanks very much. Tony -
Tony Poulos, TelecomTV (19:05):
Give us one more here.
Guy Daniels, TelecomTV (19:06):
Great. Thank you.
Tony Poulos, TelecomTV (19:07):
Which one?
Unknown Speaker (19:09):
Thanks. Quick question for Sanas. Just what are the inference requirements within carrier grade environment to run the platform with the expected customer experience?
Shawn Zhang, Sanas (19:21):
Great question. So it varies across the different solutions. If you think about the largest solution, it's going to be the language translation. So that's the real time end-to-end speech to speech language translation. Right now with Mavenir, we're running a hundred concurrent streams on one GPU. That's the RTX 6000 Pro, right? So right now we're about 100 concurrent streams for that experience. For the simpler experience, which is like the AI HD, we're actually running 5,000 concurrent streams on that same GPU. And so something that we took to heart here at Sanas is how do we build these models that are specialized, that go deep and are highly efficient because we care about scale, we care about volume. And how do we do that is that we also integrate directly into the telco network so that you're not having to pay the running token usage cost. You have a flat cost and you're able to provide that inference to all your subscribers reliably.
Nilesh Parikh, Mavenir (20:13):
That's another good question. I think the reason we're doing on-prem is to reduce the cost, but we're telcos. We have to provide high availability n+1 or full redundancy. So when you're talking about very high cost GPUs, we want to make sure that GPU is fully utilized. So for live translations, yes, we need more GPU horsepower than let's say for AI HD voice, but we are able to utilize the GPUs with other services that are non-mission critical so that, for example, service assurance or for building the ontology. So we can now utilize the same GPUs for high peak traffic for the voice translation, but the remaining time we can use it for other purposes.
Guy Daniels, TelecomTV (21:04):
Lovely. Thanks very much indeed. Tony, do we have just time for one more?
Tony Poulos, TelecomTV (21:07):
I had one very quick question. I'd like to test the system out and see if it can remove my Australian accent when I speak Greek.
Guy Daniels, TelecomTV (21:14):
Yes, please.
Tony Poulos, TelecomTV (21:15):
Very, very testing, I'm sure.
Nilesh Parikh, Mavenir (21:18):
We can try that. Yeah.
Shawn Zhang, Sanas (21:20):
So thank you so much everyone for listening to our live demos, but also experiencing yourself here at the Mavenir booth. We're here to help you guys and show you guys the magic about Sanas.
Guy Daniels, TelecomTV (21:30):
Right. Thanks very much indeed, because that is about all the time we have. So a round of applause for our two speakers, please.
Thank you everyone. Thank you Düsseldorf for having us here. It is awesome to be here. My name is Shawn and I'm the technical co-founder and the CTO of Sanas. And I'm so honoured to be joined up here by my friend and our partner, Nilesh, SVP of AI and strategic growth at Mavenir. And together, the two of us are so excited to tell you more about our vision and our mission. And that is to use real-time speech AI to break down communication barriers and build communication bridges. That's to elevate all voices worldwide with AI. And to begin with our presentation, I would actually love to tell you more about how passionate we are about speech. Number one, we believe that speech is the most emotional, most connective medium of communication because speech is more than just the words that we say. It is how we say it, where those words are coming from.
(01:10):
It's the difference that your mom feels when she receives a phone call, hearing your voice, telling her that you love her versus her seeing that same message over a text message. Second, we believe that speech is the most productive, most collaborative medium of communication. I mean, who here has been part of those Zoom, Microsoft Teams meeting where everybody's talking with their cameras off? And I think personally, I hate those meetings, but people are still talking, they're still discussing, work is getting done. And I say these things because from our perspective, speech is that largest surface area for telcos to provide value and impact to the people. And in the midst of one of the largest technological advances in our generation with AI, that opportunity is bigger than ever. But it's also an opportunity that requires action now because if you look at a few numbers, 42% of the population right now prefers WhatsApp over their actual mobile calls.
(02:16):
50% of B2B enterprise comms are moving away from their telco mobile network over to cloud hyperscalers. 89% of subscribers report call quality issues on their network in the last three months and $25 billion voice ARPU has declined year over year. Nilesh, what is Mavenir seeing?
Nilesh Parikh, Mavenir (02:37):
No, I think thanks, thanks Shawn. I think we are seeing the same shift from the operators. They're asking us, "What can you do for AI, for the voice networks?" Operators are concerned about security, spam, scam, call security. So we're seeing a shift where the operators are asking us what can we do with voice in our networks? And we are seeing a big shift and an opportunity for us to bring AI into voice and we'll talk more about that and we'll demonstrate that.
Shawn Zhang, Sanas (03:13):
Sounds good. And exactly, this is where we believe that Sanas is coming in. So just to give people more context about us, we were founded back in 2020. We have a mission for a more kinder, more understanding world. And how do we do that is that we build our own speech AI models at a speech AI lab. And why that matters to telcos is that we are deploying that directly into the core, into that network. So there's going to be no third-party cloud API, there's no data leaks, there's no skyrocketing token costs. This is what total sovereignty looks like. Just about Sanas, we are right now live with over 1.2 million active users relying on our technology. If you think about the top 20 call centers in the world, they have all been using Sanas. If you think about 20 of the Fortune 100 enterprises, they've been using us directly in their critical CX operations, including Verizon, including Comcast.
(04:07):
But what about telcos? And that's why we're so excited to be partnering with a thought leader and innovators like Mavenir.
Nilesh Parikh, Mavenir (04:15):
Yeah, thanks Shawn. We were looking for a partner who has a mature speech technology that we can bring to telcos. And I was so lucky to find Sanas who has been in the industry in the enterprise space, in the B2B space, who as Shawn explained, they have been deploying the solutions for tier-one enterprises and we wanted to have a mature speech technology that we can bring to telcos at scale. And that's what we are doing in this partnership. We are bringing a partnership to bring mature speech technologies to solve operators' problems, to bring monetisable use cases to the consumer networks as well as B2B networks. So with Sanas's help, we have been able to bring real-time language translations with 34 languages. They have deployed it in the enterprise. Now we're deploying it for the telcos. We're bringing the protected calling with the deepfake detections so we can identify the clone users.
(05:23):
We're bringing adaptive voice which allows us to bring HD voice quality with noise detections, noise cancellation rather, and a number of other features that comes with the AI HD voice that we will demonstrate. And then we have the in-call assistance that during. So with Sanas and with our Mavenir technologies, we're able to insert AI into the pre-call should the call be connected. So we are filtering the calls, screening the calls. Once the call is connected, we are monitoring the calls for any clone or any spam or any scam detections. Also, we're providing in-call assistance. So during the call, if you need to provide any assistance for any reason, so for providing any kind of services during the call, we can provide that. And then AI receptionist are solutions that we are trying to offer to the enterprises, AI receptionist, AI concierge services that we can provide using agentic model on our AVA platform.
(06:28):
So Mavenir is providing an AVA platform to bring all these capabilities to the telco operators on our sovereign AI. This is on-prem. This is not connected to the cloud. We offer both the Sanas on-prem models as well as we are working with the frontier models. If for some reason we need any capabilities that frontier models operator has a flexibility to deploy those frontier model capabilities with our AVA platform. So sovereign AI compliant to GDPR, so compliant by means if anybody's worried about, I don't want anybody using AI to eavesdrop my calls, we provide consent, we are providing compliance. The data stays with the operator's network on-prem and we do not store the data. So it's fairly compliant. It's AI native experience, it's carrier-grade obviously as it's Mavenir and it's monetisable. And by monetising, I think you heard Bijoy from Mavenir talk about tokenisation.
(07:31):
So we can tokenise the capabilities. We can build meter both on voice minutes, on the data and the tokens used. So it's fully monetisable. This solution is not coming in the roadmap. It's ready, it's here and we have a booth here to demonstrate all these capabilities. Awesome.
Shawn Zhang, Sanas (07:52):
No, thank you Nilesh. And onto the next slide. We can sit up here and tell you more about these different solutions. We have something called AI HD Voice Enhancement. On the next slide, I can show you about the different dimensions that we can do to touch up and enhance that voice. Even on the next slide, we can talk about language translation. How do we connect people worldwide globally to truly connect and understand one another? Even on the next slide, we can show us about AI fraud shield and how do we actually protect our subscribers from being vulnerable, being at risk? And I can be up on this stage to tell you more about these solutions, but what more can I do than to actually have you show it, experience and see what this looks like in action onto the actual live demo on stage. All right.
(08:41):
So what I have here is actually one of my coworkers, Luis. And so Luis, he's going to do a bit of a walkthrough through the different solutions. And so the first experience that I want to have you guys feel is going to be Sanas's AI HD voice enhancement. So Luis, why don't you take it away?
Luis Chajon, Sanas (08:58):
Sure. Thank you. Thank you so much, Shawn. So basically, AI HD, what it does is that removes all the noises from a call basically, let's say background noises and network connectivity issues. So for the demonstration, I will place some background noise so you can explain how it works. So I'm playing right now, very loud noise right now. I'm just going to enable right now AI HD so you can hear how it works. Okay? Okay. So I have enabled AI HD right now. I can still hear the loud noises here on my side, but I'm pretty sure that none of those noises are coming through on your side. I can disable it right now and you will still hear that noises. So I'm just going to pause it and you can take it back, Shawn.
Shawn Zhang, Sanas (09:49):
No, thank you. So that's awesome. The second demo that we would love to be able to show you is about Sanas's proprietary language translation. Something that we actually have partnered is actually in the customer experience side. Sanas has actually worked with America's E911. So that's America's 911 services to contact fire department, hospitals, police department. These are critical scenarios where it's where you need to be able to connect and speak and communicate and understand more than ever. Luis, why don't we show the team here more about Sanas's language translation?
Luis Chajon, Sanas (10:24):
Sure. No worries. Thank you so much Shawn again. So before I start, I would like to prove that I can speak Spanish. So I will give you a couple of lines in Spanish and then I will turn on the application so you can hear how it works. So that was just a quick introduction about how much I love to walk my cat and have some exercise as well. So I'm just going to enable Sanas right now and just to give you a little bit of context, what you will hear will be my native voice in Spanish and then you will hear the translator audio in English. I will have a quick conversation with Shawn, okay? So I'm just going to turn it on now. Well,
(11:12):
So as I was saying, I really like going out for walks with my cat and exercising, but I would like to know how the event is going, my friend.
Shawn Zhang, Sanas (11:21):
Yeah, I think the event is going really well. I'm having so much fun being to talk about what does it really look like to be this AI native telco? How's your cat doing? Tell me more about walking your cat.
Luis Chajon, Sanas (11:38):
Well,
(11:39):
The truth is it is a pretty fun action that I have, I really like going out for a walk on Tuesdays and Saturdays with my cat. It was a bit complicated teaching him to go out for a walk like a dog, but I think it is very easy for him to learn.
Shawn Zhang, Sanas (11:54):
All right. Awesome. Thank you, Luis, for showing the team here more about Sanas's language translation. And for the final experience that we would love to do on stage here is actually Sanas's accent translation. What does that technology particularly look like? Luis, if you could please take it away.
Luis Chajon, Sanas (12:11):
Sure, no worries. So as you could have here, I have an accent and I am from Guatemala. This is how I normally speak without any voice modulation or something. So what accent translation does is basically softens the accent so that you can be more understandable to other people. So just imagine you are in a call with someone that you're not able to understand the accent and it's been difficult to understand basically, but you have accent translation on your phone, so you turn on accent translation
(12:46):
And all of a sudden the accent becomes more easier to understand even from the person you're talking to and your accent reduces the noises as well and increases the quality of the sound. So I'll just give you a couple of lines as well so that you can hear a little bit more. So my name is Luis. I am from Guatemala. I live in Guatemala City. I have been walking my cat for around two years already, and this is how I sound with Sanas. So you can take it back, Shawn. Thank you.
Shawn Zhang, Sanas (13:21):
Thank you so much, Luis. Thank you so much for your time. Back onto the final presentation. We just want to conclude our presentation with what exactly does it look like to be a speech AI native telco? Why don't you take us on here, Nilesh?
Nilesh Parikh, Mavenir (13:34):
Yeah, I think Shawn, I think you mentioned about the E911 and I wanted to make a couple of points. Traditionally E911, it takes two minutes to find an interpreter for a distressed person who's calling on the line. And with voice AI, we can do that in seconds where we can provide the services immediately for a person who needs a translation or may need an accent or may have in a noisy environment, we can provide a true mission critical experience with voice AI. In addition to that, I think this voice is here, I think we are trying to now make it monetisable. We want to take voice to bank. So we have identified six or seven use cases right now that we are bringing it live, but we have a roadmap of so many activities that the customers are asking for. We talked about physical AI, 6G and everything, but right now I think if you're looking for something that you can fully monetise and utilize AI, voice is the solution that I think will make a big difference in terms of how AI can impact your top line, bottom line as well.
(14:46):
So I think I'm looking forward to this partnership and growing that. And for those of you, I'll let you close this.
Shawn Zhang, Sanas (14:54):
Sure. No, thank you. And so together with Sanas and together with Mavenir, we are creating the world's first AI native speech ecosystem directly into the telco. And exactly what that looks like is that we're here to drive those conversations with real-time speech understanding and speech intelligence like never before. So what our platform looks like is that we are empowering these conversations through frontier models like Sanas algorithms, like our AI HD, voice enhancement, language translation, accent intelligence, Sanas intelligence, but not only that, what we're also doing is that we're actually having to analyze those calls in real time to detect scam, fraud, abuse, and deepfakes to allow every enterprise, every telco, every user, everybody to go from being reactive to proactive now in real time. Ultimately, our mission is how do we have people speak confidently, to be confident that they're going to be understood, to be confident that they're able to understand, to be confident in the speaking scenario that they're in, no matter who they are, wherever they are and whatever they're using.
(16:02):
Thank you everybody.
Guy Daniels, TelecomTV (16:03):
Round of applause. Come on. Join us on stage. Come on. Great. What a great demo. As I said, it's the first live demo we've had at the forum, so I'm delighted it worked so well. Very, very impressive. Look, we've got a couple of minutes before our networking break. Tony, have we seen any hands in the air for any questions from our audience?
Tony Poulos, TelecomTV (16:26):
Everywhere. Oh
Guy Daniels, TelecomTV (16:27):
Right.
Tony Poulos, TelecomTV (16:28):
Hold on a second. I've got to get all the way over the other
Guy Daniels, TelecomTV (16:31):
side. Well, it's a law, isn't it, that the questioner is always the furthest away from you.
Yashesh Shah, Amdocs (16:39):
Hello, I am Yashesh from Amdocs. Fantastic demo and thank you for detailed explanation. We have been trying with few of the voice agents and wanted to understand how do you handle the real world challenges for real conversations, especially let's say in case of autonomous call handlers in call centre agents. So when you have a real time talk, it is possible that your agent's conversation is overridden by the human conversation or there are latencies, garbled voices, network issues where delayed information is reaching the LLM or whatever is the model behind the scenes. So how do you handle these challenges?
Shawn Zhang, Sanas (17:27):
Yeah, really great question. I think first thing I want to share about Sanas is that we have been a hardened, tried and trusted solution. So before us entering to telcos, we've been in call centres and these are our call centre agents that are sitting shoulder to shoulder, speaking with one another. So beyond just cancelling noises, you're having to cancel voices, actual human voices around you in extremely noisy scenarios, speaking to customers worldwide. And things that we're doing here at Sanas is that we've been building our speech AI models to reflect the real world reality. And so that includes the noises, the rate of speech, the accents, the languages. And I think the final thing about latency, and that's the structural advantage of being an AI native telco, especially because we are integrating these capabilities directly into the network, into the core. You're not having to have those two extra transcoding hops to go to servers and OpenAI and back.
(18:24):
You're actually doing this directly at transit, and that's exactly how you're going to reduce that latency, making sure you're able to provide a more natural experience to all your subscribers.
Nilesh Parikh, Mavenir (18:33):
Yeah. I think the reason we are putting this on-prem is to provide better latency, better accuracy, and better user experience. And obviously it's sovereign. You can experience this, come and see the demo. So when you're actually talking, you will see that you will be able to hear and understand even if you try to talk over sometimes. So there are challenges, but you will be able to see how we are addressing these challenges. So good question.
Guy Daniels, TelecomTV (19:03):
Great. Thanks very much. Tony -
Tony Poulos, TelecomTV (19:05):
Give us one more here.
Guy Daniels, TelecomTV (19:06):
Great. Thank you.
Tony Poulos, TelecomTV (19:07):
Which one?
Unknown Speaker (19:09):
Thanks. Quick question for Sanas. Just what are the inference requirements within carrier grade environment to run the platform with the expected customer experience?
Shawn Zhang, Sanas (19:21):
Great question. So it varies across the different solutions. If you think about the largest solution, it's going to be the language translation. So that's the real time end-to-end speech to speech language translation. Right now with Mavenir, we're running a hundred concurrent streams on one GPU. That's the RTX 6000 Pro, right? So right now we're about 100 concurrent streams for that experience. For the simpler experience, which is like the AI HD, we're actually running 5,000 concurrent streams on that same GPU. And so something that we took to heart here at Sanas is how do we build these models that are specialized, that go deep and are highly efficient because we care about scale, we care about volume. And how do we do that is that we also integrate directly into the telco network so that you're not having to pay the running token usage cost. You have a flat cost and you're able to provide that inference to all your subscribers reliably.
Nilesh Parikh, Mavenir (20:13):
That's another good question. I think the reason we're doing on-prem is to reduce the cost, but we're telcos. We have to provide high availability n+1 or full redundancy. So when you're talking about very high cost GPUs, we want to make sure that GPU is fully utilized. So for live translations, yes, we need more GPU horsepower than let's say for AI HD voice, but we are able to utilize the GPUs with other services that are non-mission critical so that, for example, service assurance or for building the ontology. So we can now utilize the same GPUs for high peak traffic for the voice translation, but the remaining time we can use it for other purposes.
Guy Daniels, TelecomTV (21:04):
Lovely. Thanks very much indeed. Tony, do we have just time for one more?
Tony Poulos, TelecomTV (21:07):
I had one very quick question. I'd like to test the system out and see if it can remove my Australian accent when I speak Greek.
Guy Daniels, TelecomTV (21:14):
Yes, please.
Tony Poulos, TelecomTV (21:15):
Very, very testing, I'm sure.
Nilesh Parikh, Mavenir (21:18):
We can try that. Yeah.
Shawn Zhang, Sanas (21:20):
So thank you so much everyone for listening to our live demos, but also experiencing yourself here at the Mavenir booth. We're here to help you guys and show you guys the magic about Sanas.
Guy Daniels, TelecomTV (21:30):
Right. Thanks very much indeed, because that is about all the time we have. So a round of applause for our two speakers, please.
Please note that video transcripts are provided for reference only – content may vary from the published video or contain inaccuracies.
Bringing real-time speech AI to the voice network
During this AI-Native Telco Forum session, Sanas and Mavenir demonstrated how real-time speech AI can be deployed within the operator network. They discussed HD voice enhancement and real-time language translation, alongside deepfake and spam detection, in-call assistance and new monetisation opportunities.
Featuring:
- Nilesh Parikh, SVP AI & Strategic Growth, Mavenir
- Shawn Zhang, CTO, Sanas
Broadcast live Sept 2026