
HN841: HPE Melds Apstra and Mist for Self-Driving Data Center Networks (Sponsored)
About this episode
Get every episode summarized
Each time The Everything Feed - All Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
555 searchable segments. Every word is indexed and playable.
Full transcript
The Everything Feed - All Packet Pushers Pods — HN841: HPE Melds Apstra and Mist for Self-Driving Data Center Networks (Sponsored). Machine-transcribed; use the interactive transcript above to jump the player to any line.
Welcome to Heavy Networking, I'm Drew Conry Murray. My regular co-host, Ethan Banks, has been on PTO, but he'll be back to Pack of Pusher's HQ very soon. In the meantime, I am pleased to welcome Kyle Baxter and Shredar Katari from Sponsor HPE, and we're going to talk about automating enterprise data center networks. Now in particular, I'm going to probe our guests about how HPE is integrating its abstract and missed products to bring AI ops and self-driving capabilities into data center operations. We're going to talk about how HPE's leveraging the different strengths of each platform, how the integration can reduce the data collection tax, that network operators have to pay, benefits like predictive analytics, and some real-world examples of how customers are using the platform, and more. Kyle Baxter is Director of Product Management at HPE, and Shredar Katari is VP of Engineering. Gentlemen, welcome to the podcast. First up, before we dive into the main topic, now that the HPE and Juniper Merge Requisition is done, does that unlock possibilities for an enterprise data center operator that maybe weren't available before the integration, like are there other parts of the portfolio that they could hook into? Absolutely. That's been one of the most
exciting parts of now becoming part of HPE, is the ability to have a full-stack solution, story, and offering. Back when we were individually part of Juniper, we were very networking centric. We like to think the network was the center of the world, and it's funny enough in conversations now as being part of HPE. When we say the word data center, we got to clarify that we mean network. Otherwise, people will start thinking, is there a internet service or storage, or applications? What are you talking about? That's been a transformation we've had to make as part of our thinking, but what it's done is truly accelerate what we can do and what we can offer for users by giving them a full-stack solution. It's amazing how well the teams have integrated, how well our solutions have been able to be integrated as well and create real solutions. In just a matter of months, we've been able to truly integrate things like Morpheus for integrating workloads and automations into the networking side of the solution. As well as
integrating to OpsRamp for a full-stack observability solution. That way, when something gets detected, it goes wrong on the network. We can feed that into OpsRamp and it can then correlate that against what's going on in the server and storage world to how that is impacting your applications. That's what's most important in a data center. We like to think it's the networking being in the networking team, but the networking really facilitates the workloads, whether those are applications or AI training or inferencing jobs. That's what's critical, and that's what's exciting is about being part of HP is we can provide that full-stack solution to users out there. Yeah. I think our audience probably believes maybe rightly so that the network is the center of the universe, but there are other important galaxies and stars and stuff. So yes, I can see what that integration makes sense. But we're going to be talking about data center networking. So I want to level set because we are talking about two key products in the portfolio for data center operations, abstract and missed. Can you give us a high level overview of each just so we have some context
for the rest of the conversation? Absolutely. Both were startups that were founded by industry experts that were at other big vendors that looked at the solutions that were out there and what teams were building and decided there has to be a better way. Mist was looking at campus networks where you're looking at wired and wireless devices and how do you manage that complexity when you have users trying to connect and devices and complex radio frequencies that you're trying to sort through in a wireless environment. Mist was fully focused on how do we bring a better solution to there? And they really pioneered AI operations or AI ops before AI became a household name and well before we got to mention it in every discussion now. Absolutely. They solved really the curve early. Yeah, they absolutely did. And an abstract was also looking ahead of its time and what can we do in the data center for a billion and a better way. Many people are still out there
building complex EVPN VX land networks either by hand or automating it with some level automation like things like Terraform or Ansible on top, which is really just scripting the CLI configuration. But what the team then the founders figured out is there's a better way of designing and managing networks with what we call intent based networking. And intent based networking is truly building networks based on what is your intended outcomes, focusing on the outcomes you want to achieve with your network, not that the how and the nitty-gritty details that don't really provide the right outcomes. And so both of them truly pioneered what we can do and we'll talk more about how they've merged together and really bring a differentiation solution. And street or anything you want to add? Yeah, absolutely. So simply put, abstract knows what it should be, misnose why it isn't. So the key differentiator is, abstract provides a graph
and graphical relationship that's what I called as abstract knows what it should be. And then when something deviates, when something goes wrong, you take the AI feature apply on top of the graph and it imp points where the problem is. This is a real deal. It's not a feature tour, a demo tour. This is a real deal in operations. And we're going to talk a little bit more about that, but just to call this out in particularly one of the essential elements of abstract was this graph database that allowed it to sort of understand essentially the full network state and how it was working and whether that state was aligning with what you were expecting in terms of outcomes. Absolutely. That was exactly what it was meant to do and where it provides a ton of value is to being able to build that network based on intended outcomes. And what we mean by intended outcomes is your intent is things like how many racks do you want? What server speeds do
you need? What network segmentation do you need? Not how do you go configure EVBN VIX LAN on this switch or that switch or make this protocol work and how do you get BGP sessions up and running? That's what abstract can take those intended outcomes on how many racks and server speeds and segmentation you need. And it builds that network based on validated designs that we're running every night. We're running millions and millions of tests every night on all the different permutations. And validating that yes, you're going to get the right configuration, the right intended outcomes based on your specific needs. So given that that after was essentially designed to kind of automatically help keep the data center network aligned with these outcomes, why would I want to bring in misdiops into the picture? What is the benefit of the driver for a data center operator? Yeah, that's a great question. And so intent based networking is great at building the network and finding when things deviate. For example, it knows exactly how many BGP sessions you
should and how many routes should be up and established. And when those aren't there, when say you have you know 10 BGP sessions, but you should have had 12, it immediately can can detect that deviation and alert the users. But the problem then becomes kind of what Shreeta was alluding to is, well, what caused that? Why are those sessions missing? What caused that deviation? And that's where AI can really come into the picture and where we can leverage our history where we talked about missed kind of being ahead of the game in that area of network operations is we can leverage AI to be able to detect well, why are those sessions not there? What happened? What caused them to go down? Is there a link that's that's flapping or an optical cable that's been in grade or what is causing that? And that's where we can really leverage AI's to be able to help detect what was the true cause of that issue in that deviation and how to go fix it and even potentially how to predict
before those ever happened in the future. Okay, so it is about sort of the how and the why. Exactly. Yeah. Can you then talk about how abstract and mist would like coexist in a data center? Do you see them as having different roles or operating in different domains or having different responsibilities? They compliment each other so nicely because abstracts we've talked about is what you build your network based on your intended outcomes and your designs. What do you want it to go do? And that is fantastic and it can be able to bring it up faster and more reliable than anything else we've ever seen. We have users that are network novices that that can build an EVPN VXLAN network at scale just or fast or faster and you more reliable then certified network experts that have been around for years and have those certifications and training not what people want to hear but but it helps bring that where you then can enable more people to be able to help and
you're not just dependent on those experts that have those deep certifications to go build and design your networks. You can bring other people in so that way your experts can focus on growing the business and real business outcomes and not just day to day hands in the lab of building the network and understanding what's going on. But how they compliment each other is we can take abstracts intent based networking and the right data. That's where we really like to talk about is abstract provides context on the data. So we're not just collecting the imagery of how many sessions and routes do you have but how many should you have had? Where it should be established, what routes should be up or down and what should be connected, what links should be connected, what speeds, all that information we have all that context. And so when we start sending that telemetry data to miss to do further analysis with its AI is we're giving it the right data. And we've all used AI tools out there and seeing it hallucinate when it has bad data and hallucinate
very confidently. And actually it knows what it's talking about when you're like that was obviously wrong. It's definitely very confident. Yes. But when you have the right data, that's when AI can really provide power and value. And that's what we have when we talk about abstracts graph database and its intent based design and that context of data is we can give the right data. We're not burning AI tokens trying to understand what is the data, what is the design, what should be working, what shouldn't be working. We already have all that knowledge in that context. And so then when we start feeding into AI to find out what happened, how do we solve it, how do we predict this in the future, we have a much more confident models and information we can provide. Okay. Yeah. And that confidence I think would be important because of that potential for hallucination issue, I think is one thing that's holding many work operators from really embracing AI in their day to day work. Absolutely. Okay. Yes. So the, okay, I think I'm understanding
the story you're telling about how they're working well together in that abstract is collecting a ton of telemetry about the network right down to the configuration level. And then it's able to feed that into mist and now mist we know sort of has curated context to be more accurate in its analysis. Absolutely. Yes. That is exactly what we're doing. So we can, we can take that context of data, we can stream it to our AI models and start looking at how do we root cause the problem, figure out what really happened, why it happened, and how do you go about fixing that. So can you be more specific what kinds of data are we talking about? Yeah. Yeah. Shreeta, you want to jump in here? Yeah. So the data that we have is a graph shaped topology, graph database based topology, intense dependencies, edges, vertices, and lastly the telemetry.
Right. If you look at the most product what they do is they keep the interfaces and configuration data in some kind of a CMDB or CSV files, which is not relationship centric, right? Or you can build a relationship as an afterthought, whereas an astra ground up you get a graph shaped topology configuration intent. Think of like a LinkedIn, right? Like everybody's related to everyone and how it's related first and second and third of your network and all that. So this helps AI when you apply AI and top of this, it gives you the efficacy unsurpassed efficacy, meaning it keeps the AI grounded, it keeps the hallucination almost to zero, whereas you get to the answer without prompting over and over again, something you would have noticed when you
are playing with Cloud Code, right? In order to get your answer you're going to ask like different ways, like we call prompt engineering, right? Yeah. You don't need to do all that. It just like gets you the data, gets you the information that you want, right in front of us. And we are continuously working on new MCPs, new tools that interacts with first party, second party, third party products. So it gives you the holistic view of what went wrong. And so Mist AI is able to essentially is Mist AI sort of consuming this graph database or reading this graph database? Yes. Mist AI consumes this data, this data that is coming graph data. And it arguments with what it already knows from the past. For example, HPE Juniper, we have been in this market for long enough. So we have a massive data corpus
of all the customer issues, all the KB articles, all the best practices, all of those have been fine tuned or racked into the system. Right. Now, if there is an issue it sees, first conformals, it looks for, hey, have I seen this issue before? Right. Do I see this in the in the existing corpus? If I say it, I'll give you the recommendation of the bad for you. And then it goes into multi level, if you look at multi level debugging, first thing is it looks for the existing data and the KB articles, etc. Then it looks into the out on internal corpus, which is, which is roughly proprietary for Juniper and HPE. It looks into that and provides recommendation. And then if it can't find, let's say there's a new issue, brand new issue that have you have never seen, then it gets into the log analysis and the code analysis and helps you to
point in a direction where if you hand off this troubleshooting document, for example, which it produces, right? What it did. If it hand off this troubleshooting document to a developer, it for him, it's a very easy to debug. Okay. I mean, the way you're talking about it, it almost sounds like the way network engineer would kind of work through a problem. Have I seen this before if I haven't? Let me start checking knowledge bases, checking logs. That's right. That's right. Okay. Exactly. That's a very good way to think about it is like another season network engineer working alongside with you and making your problem solving much faster in getting to the right, the right solution and how do you get your network back up and running? Get the network out of the way. So the applications can start generating the business you want it to go do. When we talk about intent, it's sort of described as a high level business logic. Can the AI also
understand intent and see if there is a deviation from that intent or is it doing something different? Yeah. Let me take this right. Let me step back a little bit. Let's go through the journey of a troubleshooting fuel. Let's say something happens at the AM in the morning and you know, your page that goes off, you wake up and start looking at where the problem is. Most of the time you are not only looking at the system state at that point in time because things have already gone bad by the time you logged into the system. So you have to go back and look at the log files. In fact, I call that as you become this midnight early morning, V-averse log bundle archeologists. So you go into the logs and you to dissect the logs and then you have this relate that with the time it happened, the time the user, your internal user,
your customers have filed an issue. I'm not seeing this application access and things like that across the other side of the world, for example. And then you are dissecting the logs itself and you are digging through it, the tarballs that you got. Now imagine this happens at scale. For example, in AI, let me quickly pivot to AI cluster scenarios where Juniper plays a critical role. I think we are the only company which offers scale apps, scale out and scale across Ethernet for the AI factory and AI clusters. You are looking at thousands and thousands of GPUs connected and somebody has raised a ticket saying, hey, my job training is slow and things like that. So how do you actually first and foremost get to the bottom of the problem? If it is a not a networking problem, you want to prove it's not a networking problem. It could be storage
issue. It could be Nick driver is an issue or a server is an issue. So first you need to see where the problem is and what is not an issue and doing all this is daunting. So this is where we call the data collection tax. So this is where it becomes pretty daunting and intimidating for any network admins out there. So what if if I give you a simple answer, let's say this AI use case, right? Hey, why is this training job ID XYZ is slow? If I can give a straight answer that, hey, there was been like conditions I've been seeing in the network. It is overloaded. Maybe another job is running mistakenly on the same links that you have configured. Links are not properly partitioned at etc. If I can give you this and also give you recommendations. So we have some innovation that we have done on load balancing technologies in Juniper. So what if I can give
you a recommendation, precise recommendation with command sets that, hey, these are the recommendations you can implement. Wouldn't it be great at 4 a.m. in the morning, right? So you can you can give this to your operators and you can probably go back and catch up on your rest of the sleep. So this is the journey that we are in. So we are trying to reduce the double taxation I call right where collection is one tax and making sense of the locks. So we are reducing the tax. So as a network operator, I can start prompting and the AI will go and start doing that data collection for me. So I like find the logs I need and so on and start looking at all of the other data. Actually, that's a great question. We collect whatever we need in a continuous basis. Okay. Right. As a matter of fact, we auto triage some of the issues. By the time you wake up, you look at your screen, the system has already triaged the problem for you. It's like you,
you know, you go to a hospital like there's a triage group. So the triage is behind the scene. The triage is always running, right? And then by the time you go to the screen, you will see a dashboard or AI agent dashboard, you would see what problems have been encountered and how it has been triaged. Now you can ask more questions and dig into more details of it or you can simply, hey, collect the log bundle, single-click, file a ticket with HP Juniper and boom. It goes to our JTAC and they will call you in next five to ten minutes. So sticking with that sort of log archaeologist metaphor, you said like it's like an archaeologist stepping onto a site where some grad students have already done all of the very fine work to expose whatever it is you were digging up as opposed to starting with just a pile of dirt. Exactly. Exactly. So that's the key here. Yeah. Okay. So data collection is obviously very important to make all of this work together.
How is how are missed in Appstra communicating? Are we talking APIs, MCPs, something else? Yeah. So we have both. We have APIs. Appstra is already rich with APIs. And it has, you know, Ansible SDKs, Interofarms and all that. So with every single operation, Appstra has an API. Now what we have built is MCP that runs on top of Appstra. This MCP now has been added with tools and skills. So just putting together an MCP doesn't help these days. You want to ground that MCP as well. So we have written tools and skills. And how did we do this is essentially go through some of the troubleshooting real world troubleshooting we have encountered in the past. And we have written tools and skills that helps the MCP to invoke certain back and APIs, if you will, right? And
provide this data to the cloud. Now when the when the data comes to the cloud, then it will do more correlation because we also have a memory database in the cloud that does more correlation and correlates with applications, for example. And a few other things. So can you give some examples of some of the tools and skills that you've built into your MCP server? Yes. For example, BGP related issues, BGP neighborhood related issues, right? Now if you simply ask an agent a debug BGP issues, it's going to look at the common knowledge it has from internet and it can take you in the route where you may or may not get to the bottom of the issue. So with our experience in HP and Juniper with yourself working on BGP and building
standards for BGP as a matter of fact. So we have tools and skills that helps the agents and MCP to stay grounded, stay course, stay laser sharp in troubleshooting the BGP related issues. That's an example. Similarly with link flaps, same thing, right? There are 20 different reason links flaps when happen. But you want to get to the bottom of it soon. So for that, the tools really helps to play a critical role, keep the agent grounded to get to the bottom of it. Okay. All right. Can you maybe walk through some examples of real outages where customers have leveraged, missed, and after to resolve them? Yeah, let me take this. So one of the area where customers like us and loved us is let's imagine you have a VM where one of this hypervisor
environment and you have some buckets storage that is either NFS or ISC or ISC. Right? That's VMDKs, for example, the objects for the VMs are residing. It is incredibly different, incredibly hard, I should say, to troubleshoot an issue or flow that traverse through the fabric from your storage to your server or server to your storage. So not only you get to the, if something goes wrong with the story, let's say a couple of links are down and it's a multi-path, load-balance link, a couple of links are down, switch down or leave down, things like that. You are going to get a degraded performance or if it is even flaky link, right? If it is dead in water, you can easily identify it. If it is flaky, it's even harder. So how do you go about troubleshooting this?
It gives you an ability to find the path, first and foremost, the path from the server to the storage and also checks the intent of aftra along the way and where it has deviated and where it has seen problems have occurred and then maps it to your exact ISC and NFS mount volumes, which is incredibly valuable, right? Well, we are in the fire to debug issues and get your application up and running. So this is one example, right? The second example I can call out, which is one of our customer face is in the AI cluster front, right? So the flow stalls, thousands of GPUs are there, the flows are stalling and creating problems with the training. How do you go about troubleshooting these issues? So we have for example, integration with
Javskaydulur's Lexlun, where we understand the training, we understand how the training is distributed in the network, and then we understand the flows in the network as well. So with that, if there is a job training causing an issue, you can simply ask, hey, why is my job running slow? Or you can go to our system to see which job has impacted with which congestion at what point in time. That's key, right? Because it's sporadic, it comes and goes. So you want to identify and catch it when it happens, the congestion happens during the microburst because that's how AI behaves, right? It's all a lot of microbursts that happens. So you want to catch, it's like racecar going around the circuit, right? So it just happens so fast during the turns and twist and turns,
you want to catch it like what exactly was causing that congestion and you want to take an action on it. So then you go into the recommendation where you can, hey, enable distributed path forwarding and various other load balancing techniques. So these are some of the issues that we have seen. There are several of them I can quote, right? Optics issues and interesting things, but overall, these are some of the real world examples that we have come across. In our prep call, another thing that we talked about was this notion of predictive maintenance, as opposed to responding to events, abstract and missed working together can help maybe get you ahead of potential problems. Can you talk about how, give some examples and how the system is actually able to do that? Yeah, so let me start talking a little bit first on the use cases and then treat our lichu, get under the covers a little bit and what's under the HUD. So we started with two main use cases because we've talked about a lot of the rich data and tlemish we have.
So it was a very natural next use case to start looking at how do we predict problems and flip the operator's experience from being reactive to being proactive. And the two use cases that we built and we have live now are centered around optic health and system health. So optic health is your optic transceivers and so there was all sorts of data we're already getting from DOM metrics from temperature and voltage and bias and in areas you can get on the the line from CRC errors and feckers and other things and statistics that you can get in histograms and all this stuff that we can build and look at to be able to help predict when do we see things starting to begin to degrade that could cause issues now and will likely cause issues in the near future. And then system health is all around the health of the switches. So that's what we mean by system as a switch. So it could be a QFX switch leaf spine whatever it is.
And we're looking at common things like CPU and memory to see if there's a potential memory lake for a specific process that's beginning to creep up and is eventually going to cause that process to reboot packets to get dropped and then unknown error that you're troubleshooting the next to going right. I don't know what happened. Something happened. It's not happening anymore. I intermittent issues are the hardest issues to troubleshoot and solve and it's like the bane of existence for everybody that's tried to troubleshoot something. Jason goes to network problem or software problem is just intermittent problems are hard. And that's what we start to see when you start seeing memory leaks or optics degrading is you get these intermittent problems. And we can be able to use AI to help look at how do we predict when these are going to happen in the future and give you the information and should are going to elaborate a little bit more on this here in a second on what is all the information to help build that trust on why we're making this prediction why the model made this prediction what's going to happen what
could even be impacted because we also have all the context on the network and know what applications and jobs are running. We can tell you that hey this is going to be a pretty critical impact because you have these key jobs that are running or these key applications that are running on this optic that is about ready to fail on you. So trade our hand off to you talk a little bit more on depth on some of the what's under the hood. Yeah. So the simplest analogy I use is the space and time dimension. So so far what we talked about gives you the space dimension. But then you want to have the time dimension. So that's where model comes in handy. If you really look at it the dog metrics alone there are about 3035 metrics that you can collect from digital optical metrics depending on the optics. On top of that you have system level metrics like what is your ambient temperature is what's your various other characters how's your CPU looking and things like that right
and what's the other side telling you like okay you are you are you are measuring on one side of the switch what's other side's health like if the power are compensated in the way optic works as they compensate each other and optic a lot of I would say rule based intelligent built-in already like that's why like forward error correction, bit error rate corrections and all that. So what happens in an optic situation typically is optics continues to perform it doing fine but then like starts doing errors you don't really pay attention to those right who pays attention unless it fails right so it goes and then continues and then and boom one day we call this clip zone right it just drops and it stops working and then you are doing the you are in the process of scrambling okay which optics I need to put do I need to put the compatibility compatible optics and things like that layla there may be second party third party optics etc etc right so you are really like spreading your hair when something goes wrong in an optics so we want to detect this
way ahead of time right so if you look at our system it tells you the overall optic health and also tells you at what what do we see in the optics that we think it might fail in next one day next two days next two weeks three weeks in time horizon so that's what I call there's a time dimension so under the hood we use both mission learning and deep learning models you know without getting into too much details of it and of course you know with with the augmentation of statistical modeling so we use all this three different models and more and data scientists the job is to create more features or work on new features on these modeling techniques and provide fundamentally a way to detect the optic failures that before they happen before that happens so that you can you can give a heads up your supply chain person in your company hey I may have to order
like no golden optics right these days I think you need a time machine if you're going to give somebody a head up but yes exactly yeah supply chains and nightmare these days yeah but I take your point like all of the potential metrics you're describing even just four optics I'm imagining a computer screen with dozens of dashboards and a human trying to look at that and make sense of it doesn't really stand a chance but this is exactly the kind of thing that machine learning is for being able to gather and correlate and analyze and keep an eye on those second by second minute by minute metrics that you know if you've got the right context and the right predictive math can say like I'm pretty confident that this is going to go down in seven days based on all the other data I've seen about this kind of situation is that the idea yes exactly okay and do you provide any kind of like scoring to say I'm you know sort of 95% confident 85% confident about this prediction because that's also important I don't want to chuck a good optic because it's expensive
correct correct yes you nailed it actually so we have a confidence scores right like we have like multi-dimension metrics the confidence scores the time to failure and the blast radius for example that's interesting yeah let's say optics facing not side of the spine for example right it can take down multiple networks at same time right so we want to have we want to give a view of hey the blast radius is pretty high so even though I'm only 70% confidence you might want to start seriously considering before you go for this weekend who knows right so things like that so we want to give the blast radius and how many applications going to get impacted it not just a fabric so since we have the application we know exactly what applications are going to get impacted if the optic goes down I think that's another interesting part of this because saying optic down sounds like a problem but if it's
a set of applications that okay we can live without it for a day or two or whatever versus how much it would cost us to like you know take the switch out of commission to reinstall the optics or whatever you have more context to make a business decision when you can tie it to applications affected yes yes exactly right interesting and how do you see folks using the predictive analytics around CPU memory and storage so well first like the two areas that we are focusing right now one is the systems like Kyle mentioned systems areas so what is the system system is first let's start with ground up right CPU memory storage right that is sitting in the box right that's like you know hardware and then we get into the platform level stuff right how's your you know platform drivers are doing for example are there memory leaks that are processed leaks and things like that right so these are things that we the predictive analytics continuously looks out for right now it so that's
on the fabric health and fabric systems right now where we are going is we you should able to ask a question that okay I see there is a memory leak are there patches available for it right are there because because we like previously mentioned we have trained our internal corpus and the release strains you know product release strains are versionings and all those in our internal trained LLM models we can quickly correlate so say hey this release under certain conditions right you may run into let's say if you configured certain configuration into the box it may lead to leak for example right so these are the things that we want to
and it's fixed in particular release right we'll give you a fixed recommendation and so you can go ahead and apply this apply this fix us so so there is a link between what predictive maintenance does and the recommendation engine we talked about early on okay so we've been talking and sort of the examples we've been given you know the apps and missed can help a network engineer who's coming into a situation have better context around what happened and maybe what to do next maybe even recommendations you know from the system on next steps but we're also I mean HP in juniper have been talking about self-driving and network automation for a long time which to my mind means like maybe I'm not even getting in that 3a I'm call maybe missed an app store have just kind of fixed it for me are we very yet we are very close we are right on the the crest of getting there because where we've been talking about this whole session is having the right data and
getting to the right insights yeah and getting to that efficacy that treat our talked about earlier that is absolutely key and how we build trust and how we can then get to the point of all right we've seen this a hundred times a thousand times and every time it's the same solution why don't we just go ahead and automate that for you and and just take that off your plate so that way the next time it happens at 3 a.m. you don't get paged we just do it for you and then we can you know log that event we can you leave me a little bit slack to say yeah yeah yeah yeah yeah so you're sleeping yeah so once you get up and you're drinking your coffee and you're like oh look hey it fixed it here's why yep that makes sense all right cool I'm good let's let's keep doing this and it's and it's going to be an opt-in model because we know everybody is cautious as they should be sure right it's the same thing with self-driving cars and and we've all seen the evolution there where you know at first we were like oh my gosh this is going to be the scariest thing ever I'm never
getting in one to now they're all over the place and and and they're they're just you know used everywhere and so that's the same evolution we're getting to is we we've started with the right platform the right data and the right capabilities and then build the trust that say hey look we can do this over and over and over again and build that efficacy that we are providing the right recommendations and so let's go ahead and start automating those and so that's where we're going later this year is to start rolling that out in data center for the first few use cases where we can look at how do we begin to bring those self-driving actions to data center networking operations so can you talk about that that trust building process how how am I as a network operator in my day-to-day interactions with mist and abstract it's sort of starting to learn that yes I can trust the output from these systems yeah so I'll start a little bit and then Trader you can you can add to it but it but it's all about showing the users the right information so it's all about
presenting what we're finding why we're finding it information so some ways it may look a little verbose you know and and you get all these dashboards information but it's to build that confidence and trust that hey yeah we did run these commands we did collect this data and get this insights that why we made this recommendation yeah and to be able to provide that that whole trust into what the system is doing for you I think that's actually I'd rather get too much information than not enough and have like footnotes essentially they actually go back to actual dockiness yeah so okay should I did you want to weigh in on that yeah see trust is and when it shows it to work right and right often it's right often enough at some point you will stop checking so are we there now we are not there but we are at the point where we want to earn the trust with the customers right by by better reasoning even on the predictive area
we want to give you reasoning right in fact you can really geek out on optical metrics that's the capabilities that we have offered which no other product in the industry offers right you can go and see how the signaling look like and how the you know is a negative one DB or positive one DB and and the power and variations and and bit error rate all this in a correlated metrics you can see all those graphs if you like to right if you like to because we're going to our answers are based on what we found on all those 32 different metrics right and we will tell you okay this is when it's going to fail right but then now you have an ability to double click and go into each of these 32 metrics and see what is happening right and and we want to give you at we want to give indication that this is the reason why we we arrived this conclusion that I'm 80% confident your optical is going
to fail you know on Sunday so no yes I mean like as more and more of these confirmations or I should say like we can bring into you at some point we will we will again you trust and and the one of the pointers the data is is a key right we need to have large set of data to do our trainings and modeling without data model modeling is pretty useless yeah so Juniper internally we have AI labs where we'll be done that we have all the you you call any Nvidia AMD advanced GPUs we got all those in our labs sure and we use it for uh essentially testing various benchmarks of training let's say you know um
Kimi 3 came out we want to see how it does for example at Gwen how is it doing how is it doing with inferencing how is it doing with training and etc etc so so that we are in the position to a recommend customer hey this is what we this is these are the models we have tried and this is the performance benchmark that you can get now it's slightly off topic but this is what we are working with many of our new cloud vendors to give them the confidence that how we understand the training in inferencing area right now they the one more outcome that comes out of continuous experiments going on the lab is I get the data from the lab like I get how the optics are performing when you're when you're running for these benchmarks and that's a huge advantage that we have in addition to that we have thousands of switches in our labs sitting in various facilities various labs we we
have a data lake that collects all the data and we train on the data again a site point right all these data used for training and then that's what that we can do a more trustworthy inferencing for sure yeah yeah no I would want to know that that's exactly what you're doing I'm curious you know if we if customers become comfortable enough to maybe flip on a few automated fixes and responses how does that manifest is it like a missed agent logging into devices and making changes is it missed saying to abstra hey make this configuration change what's the actual mechanism it's the the latter so we'll leverage like we're talking about APIs earlier um abstra has robust set of APIs for everything you can do and so once missed in there in the agent determines that hey this is the recommended step that we've seen a thousand times to go automatically solve go do it it'll use the abstract APIs to automate that action and make that part of your intended design and and gets
back to that that loop we were talking about the beginning that now that is part of your intended outcomes and what your intended network should be looking like so then it becomes part of the state every network so then we can continue to compare that against that new state that we've automatically fixed and resolved for you okay is the automated system also able to roll back if for whatever reason it changed and have the desired outcome is that something that it's capable of absolutely so within abstra we have a feature called time booger very similar to time machine if you're a Mac user and where snapshots your entire fabrics this isn't a switch by switch hey we've stored the configuration this is looking at fabric wide snapshot that we can roll back to any known point in the past and be able to do that and basically a one click or one API call from an automation perspective and so absolutely it's a really easy way to to roll back whether it's an automated change or from a user perspective if they see oh yeah I see why you automate that but actually
I didn't really want you to do that I'm going to go roll that back and if an automation system makes it change does it also follow the same workflow that a human would in terms of I've opened a ticket I've worked through the ticket I've closed the ticket I've updated X other systems associated with you know my goal yes yeah and that's one of the things also we've done as being part of HB is we've now we have more integrations and we've integrated with things like service now which is probably the most popular yeah yeah yeah it's too out there from ticketing perspective and so same thing we can automate and you know either open tickets closed tickets when when actions are resolved and automated and be able to update the status of them so that way when when the users all wake up in the morning and they log in that's it's all there for them the history is there and it says yep it was resolved by by Marvis and and here's what we did and why we did it and we've closed the case now okay so just a couple more questions before we wrap first Kyle you mentioned at the beginning miss did start as a product for the wireless campus then moved into wired campus now the data center
so I'm clearly seeing a progression here is there a long-term objective to build a platform that spans network domains absolutely and that's really what we've done with missed as we've turned it from a product where you mentioned it started in the wireless and then the the wired land space into a platform and so we're leveraging that same platform the same capabilities that rich history that missed his built of solving networking issues yes the data center doesn't have wireless access points but it's still networks and packets moving that we can use the same framework same use cases a lot of the same learnings to be able to apply to that and so by turning missed into a platform we've now can expand it to other areas of networking so including data center including the WAN including even security domains and so then once we connect all those together in the cloud the next easy obvious next step is well why don't you do end to end and that's exactly where we're going to be able to correlate because the most common problem that that users run into is user
reports an issue and then you got to go ask three different teams is it the the wireless having an issue oh nope the wireless looks good oh is it the WAN oh that looks good or there's an issue or then you get to the data center oh where's the issue and then you get the finger pointing well it's not me it can't be me it's you it's so that's the problem we want to get to and solve is instead of having to swivel chair to multiple tools multiple teams multiple platforms how do we give our users one solution where they can see the entire domain from the user to the application and the workload and troubleshoot from that end to end view yeah just to add to it if I may yeah just you know in the past it used to be you have wireless admins wireless admin or campus admins and then you'll have data center admins and then you have security admins and WAN admins and things like that so what we are seeing lately from our customers obviously you know
due to resource crunch or expansion of the data centers which has exploded in past few years or so they are looking at a site level admins right meaning if you are if you're a site level admin there's no one else there to punt your problem too right so you go to a portal you ask hey why is a kiles application not reachable and kiles saying his application is not reachable for example right if I ask a question regardless of which domain the problem is right it could be wireless it could be wired and campus the switchers sitting the closet in the storefront or in the in the van links are there's a firewalls before it gets into the internet and comes out back into the data center so you want to have the complete handle of troubleshooting technologies built in
so so you get the answer you want right see maybe you will not able to fix everything right um you know maybe you don't have some of the permission to fix a storage problem for example that is sitting in the core data center somewhere in the remote location but at least you know that is the one that is causing the problem it's not the wireless it's not your local site right so this is very powerful many customers are asking us for this kind of an end-to-end strategy and end-to-end solution for across domains so you've got competitors there are other big network inventors who are building out or have some kind of AI or automated management platform there are third-party automation tools there is the notion of you know bring your own AI model into your ops so what should customers be thinking about as they think about data center network automation
yeah so when we talk to network operators and interview them and so what are the common problems they run into usually they they fall into three different categories speed reliability and insights speed it's it's the challenger how do you keep up with the the fast-paced industry out there AI is making everything move so fast that what is new today is now old tomorrow and so you have to keep up you have to be able to deploy new services deploy new networks all that quickly and that's a challenge when you're trying to do it either by hand or trying to script around it or build your own tools or even other vendor tools out there is they just can't keep up with the industry needs and networking to be able to deploy fast and to me that also then goes hand in hand with reliability because if you're just deploying things fast if you trigger figure out how to automate things quickly how do you know you're getting a reliable solution
and that's where then the next point we run into is is people have automated something but then it's always breaking it's always causing problems and outages and and then they're having to redo work and go back and revert things and figure out what happened and that's where our solution truly differentiates is by looking at intent-based outcomes we can give you that reliability we can use the AI appropriately and intelligently to be able to make sure that yes it is continuing to work as you intend and when it doesn't why and how to go fix it or even predict it like we were talking about and then that insights is that third category is a lot of users we talk to just have no idea what's going on in the network they don't have all these this information and they're lacking of what is actually going on especially when you're talking about how do I know what applications and jobs are running my network and what are they utilizing my network they might have good information on the network but they're lacking in that application in-site because whether that's
you know some other app tool that some other team has or they don't have that access to it network operators tend to struggle with having that right insights of seeing the whole solution right and that's where we can get that value for our users and and network operators out there is giving them that insights and understanding of how their network is being utilized what is using it what jobs are running what applications are running and when things go wrong what is impacted what's that blast radius how severe is that issue what do you need to go prioritize do you need to just go drop everything and pull all hands on deck or is it that's something we can deal with tomorrow and or next week and so that valuable insights is differentiating in how network operators and can operate in their day-to-day lives that is truly something that we are out there to help solve within hb in the solutions that we're building and being able to provide from traditional application-based workloads to AI data centers where you're building training inferencing jobs the the whole solution we have offerings for everything out there treat our
anything to end no Kyle covered it all the three-billars right speed reliability insights so these are the key areas where you know we excel up all right well that brings us to the end of this episode are you guys online do you blog do you do do you have a tiktok or anything or work on folks find you if they want to reach out for me it's linkedin so that's where I'm active I post new things that are coming up interact with with our other users and network operators out there so look out for me on linkedin okay treat our yeah same here look out for me and linkedin I have written a few blocks in linkedin in this area a ops prediction how helps how agent stays grounded and hallucinates less with the graph and things like that so please read up and happy to get your comments yeah I'll also
drop a link into the blog post you wrote treat our about that that notion of data tech collection which I think was was really insightful so that'll be in the show notes that accompany this podcast you'll also be able to find hpe on linkedin and on x and you can we'll have links in the show notes to learn more about hpe networking and if you want to see apps for data center director yourself and sign up for a virtual lab we will also have that link in the show notes so take that opportunity to get your hands on it play with it for yourself that does wrap up this episode thank you Kyle thank you Shredar for being here and thanks hpe for sponsoring sponsors make everything we do possible here at pack up ushers if you are listening happen to investigate hpe after and or missed please tell them pack up ushers sent you that will help us a lot and also to you listener you are the reason we're doing it you matter the work you do is important we're here to help you develop your career and support you so thank you for everything you do don't forget there is more to pack up ushers than just heavy networking we've got more podcasts we have a youtube channel we have our weekly newsletter human infrastructure if you want to be a little more social we
have a select channel you can join to hang out with other network engineers and talk shop all available all for free at packupush.net and always remember that too much networking would never be enough
More episodes
More from The Everything Feed - All Packet Pushers Pods

TNO072: Connectivity and Community with Jason Gintert
The Everything Feed - All Packet Pushers Pods

LIU022: Chris Grundemann – From Pulling Cable to Network Automation Forum
The Everything Feed - All Packet Pushers Pods

D2DO312: Networking at Scale: AWS Transit Gateway War Stories
The Everything Feed - All Packet Pushers Pods

PP125: News Roundup—Cyberattack Impacts Pacemakers, OpenAI Publishes Eye-Opening...
The Everything Feed - All Packet Pushers Pods