
IPB207: Flying Blind: Monitoring Might Not See IPv6
About this episode
Get every episode summarized
Each time The Everything Feed - All Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
339 searchable segments. Every word is indexed and playable.
Full transcript
The Everything Feed - All Packet Pushers Pods — IPB207: Flying Blind: Monitoring Might Not See IPv6. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Welcome to the IPv6 Buzz, where we dare to dive into the 128-bit address space wormhole I'm Ed Horley. And I'm Nick Barolio. We discuss everything IPv6 on the show from strategy, design, deployment, operations, and even internet standards. And I'm Tom Cuffine. We've spent 20 plus years working with the IPv6 protocol. We work on getting IPv6 working. And we're here to share some lessons learned and how to avoid common mistakes. Hey folks, welcome to the IPv6 Buzz. Glad you joined us. Today we're going to be talking about monitoring because we think there's some issues overall for the industry around monitoring. And we thought it would be a good chance to chat about the state of monitoring for IPv6 overall. And the things that we should be paying attention to or thinking about in regards to IPv6 and monitoring. And what turning on IPv6 within your network environment may have impact wise on monitoring, which I think is probably the important stuff. Because you probably have a set of standards about what you're doing from monitoring today
for network visibility. And you want to know what gaps you might be addressing or dealing with by turning on IPv6. And so we just wanted to talk about that real quick. So, you guys, you know, what do we think is the main key thing that we want to sort of touch on about about this and what's broken. I think the biggest thing that steps out your jumps out to me is sort of the your monitoring platform may not be V6 aware, right? In terms of either being able to ingest a data stream flow and that flow or something else of that nature as flow from the IPv6. That could be one problem. And then the second problem being that anything that's looking at that doesn't really understand IPv6. And so it's either ignoring the packets or doesn't know what to do with them. That would be the first thing that jumps out to me. And I think that's a pretty common issue for folks who are working with vendors. I don't know Nick. Yeah, I think you've had some experience in this particular area. Yeah. So there's a bunch of different aspects of this from
my point of view, right? You have, does my monitoring application or suite of applications communicate over IPv6? Like can it talk to the network elements over IPv6? Can the network elements communicate over IPv6 to export their data or to allow polling of their data? And does the monitoring platform understand what that data is and can't do something with it that is analogous to what it does with IPv4? Right? And then there's sort of the subtleties of is there, you know, parity between what the application or the network element or whatever it is you want to monitor is there parity between what it gives you with IPv4 and what it gives you with IPv6? So I guess there's four things that come to my mind with that. And that is a difficult matrix to to have as complete in my experience. There's generally one part of that that isn't quite right
or it isn't quite there. And that's just, I mean, that's much better than it used to be, right? And you can run a network that is wholly IPv6 only and have pretty much everything, but you got to do your homework and do your testing. Yeah, I think for me, I found at least early on that the majority fit in the category of they just didn't support IPv6's and point to ingest data. I think that's come a long way from, you know, a decade ago or even five years ago, where much more of the platform support is there from the devices themselves to be able to send data via IPv6 and then for the endpoint to actually ingest that data that's coming to it. And then secondarily was the time it took for vendors to catch up with doing IPv6 specifics sets of features and capabilities, right? Of understanding the discovery or understanding other aspects of the packet adjusting type that they were getting, whether that's multicast traffic or something else of that nature. And so I think most of that is pretty caught up
in terms of at least those two categories. I feel like they're in a reasonable state right now. You can find a vendor to solve this. It may not be the vendor you used today, but you can definitely find a vendor who can solve these two particular aspects off the checklist. I guess my issues are probably more fit in the third category, which is things like event correlation and talking about events that might be happening over IPv4 and IPv6 that are actually the same event. They're just being split across the two different protocols. And can you event correlate those two items, regardless of which networking protocol they came in on? Now, there's a lot of systems that can do this, like as an example, an IPv4 with two different IPv4 addresses. They can still event correlate and say that's really the same thing that's happening there. It's we're being exploited across two IP addresses. And this is actually a more complete holistic view of what's actually happening on the network. I have not seen a lot of systems that have the capability to do this across v4 and v6 in the same way to sort of build a more unified view of what's going on in the
network. And if you have a bad actor who is doing lateral side moves or do extra-tooling data from one protocol to the other and then sending it off, I haven't seen as much around the sort of security capabilities of products to be able to do that. And I don't know, no nick or tom, what your experience has been, but that's been mine. And so far, I feel that's a pretty big gap in the industry, right? This is for v4 or v6 because obviously it's a dual protocol problem. It's equally bad for both protocols. But I think because of the fact that folks are very commonly deploying dual stack that this becomes a more clear and present issue for folks to really sort of understand what's happening. I guess one hedge based on where things are sort of headed with IPv6 mostly, the fact that the infrastructure kind of remains in a dual stack configuration, which then supports whatever monitoring regime that you that you have in place, not without having tibb worry that it's going to work over IPv6 because as you're pointing out, there's kind of feature
functional parity question about whether you can deliver data over IPv4 or IPv6 and then of course treating that data as telemetry to make decisions about changes to the network and that sort of thing. I think that's kind of reassuring. I'm just saying that if you imagine a scenario where everything has to move to IPv6 only in both the infrastructure and the edge, then you run the risk, I think, of having monitoring applications that can't handle an IPv6 only environment in terms of delivering the data. But I think, yeah, so there's that. But then I guess maybe the larger issue, I think, in my mind, is with enterprises, and maybe you guys have a different perspective on this, but I'd like to hear the opinions on and from the listeners as well. The monitoring regimes as they're constituted with the enterprise software and infrastructure development, how baked in are those to the point where it's not even really feasible if you have a monitoring platform that is not performing on the V6 side and doesn't have what you need it to have,
how feasible is it to move away from whatever workflow you have related to monitoring today? Just basically re-build the system in parallel or swap out a system for monitoring that's going to give you the data that you need with IPv6 as you deploy it. So I guess I'm just wondering how much sort of vendor lock in with the substandard monitoring with IPv6 is happening in enterprises at scale. My experience with that has been that even when you have a fully featured monitoring system moving to a new one, especially if you have a network of any given size, it's extremely difficult. Yeah. Like it's very ingrained. It's probably been there for however long. I've seen systems that have been in place for 20 years that it's you're not taking it out. Like if it doesn't do IPv6, you're basically waiting until it does. It's such a heavy lift because at least from my experience, and again, this is maybe not the this is maybe atypical, is that monitoring doesn't necessarily have
as many eyes on it as probably not a dedicated person that does it or set of people. It's the engineering team maintains the monitoring framework because they need it for their on call rotation or they need to be able to feed it to the knock so that the knock can do the you know, the triage and it's very difficult, extremely difficult to change them. Again, maybe atypical, but that's been my experience more or less my entire career. Yeah, I would have to I would tend to agree with you on that one Nick. That's been my experience also is that changing monitoring platforms is incredibly number one. It's a very large project overall. It's actually not a project that can be cut over. Typically, it's one that you migrate services as it makes sense and you may have certain sets of dependencies, third party dependencies with many of the monitoring platforms because they may have specialty integration for particular applications or product suites. And so you're you're stuck dealing with this particular monitoring platform because you're
the one that you might want to choose doesn't have a direct integration in the same way that your current one does and it solves a particular problem or alerting or you know, something super structurally important organizationally within, you know, your business. And I think those are the the big challenges of monitoring platforms because you know, they're and many of them are tied into or alerting in performance, you know, thresholds and so that data becomes super critical. And you have folks that basically get triggers to other applications or services based off of what this monitoring platform is reporting and you know, replacing that side of capabilities is actually pretty difficult in a lot of environments. And so I think those are all super valid points, you know, and that your observation is pretty accurate with mine in terms of working through enterprise customers and what we've seen many of them do. And the challenges they face when trying to switch out to different monitoring platforms themselves. And there's a there's a huge swath of
tiers of monitoring platforms to from the SMB side all the way up to mid-level commercial, all the way up to super huge enterprise and then to search water, right? There's just different categories of what's available. And then there's all the shops that are building their own on open source tooling and putting together all the right sort of combinations of things they need to do. And you know, that's that becomes a whole different set of parameter challenges for folks in terms of, you know, what is built into the system versus what's supplied by a commercial vendor support and what key insights you can fair it out of there is really dependent on what the code is, you know, makes available to you. So if you're putting together your own stack from open source, you might be in a better position to be able to get things done that you want or to be able to swap a particular component out that might allow you or give you better flexibility. But you're not going to be able to tap someone on the shoulder and ask for features in the same way that you could maybe with commercial product. And or to further complicate things, maybe you're
running a commercial application that you've written custom things, you know, custom hooks into that don't translate into the next thing or you've modified an open source project in such a way that it is not upgradable or it isn't, you know, you can't easily create the same type of capabilities in another platform or even worse, you may have multiple monitoring platforms that exist for, like let's say, the network. They have their own monitoring platform and then the systems people have their own monitoring platform and then, you know, the folks that are running Kubernetes and whatever, you know, what other containerization platforms or the application people have their own or any and all of them together. And then it's just a popery of pain to try to, you know, to try to replace it all. Right. Yeah. There have been a couple vendors that are trying to go up and down the stack in a more complete way, you know, what I'm thinking of like Dynatrace and
and handful of others that, you know, go from the application stack all the way down into the networking stack to try and get better key insight and give you application awareness and things of that nature. Most of them require agents in order to be able to do that, right? So there's the whole debate about do you have access to actually put this agent on the system is both server and client systems to be able to figure out exactly what's going on and, you know, there's a whole slew or whole set of category of problems to go along with that. And, you know, a lack of network insight can cause also to problems upstream that look like application problems, but actually network problems and the reverse, right? So, so try to fair all that information out of the event correlated to your point is really complex. I think the challenge with V6 in these particular environments is the fact that if you're ignoring that particular network protocol, you're only you're only seeing what V4 is really doing on the network and so you have no idea, you know, to dealing with, you know, QS
and congestion control as an example, whether you're filling up your cues with V6 related traffic versus V4 related traffic and how that plays out and what sort of performance differences you get. The same thing holds true for, you know, overall bandwidth performance. If you're not measuring the V6 related traffic, why the heck are they, you know, at a later two, why am I seeing one number and a layer three, I'm seeing a vastly different number in terms of performance? Well, I mean, that's turn on V6 like where do the V4 traffic go? Right. Yeah. And that's that's a real world issue. And so you need to make sure that you're doing all the prep work and the right work in advance, but then making sure that you're captioned the data to know whether you actually move that traffic onto V6 or whether you just dropped it to the floor. So it's an important issue too. So I think those are some of the challenges that I've seen from just trying to get monitored and work flat out just outright. And then there becomes another stage of like more nuanced IPv6 monitoring problems that are that fall into the stage of like, I've got something
weird happening with IPv6 and I'm trying to get some insight from the tool about what's happening, but it's not able to give me anything because it doesn't really understand the packet payload. And so I'm doing, you know, I'm having it provide me dumps to try and do, you know, to try and look through with Wireshark and figure out step by step exactly what's going on or things of that nature because they just don't have good intelligent insight tools that are prebuilt into many the platforms unlike with V4 where it may just say like, oh, well, it looks like it's this problem, right? For you, there's less just inherent identification of common use case scenarios and issues that go on within the environment. And that's been my experience. There's just much more limited, sort of structural security, general network, misconfiguration, and stuff like that that can get caught by the platforms for V4 that just aren't necessarily available for V6. And I don't know if that's been your experience to Nick Tom, but that's been mine. And that's sort of what I feel like
is the current state is that is working on that space is going to be critical for monitoring systems. Yeah, I mean, I definitely think that that is improving, you know, better than it was prior, but the gaps are still there's still pain points for people for sure. Yeah, I don't, honestly, I don't know, I don't know this data of things for, you know, NetFlow and IPFix from a vendor space perspective, you know, but those should be relatively good for just pure data flow, right? If you're using like, you know, any of the ingestion tools, you're going to be able to see the V6 really traffic. I just don't know what's the visibility gap is really, I guess, the big issue of how do you present the right visible data for, you know, cloud environments or ISPs in general, and what they're providing from a performance. And are they mishandling, you know, extension headers or are they doing the right thing by forwarding certain traffic types to you, right? Across, across there, are they truncating, you know, extension headers
that you might need? So those are the sorts of things that like it's super useful to have a tool that bubbles that information up to. And I'm not as familiar, maybe neck you are, but I'm not as familiar with tools that are really doing a great job with that. And, you know, the whole debate about dropping the extension headers to the floor and firewalls and are you getting that data from your firewall or not? Like becomes a whole different secondary issue, right? Have, yeah. Am I getting insight to tell me that I'm actually having that thing happen on my network? Am I actually dropping portions of extension headers intentionally or unintentionally based off of default parameters that are set up within a firewall vendor or, you know, in a security device? And that's the sort of stuff that becomes more problematic because it gives inconsistent behavior of how V6 behaves. And that's, that to me is a bigger issue, right? Yeah, there's been a never ending debate about extension headers in the wild, you know, in the internet. Are they, do they exist? Are people
using them? How do we know if we're seeing them or if they're being filtered until less are sent the same questions been raised with like flow labels? Because there is some creative use of flow labels like fireflies and things like that within, you know, the scientific community, but there's no reason I couldn't be used elsewhere. I have seen, I have seen flow label usage increased dramatically over the last several years. And it's probably the last two. My experiences is not being filtered. And it is being experienced too. When it is being filtered, it's, it's a security appliance that has some default and no one knows it's happening, right? It's just like silently being dropped. And I think that's probably similar with extension headers, but, you know, I haven't done enough study to know is that accurate or not? Extension headers were, extension headers were given a pretty negative
industry, not industry standards group, but the industry was pretty negative on extension headers across the global internet. And there was a lot of work that was done, you know, sort of outside of the, the standard spot is around saying like, this is the stuff that needs to get dropped. Why should this ever make it out? Why do we need more than XYZ amount of extension headers? And that data was sort of fed back into the standards groups to try and put better boundaries, I guess, or controls on extension headers. And I feel like that work really doesn't exist for flow labels in the same way. And I don't think people feel as negatively about flow labels in terms of what they could potentially be exploited for or what, or, and their use cases more narrow, I think the extension headers extension headers by definition are, you know, are extensible. So they become something that becomes a footprint for attempt for exploitation. And I think that's what the major concern was, or at least that's my impression for, you know, sort of sitting back and watching what
was going on in the community. I don't know if you saw things differently, but that's how I saw it. No, I agree. I mean, I think it's interesting that I do see, you know, just doing random packet captures on V6 for other purposes, you know, just noticing when there's flow label data in there and wondering what's going on there. You know, if that were filtered out, I mean, you make up, you bring up a good point that is different than extension headers, which, you know, have the potential downside and vulnerability of the traffic being manipulated in certain way versus, you know, just a flow label value, which can be acted on or ignored, but not to that depth of, like, you know, really rerouting the traffic and causing issues in that regard. But I do wonder, you know, if you filter, if you have an appliance that is set by default to filter that information out that if there's like breakage happening on the back end that you're not aware of, you know, that the application is relying on that, that information. So it's kind of a mystery, you know, what are those values? What are they being used for? Well, and this gets back to my, one of my original points, which was that we don't necessarily have
parity for the capabilities of the monitoring and gestion systems to see all of this related information to bubble up in the way that's useful to you and then to provide some insight. So in the insight being like, this is how you fix it. This is who you need to go address it with. This is, you know, this is, it looks like it's, you know, maybe this appliance that's doing this, like, you're not necessarily getting all of that data in the way that you would like to get it. And I think the V4 side has much, much further in regards to, you know, sort of solving those related problems versus where we're at. And then we always end up sort of talking about data normalization with this because you, you know, you have, to Nick's point related to, you can pretty much cobble together something that's open source that's going to check all the boxes if you're pretty intrapid with how you're developing your monitoring system. If you have the freedom and the resources to do that. But then, you know, the more data sources that you have that are from disparate open source projects and you end up with like potentially different presentations of IPv6 addresses, which in V4, you don't really have the same issue. So you end up with like a lot
of different data, a lot of sets of data that, you know, from a V6 standpoint aren't normalized where you can do an Apple's to Apple's comparison on one dataset versus another just based on how it's presented and how do you get it? How do you get around that? You know, if you're, if you're using a tool like Splunk, you know, that has like some normalization built in, but is it sufficient to take in the data from all the different sources that you have it from and to normalize it in a way that makes it actionable? Well, I also think that, you know, it's not necessarily a totally fair comparison because IPv4 is so much more rudimentary compared to the header structure of IPv6. And let's be totally honest, there's not a lot of people that actually understand a lot of the nuance of the flow label and the extension headers, right? Like even within the expert community, like you go look at some of the ITF lists, like the extension headers are, I mean, I personally,
you know, I understand them enough to realize they're sort of black magic, you know, and I've been doing V6 a long time. So, and I don't totally know everything about them at all. So I think there's, you know, there's the complexity issue that there probably would never be parity simply because V4 doesn't have those attributes. And so it's not like I'm porting over a feature. I'm writing a completely new parser for this. And I think most of that, especially within like, also like for S flow, for example, it has to support IPv6 to officially be, you know, blessed as this does S flow support. Like if it doesn't do V6, it's not officially S flow. If it doesn't, if it's not in the packet structure, now it may or may not be able to source that data from a V6 address or sync it to a V6 collector. But, and I have seen that even as recently as like last week. So that's
still a thing. But, you know, there's not a lot of gear out there that still does, I mean, I'm sure there is actually someone will correct me. Gear that only does like net flow version five, which doesn't support IPv6. It's going to be version nine or 10 IP fix. And that by nature does does IPv6 export and supports that packet structure. So I think that's changed a lot and it's getting better. But it's more the the collector parsing logic may not have all those pieces or it may just be significant and more complicated to pull it out of there. Right. Or it's very diminished in the data value that it gives you. It's just like this is the extension of her, but it doesn't give you anything. Oh, yeah. So you're like, like, what does that mean? So the data insight is the is the big delta. And I agree with you, extension headers are not are simple and concept. Very complex and implementation. And this is one of the reasons actually that so many people
are so nervous about it versus a full label. You can to be blunt, you can stick whatever value you said you wanted to fully label and use it and use it for that purpose because it's really you know, it's designed for things like entropy across, you know, M lag and you know, things that need to have a data value in order to, you know, improve performance or, you know, provide a way of tracking a particular flow. And so, you know, there's there's useful information within that. And there's some hierarchy that they've worked out within the full label to help, you know, do correlation for inclined devices and things of that nature, which are pretty cool. But the reality is it's I think is much easier to understand what's going on with the full label than extension header in many ways in terms of at least the implementation side. And I don't think the industry's got consensus around extension headers. And in any really good firm established way, maybe there is on the service water side terms of how they want to deal with it and what they expect to see, you know, in terms of, you know, internet pairing and traffic
forwarding. But in terms of enterprise environments, I'm not aware of a well written document that talks about exactly how extension headers should or should not work within within enterprise. And I may just be not up to date on my RFC, you know, references and operations. But, you know, I put that as a to do list for Nick for, you know, go find someone to write that. I see two people right here. So, so I think that I think that would be, you know, super helpful. But, you know, and so maybe we sort of rounded out, I think the final thing is, you know, besides all the event correlation and data monitoring is, is really, are you gaining the insight to be able to provide reports that are useful about talking about IBV6? Because that's sort of the final, you know, lag in the stool. Of if you can't provide good insight data, and if you don't have a system that's summarizing that,
and any useful fashion for when you have to deal with it out of it, or when you have to deal with, you know, security incident, or when you have to deal with, you know, some other network or systems related problem, you know, it's great to have a monitoring system. But if you don't have something that's providing you the visibility and reporting out of it. And I think there is still a gap in the reporting side for V6 to have something useful that bubbles up from that, especially around the things like, you know, client device types, because if you have, depending on your environment, if you're doing slack, you may have, you know, permanent privacy address, you could have temporary addresses, you could have, you know, global unicast addresses, you could have VLA, obviously, have link local address. Do they all correlate all that related information along with IBV4 to a singular device that tells you, this is a device that owned it, and it may own these addresses over a certain duration of time and, and have a time mapping that gives you. And our, you know, is it correlating with whatever the address allocation method is for that particular
network segment and understand what's going on there. And I think there's a lot of missing data from a visibility standpoint that isn't taking a bunch of that logic into account. And this is just my experience of sort of looking at systems and being like, yeah, I don't necessarily trust, you know, that a system is, you know, scanning for MAC addresses makes no sense anymore. Scanning for V6 addresses makes no sense anymore because we've got MAC rotation, we've got V6 address rotation, right? So how you go about getting that information and getting the right insight and tying that to identity becomes a much bigger, bigger structural issue. And I think those are important things that need to get tackled. Yeah. And I'm going to be totally honest, that's a big impediment for a lot of organizations that have historically done MAC address to, you know, layer three address correlation for identifying a, identifying a piece of equipment. And thereby the user that's on it, right? Because it's let's say it's a laptop or whatever.
And that's been a big impediment for IGB6 is because there is not a good way to do that, especially because they typically do that with a DHCP fixed address. And as soon as you introduce a single Android device, it doesn't, it breaks down, right? Because they don't do it. Right. Even with DHCP V6 with the rotating MAC address issue, right? That still becomes, it still becomes a challenge even in the V4 world. So I think they're starting to realize that, you know, identity certificates are more important. Like there's other attributes in that, you know, MAC addresses and IP addresses, whether they're V4 or V6 are really just pieces of metadata that go along for the ride and can change over time. And I think that's, that's an important thing for folks because there maybe we can do a separate show just talking about that. We can, we can talk through that. But on the monitoring side, I would encourage folks, you know, to really do some serious investigation
around what V6 capabilities your platform is and what you can support and what you're looking at and what your need requirements are because that's obviously a big, a big part of it. And so just up down link interface stuff has been out of vogue since, you know, the mid 2000. So, so you need, you need something more sophisticated that really gives you much better insight. And so my expectations are, you know, that gets back to the original problem of, you know, future parity versus functional parity and understanding where you fit for each one of those and what you need to do. And I guess we need to do, we'll probably need to reintroduce and do another show about future parity and functional parity just to make sure that our ones up to speed on all of that. But, you know, keep an eye out. I think the monitoring scene is improving. There's been some interesting sort of work that's been done in that area. And maybe with all the AI insight, capabilities of what platforms can provide today that will get better and better data in this, in this category. And I think that's one of the things that might be super interesting to see
develop over the next year or so. Yeah, I'd like to see more of that as well. Yeah. Cool. All right, you guys, there you go. We got one wrapped up monitoring. It is. Do the monitoring with the six. And if the audience has any, your personal experience around this and recommendations for folks we'd love to hear from you, you know, just do the pack of pushers dot net slash F you for the follow up. And we'd love to hear what your experience has been with monitoring IPv6 and whether that's a still a problematic area for you or if you think you you get a pretty much nailed down the solve. Yeah, let's throw it down the gauntlet with the vendors too. Somebody has an off-the-shelf product out there that does v6 end end and provides all the actionable data that you could need. Yeah, come on the show. Come sponsor. Exactly. Thanks a lot.
More episodes
More from The Everything Feed - All Packet Pushers Pods

TNO072: Connectivity and Community with Jason Gintert
The Everything Feed - All Packet Pushers Pods

HN841: HPE Melds Apstra and Mist for Self-Driving Data Center Networks (Sponsore...
The Everything Feed - All Packet Pushers Pods

LIU022: Chris Grundemann – From Pulling Cable to Network Automation Forum
The Everything Feed - All Packet Pushers Pods

D2DO312: Networking at Scale: AWS Transit Gateway War Stories
The Everything Feed - All Packet Pushers Pods