Skip to content
TrackPodcasts
technologySep 17, 202633:44

IPB208: IPv6 Address Management Is Broken

Get every episode summarized

Each time The Everything Feed - All Packet Pushers Pods publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“Eligible new customers can get one month free on one eligible Megaport port and Megaport CloudRouter to test private connectivity and cloud routing in their own environment. Visit Megaport.com slash buzz to get started.”From the transcript
Traditional IPAM and DDI tools struggle to handle IPv6 at scale because they were originally designed around IPv4. Today Ed and Tom examine the visibility challenges posed by SLAAC  addresses, broken DNS integrations, and massive prefix allocations. They also share potential solutions and strategies for enterprise IPv6 deployment. AdSpot Sponsor: Megaport If you’re supporting cloud... Read more »

Hosts & guests

Transcript ready

481 searchable segments. Every word is indexed and playable.

IPB208: IPv6 Address Management Is Broken

The Everything Feed - All Packet Pushers Pods

0:00
33:44

Full transcript

The Everything Feed - All Packet Pushers Pods — IPB208: IPv6 Address Management Is Broken. Machine-transcribed; use the interactive transcript above to jump the player to any line.

This episode is sponsored by Megaport. Eligible new customers can get one month free on one eligible Megaport port and Megaport CloudRouter to test private connectivity and cloud routing in their own environment. Visit Megaport.com slash buzz to get started. Terms in Eligible will he apply. Welcome to the IPv6 Buzz podcast where we dare to dive into the 120-bit Adjerspace wormhole. I'm at a really Tom Koffinazom with me and Nick Baraglio is out on assignment as usual. Beautification making himself look better. It looks, it looks maxing. It's what everyone's doing. So we wanted to chat about whether you're your IPAM and your DDI tools are sort of ready for IPv6's scale. I think most folks probably are using some sort of DDI solution set or maybe doing a one-off

and just doing maybe some DNS and DHP configuration, maybe not doing the IPAM portion. It's one thing to get it up and working in a lab and just to get it functional, which is definitely a doable thing. You can absolutely get that up and working. There's a whole different set of challenges that come along with trying to do it at scale. And I guess we should define what we mean at scale. For most enterprises that are running a distributed DDI platform configuration of some type, whether that's a commercial product or they're doing it themselves sort of bake from open source, the reality is that coming up with high availability, redundancy, how do you build and redundancy, how do you deal with all the related dependencies that go on with allocations and scope management and hygiene and address tracking and all that other stuff. Actually becomes a really hard architecture and design problem to sort of figure out about

how you want to do it at scale. So we wanted to just take a little bit of time and talk our way through some of the things that we think are maybe major issues or things to consider in sort of working your way through that. So I guess we'll kick it off and start from there and really start to talk about maybe some of the first challenges that go on for everything that we sort of talk about with V4 and V6, which is sort of the scale difference. There's just a fundamental scale difference from V4 and V6. So if we talk about that first, what are the things that are different? And maybe just some of the obvious ones for things like DNS, probably not going to use origin and just pre-build all that much of stuff for or do the same thing in your DGPV6 platform. So we build to the 64th addresses in a database and figure you're going to assign them.

I think you're going to want to have that beat dynamic and not something that you set out. Handing out addresses one by one, which we're probably very familiar with doing in I PV4, is actually knowing that we can assign all the addresses in a slash 24 as a relative to easy and reasonable tasks that you could do even in a spreadsheet. You don't even need an actual tool to do that. That becomes an impossibility in the V6 world. I guess that would be the first place to start and that's where the tools make the most sense, right? In terms of actually leveraging those tools. I think there's another more fundamental problem in that you don't have kind of operational model choice in the V4 space that impacts scale and the way that you do in V6 where you have to make decisions about how you're going to allocate IPV6 addresses and you have to use the technologies that are available to allow you to do that, which are Slack and DC PV6. That environment is by no means super deterministic for all enterprises in terms of, well, just

turn this on and it'll work in the way that you've been managing your IPV4 resources and endpoints with IPV4. You have to make some operational choices that have pros and cons that mean that a lot of enterprises, I think, that we've run into in agencies, they haven't really gotten to the point where they've made firm decisions about what that operational model is going to look like. Even before you get to the scale question, you're kind of stuck in this limbo of which operational model we're going to use. Now, it's too good that we have IPV6 mostly coming on the radar here across the board with endpoints, OS endpoints, but it does, for their complicates the question, it's like which... Well, I guess I would argue it actually provides clarity to the question because it forces you in one direction if you want to use IPV6 mostly. You can only use Slack and therefore at least in the current implementations in the

RFC standards, that's the only option that you have. The only tool in your tool belt, if you want to solve that and make use of it, it forces you in a particular, I guess, design pattern in order to make that work. Now, I'm not saying you can't use Slack in DHCPV6 within your design and it may be specific in particular areas and maybe in your client access, you're using Slack just because you want to take advantage of IPV6 mostly, but you might choose in other areas. And those other areas are yet to be determined, I guess, but you might choose to do DHCPV6. I think for me, the things that get more complicated and things like data center design versus client access edge for LAN is in the data center world, the decision between Slack and DHCPV6 really is in a discussion that anyone really has, I don't think.

And partially, because of the virtualization platforms, really taking control of the address allocation, taking control of how that stuff works. And so that's taken out of your hands and you're not worried about that. And that's true in the cloud environment, it's also right. So any of the public clouds, they have this shim that sort of sits in between and you're doing allocation work from a pool of address ranges that it's managing for you on your behalf so that the automation tooling can go out and do all the work there. And so you're not necessarily doing this in as much of a direct way as maybe choosing to do Slack or DHCPV6 would be in a LAN environment for client access. Arguably, you're bringing your sense and understanding of how to manage scale with you to the whole project when you're in the data center environment. Yes, yes. In a completely different way, which leads probably to greater success in terms of scaling. But to be fair, it's also maybe a more deterministic environment.

And I think the wildcard here has been the lack of IPV6. I mean, we've really only been talking about IPV6 mostly in the V6 community for a relatively short period of time. How much is actually percolated out as a potential design option of two enterprise network managers is another question. I think that's going to change obviously over the next year as it becomes widely available in the Windows platform. And then to your point, the scale problem that it solves will become quite evident. And DDI will certainly play a role in that, although as we were talking about earlier today and settings with clients, with customers that are kind of not sure what that's going to look like in terms of how to track those resources. Because we're kind of back to a big operational jumble again with if I can't turn up the CLAT with the DDI, V6 address, which I can't. There's no support for that currently.

Then I'm stuck using Slack. And if I want that tighter control over endpoints with, you know, that I get from DDI, V4, you know, then I'm back to running two environments. And that's super impactful on scale, of course, and just the operational management question. Yeah, I think the option management question is the real key scale question there of, can you stream the data out that you required to identify endpoints and to know, you know, what's what? Now, I've always argued that, you know, addresses, whether they're MAC addresses or IP addresses and IPv4, IPv6 does a matter which should really just be pieces of metadata that are used to inform what a, you know, an end point user or client access device is doing and that, you know, having a zero trust platform or, you know, solution that, you know, provides you or stitches together all this related information is a better way of going about that. And this is just a piece of data that goes into that platform to express, you know, this is

the CLIOS. This is the user on that device. This is the device itself. It's MAC address. It's IPv4. It's a mobile IPv6 addresses, you know, all the related pieces of telemetry that you want about that end client that determines what trust platform, what trust levels they should be getting in access to and that stuff should be dynamic and on the fly. Whether you have a stateless auto configured address or a DHCPV6 address really probably doesn't matter in those scenarios. If you're in the right architecture, the reality is is no one's that I've seen is really there yet. Right. In terms of deploying or building any of those sort of solutions at scale or not at scale. And so that becomes a whole different, you know, conversation of like, you know, what do I trust from an information basis and do I have the right tools to get the information out of, you know, my network devices to get all the slack information so that I can identify

my end points and what, you know, am I enabling them to have, you know, temporary versus privacy addresses, you know, how does that work? Do I have managed client devices in my network or do I have all unmanaged devices and IoT related devices? And that becomes the whole, you know, do I trust the end client for doing registration for things like dynamic DNS and reverse DNS? And you know, if I've got manage devices, yes, I want to trust them because I want to be able to reach out to them to manage them. And so I need them to register there, you know, their host name, you know, write a QA record to their IPP6 address, they're going to make use of so that things like, you know, whatever, Hector Director, you can reach out and push a group policy down to that particular client device right when it comes on the network. Things of that nature and that looks different than like the purpose and use of like temporary addresses and Slack. And whether I register my Slack address or whether I register my DGPP6 address, I don't

know if either one of those matter, right? In terms of the entry, I just need to get to the device itself. So the real question is, is there operational models that support doing all that? And today I would say no, it's just to be blunt. You know, if you, if there's some RFCs that have been written to talk about, like, you know, I got a Slack address and I want to go, you know, notify DGPP6 that this is the address. I'm using. And so there's a RFC that has been written that does that, but I'm aware of no client and no DGPB6 server that actually works with that particular RFC. If there's experimental code, I'm sure there's experimental code because most of the time for the RFCs there are, there are, but I'm not aware of any commercial client packages that support that. And so yeah. And so you're stuck with, okay, well, I've got to figure out this Slack correlated information to it and client device, you know, MAC address or, you know, something else that I'm identifying

with and where do I get that data from? And that comes back to the pre-sup episode that we did about logging into Lomitry and Visibility, right? And can I stream that data off of my networking platform to give me those sorts of insights and do I have a platform that I'm ingesting in and that can then bubble that information out in a report and some sort of some sort of report, real-time report or close to real-time report that gives me that information that I can then feed into, you know, whatever, you know, security related tool sets or whatever compliance reporting tool sets that I need to, you know, sort of have on hand in order to, you know, you know, deal with the issue of like, do I know who this person is on what computer or device with what OS with what IP address is with what MAC address is right? Can I get something that mine's all that up for me? And that's problematic. And today, I'm not aware of any of the DDI platforms providing that set of glue in any sort of comprehensive way. I think there's pieces and parts that exist in these platforms and these solutions, but

I don't believe anyone's got a real good end-to-end solution offering that provides this sort of insight. Now, I could be wrong. And if I am wrong and you're one of the vendors, you know, reach out and come sponsor a show and tell us how you're awesome. Absolutely. Come on the show. Let us know how it all works. How it all fits together. Yeah. Yeah, for sure. But I'm not aware of anyone that has that sort of integration and even correlation information between the dual stack environment of V4 and V6 and matching all that stuff up together in even the IPM tooling to understand and correlate that information. And I think that's a real problem because how do you deploy at scale and figure out where you're getting exploited or where you have operational issues or why you're having log in authentication problems like all these things become exponentially harder or had scale with the address base that you have, but also with the distributed portion of building

on these platforms. And I don't think anyone really has a great solution for at least for DHPV6. You know, high availability in DHPV6 is not really super pretty. Right? You know, the split scope instead of a scope for data to be. Yeah, split scope instead of redundancy is not a great answer. We used to have that in V4 and everyone hated it, right? No one loved doing that. And so, you know, for some reason we decided to carry that no love forward. And again, keep the same design, you know, early design model. And I have no idea why. It's really just a horrible solution in terms of split scoping idea. To be fair, I mean, there had been efforts in the ITF to define what that looks like. To have DHPV6 redundancy and proper failover. We haven't seen a whole lot of implementation of it. I, it's been a while since I've looked at Kia to see what's what's there. It's in terms of like support for the equivalent of DHPV failover and IPV4.

But the larger point is that, you know, if you're just, you're used to buying commercial, up to shelf stuff and you've got a DDI solution. And, you know, it's like I want to be able to check the box that DHPV, DHPV4 failover is going to, it's going to, V6 failover is going to work in the same way that V4 failover works today. There it's not that trivial and not that easy to do. And so then, if you add yet another wrinkle to the pile of operational models that you have to sort through in terms of, you know, making sure that. Yeah. Yeah, we get further and further away from having operational consistency between V6, right, in terms of how that plays out. So just, it's just a frustration more than a technical hurdle, I guess. You know, it's, it's workable. It's just not particularly elegant and it doesn't match the models that's that folks want from a V4 perspective. Right. You expect that stuff to be solved at V6. Why, why are we carrying over older older solutions? This episode is brought to you by Megaport. Testing a new cloud or network design,

Megaport's private connectivity and cloud routing could be provisioned on demand across cloud and data centers. Right now, eligible new Megaport customers could get one month free on one eligible Megaport port and one eligible Megaport cloud router. Visit Megaport.com slash buzz to get started, turns an eligible to apply. And then I guess, I mean, we talk a little bit about, you know, just the operational reality of the size and scope of address space because we sort of let off with that a little bit, we didn't really dive into sort of talking about it all that much. I mean, how do you even plan around that? Right. If you're an operator before, there's actually some usefulness of the constraints that IPV4 puts on you, right? Before it forces a set of discussions that declare like, I'm going to use this address to do this function or I'm going to use these sets of addresses to do this scope of function or I'm going to use this entire subnet to, you know, turn up these related client devices.

And in V6, we really can't have that same conversation, right? It's really all about prefix allocations for particular functions or services that you want to deploy. You don't really have the like, I want to assign this address to do this function sort of thing. You can, but, you know, it's pretty limited in terms of what you want to do. You might want to do that on your server, a little balancer. You might want to do that on your routers for things like router IDs and bootbacks and stuff like that. But outside of that, I can't think of any reason that you really care what the address is on the client device itself, just let it get an address and go. As long as it's registered in DNS, who cares, right? Yeah. The scale, and then you get back to the issue of like V6 scale when you're talking about this scenario and that, you know, I can, in V4 land, if I want to do any CAS as an example, I have to have a small prefix, a small prefix I can have is a slash 24, which is pretty small, you know, from an address utilization standpoint. And then I'm going to get in trouble here with my own standards by talking about prefixes

in terms of size in relation to host, host addresses consume because we, you know, we were always warning against that for good reason on IPv6. But in V4, you know, you don't break out into a cult sweat thinking, well, I've got to set aside a 20 publicly routable 24 to be able to do, you know, when any cast on like publicly, if it's the same thing in V6, I need a slash 48. And so, you know, already it's like, oh, so I'm going to use 65,000 slash 60, the equivalent of 65,000 slash 64 subnets for potentially a single service address to do DNS any cast or some, you know, some service any cast. And that's kind of a hard hump to get over initially if you're not used to how we use IPv6 address space, you know, typically. And then, yeah. And then it becomes the problem with how do I report this in my DDI platform, right? Like how do I express this information?

Like, you know, when you look at it on a simple architecture diagram and say, like, okay, well, this address is servicing that role in V6, right? And you're able to maybe make use of zero compression, right? And, you know, double colon up and make it nice and short to have a efficient maybe DNS, you know, any cast DNS name server that runs both either publicly or even internally for your corporate environment, you're burning a whole slash 48 to accomplish that in many situations. And so, you know, is that a concern from a, you know, and can your platform describe that particular attribute of what you're doing there without making it look like you just burned, you know, buildings and buildings of addresses, which is not a, which is not a comparable operational, you know, you know, feature or function that you want to necessarily have displayed that way, right? Like, it circles back around to just a general point about DDI and IPM in particular related

to V6 and how that information is presented. And more importantly, what that, what how the information is presented implies in terms of how you should be thinking about the resource. And so to that, to that exact point, when you talk about having, you know, potentially burning billions of hosted dresses, you know, I'm just, I'm just throwing away, you know, a billion hosted dresses, which, you know, of course, I'm doing that if I assign a slash 64 anywhere I'm doing that, I'm doing the same thing. So, the point, the point being that the IPM platform hasn't really incorporated the scale of IPv6 and the way that it visually presents the data in a way that's meaningful and actionable to most architects and most engineers. And again, it's, it's as we've said many times before, it's not, it's not for lack of trying, per se, it's that it's such a paradigm shift in terms of the scale and involved that it's just a really difficult sort of problem to solve and to make digestible for

enterprises for enterprise and network architects and engineers. Yeah, it's both, yeah, it's both a UI problem from a display and it's a functional operational problem in terms of it's, it's both and, and they reinforce each other. They reinforce each other. There's, there's not a simple solution of saying like look at this, you know, graphal depiction to simplify what's going on over here from a operational basis that you have conception of what's going on. Right. That hasn't matched up. There hasn't been a tooling set that I at least I've seen outside of maybe some of the stuff that we've built to help organizations figure that out in a better way. And it's, it's, it's very, very problematic because it has some certain, um, limiting factors in terms of, you know, scale architecture because if you're doing this consistently across the board for site configurations, it can look like you're incredibly wasteful from an address perspective.

If the, if the, if the DDI platform is really presenting everything from an address basis versus a prefix basis and then is not making the distinction of, you know, particular Haggerst architecture, you know, design decisions that were placed and highlighting what those are within the architecture itself. If you can't see that, then it can be very problematic in terms of understanding how you're going to scale out a particular design. And, and, and one of the problems that goes along with V6 is, um, from a, an operational model is like, where do I send my request for DHCP V6? Where do I send my DNS related information? What hierarchy am I supposed to be using? Should it be matching what I'm doing with IPv4? And as with everything, it's complicated. It really depends on what you're trying to solve for and then what has dependencies on those things. So things like, you know, Microsoft SAC to directory and you're tying that in.

Do I need to pin certain things in certain locations in order for that to be consistent? Am I building a, you know, much more, you know, distributed regional, you know, architecture hub and spoke more, more, more around that design or is everything distributed or is everything sd when and everything's connected back to a centralized data center and this whole regional ized concept goes out the window and I'm not really worried about that. Or am I running these key sets of services in cloud and I'm really doing local internet hop off and, and I've got direct access for these resources. So there's, there's so many different things to sort of think your way through. And in IPv4 with the use of network address translation, some of these complexities get hidden or you don't have the tooling to necessarily, you know, have the same insights about what's going on because you're, you're just doing the translation and, and, and assuming that that's a boundary set, both from a security dynamic, but also from an architecture dynamic that, that stays consistent across your, your design.

In MV6, that doesn't necessarily have to be the case, right? We don't have that constraint. And so it becomes a little bit more complex to understand the scale of like, you know, who owns what portions of my address space to manage from an organizational standpoint? Should I be, should I be allocating these addresses to be run in a minister locally by a local team to do things, right? Which I probably would not do in IPv4 at all. No, well, yeah, and you can, you can basically, you can and you properly are stuck front loading capacity management related to available prefixes like in the sense that you're allocating of, you know, if you're doing your address plan, right, and you have the right size allocation, you're allocating up front in a way that, that limits or prevents the, the standard practice in V4 of having to go back and get additional prefixes because it's a space constraint and glue them together from together. Yeah, try to, you know, number out of some resources to, in order to create some contiguous

subnets, all that fun stuff that we get to do with V4, you know, in V6, you know, the idea that, okay, it's not just, it's not the concept of, I'm going to put a slash 64 in an interface and never have to think about it again in terms of host addresses. That also applies at the, at the prefix level, you know, for like regional breakout for things like routing summarization and, and regionalization of routing and, and potentially security boundaries and things like that, where, you know, if I allocate initially a slash 28, you know, for an entire section of the network, where it just completely does away with the, the need for capacity planning for, for future prefix use, you still need to obviously a process to, to track what's been allocated and, and where it's being used, but the fact that you've front loaded the allocation, you know, of all the prefixes that are available, you know, you're never going to have to go back and get more in the way that you do in V4. And so again, now we're back to, you know, completely different operational model that, that, that, that more we have in V4.

Yeah, yeah, I agree. And then as you mentioned, like the tracking and the architecture, become, become really tightly coupled, right? In terms of understanding where resources live and, and what you're signing out. And I guess the, you know, sort of mentioned it at the, at the top of the show, that the, you know, the invisibility in some ways of, of stateless address auto configuration slack, not providing the updated information into that system. So now you have, you know, at scale, and this becomes very problematic, right? Where do you write to to, to get this information, you suddenly have to build an at scale system that can pull data from all the layer three platforms that you run in your network in order to, you know, get that visibility. And is that a scale problem that you're ready to solve and can your devices handle that and can you ingest that sort of data to sort of figure that out? And that's, that's problematic because slack scout, you know, whole set of, you know, both temporary and privacy, you know, time out thresholds that, you know, can invalidate

addresses very quickly depending on what you're doing, like some of the mobile platforms, right? So it's, that, that becomes a whole challenge of, can my DDI platform scale to, you know, ingest and provide me that information or is it even capable of providing that information? I would say today, it's probably not capable of providing that information to you. If you are running slack, you're going to have to find a different method to go query and get that data, right? Yeah. There's nothing there. And so it's almost self-defeated because, you know, the high and the IP, you know, IP address management, that management part becomes less and less management. I guess because you really don't have the information in there. So I think that's a real problem for the industry overall to start working on. And I think that's a unique challenge, maybe, for how enterprises are going to deal with stuff. I don't know if I'm concerned with that for something I'm concerned about for home,

something I'm concerned about for service providers, but certainly for enterprises, you know, of any size, this becomes a bigger, bigger, bigger issue. So I think those are just some of the things from a, you know, from a IPAM and DDI, you know, sort of, you know, platform basis that you need to be thinking about carefully and trying to really work through where your scale problems are going to be and which ones you want to own, which ones you want to farm off on other people. I don't know if you have any sort of, you know, closing thoughts about, closing thoughts of famous, famous last words. Yeah. You know, I think with, with where DDI is at in general in terms of what vendors are offering today, you know, there's obviously a tremendous amount of IPv6 support. And as we talked about, you know, the bigger challenge oftentimes is which operational model are we going to use? And then only through making that selection do you tease out where the, where the deficiencies

are in a way that may prevent you from establishing, you know, a robust practice that, you know, relying on things like automation based on having that operational model, like I'm very clearly defined. So and I think I've always believed, hope springs eternal. I've always believed that the vendors, the DDI vendors will catch up to the brave new world of IPv6 at scale deployment. But, but in their defense, you know, it is, it's been such a long strange trip, just, you know, trying to get to the point where we even recognize that something like IPv6 mostly is, is, you know, a potential, you know, lever in a way to, to accommodate IPv6 adoption at scale that, you know, we really haven't, haven't had. So it's, it's not entirely their fault in terms of that fact. Yeah, it's, it's a demand side issue, right? And if you don't get the demand, you're not going to build the, the product support until you, you receive sufficient demand for that particular need.

And so, you know, they're all product companies here. They're, they're all, they're all profit motivated, just like anyone else. It's going up. It's been, yeah, so they're going to spend the resources and, and the energy and time on the things that customers are requesting. But I do, but I do think, you know, IPv6 mostly is relatively new for my technology basis. I think the rapid adoption of it will definitely change the landscape of the available services. And the scale problem will remain. I don't think that's going to go away. I think that still falls in the realm of, of architects and designers to really figure out exactly how they want to, how and what they want to support within their environment. And then making careful decisions about the high availability, the redundancy, the, the distribution, what sets of services and, and, and what sets of, of technologies they need to support within their environment will really make that determination. And though there'll probably be a couple of standards sets of, of reference architectures, how about that that folks will use to, well, to solve some of these problems.

Yeah. And then that's, that's something that would be nice to see from, from DDI vendors. You know, and that's something that's, there's a dearth of because of exactly the factors that we've talked about. But, you know, it's helpful when you're, when your DDI vendor can publish a reference architecture that, that, that ties all these disparate parts together and, and, you know, gives the, somebody that's, you know, really struggling in the enterprise at a level where they don't have a lot of authority, but they, you know, they have access to the platform. They have a lot of job roles and responsibilities that they're, they're trying to accomplish in having that reference architecture to work from, would be, would be really helpful. So I think, I think we'll get there. Yeah. Not tomorrow. I agree. Yeah. I agree. And I would love to hear from users. If you, if you're in an enterprise deployment with V6, how are you solving your DDI sort of scale out? How are you getting your telemachia? Are you getting your information? How are you tracking? How do you know the accuracy of the things you're tracking? Love to be able to hear from you. So, you know, hit us up at packerspushers.net slash FU for a follow up. And we love to hear what you have to say about it.

More episodes

More from The Everything Feed - All Packet Pushers Pods

View all episodes →