
Get every episode summarized
Each time Late Night Linux Family All Episodes publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
“Thanks for choosing to listen to this late night Linux family podcast. We're only able to do this because of the people who support us. Learn more at latenightlinux.com slash support or support us by telling a friend or colleague about everything we do.”From the transcript
How we approach the development and testing of our infrastructure as code, and how it differs at the homelab and professional scales. Plus our Ansible best practices.
Support us on patreon and get an ad-free RSS feed with some early episodes

Subscribe to the RSS feed.
Get every episode summarized
Each time Late Night Linux Family All Episodes publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
443 searchable segments. Every word is indexed and playable.
Full transcript
Late Night Linux Family All Episodes — Hybrid Cloud Show – Episode 65. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Thanks for choosing to listen to this late night Linux family podcast. We're only able to do this because of the people who support us. None of our show is a pay world, but our Patreon supporters can get most episodes early and never hear any ads. Learn more at latenightlinux.com slash support or support us by telling a friend or colleague about everything we do. Hybrid Cloud Show Episode 65 I'm Aaron I'm Gary I'm Sean And I'm Shane Andrew sent in an email that I think could be a broader topic of conversation for us, but let's start with the email. He says, I'm looking to convert my home lab to infrastructure as code and configuration as code using get-ops to manage changes. I've tried this a few times, but I always get stuck on the development slash testing workflow. I'm not an expert, and while I'm learning I end up with lots of test commits or changes that don't work, then those commits are pushed at any CI, CD, pipeline, fails, and my get-history
becomes cluttered. How do people typically handle this? Do you test changes locally before committing? Do you maintain a separate test environment that mirrors your home lab production environment? Or is there another workflow that keeps your get-history clean while still allowing you to experiment and make mistakes? Get cleanliness is so important, and I'm glad that he's thinking about this because I definitely don't remember what I did when I was experimenting a week ago with something other technology. I do appreciate and understand the problem he's trying to solve. Learning get and learning how to do the correct get workflow appropriate, I think is step number one. Learning how to squash all of your merges, that way you have one single commit that you can adequately describe is a good path. Branching, of course, is always a good strategy. Create a branch for this thing you're trying to do, do all the commits on that, and then once you've got it solidified and past testing, then you can squash all the merge together and then merge it into your branch. As far as separate test environments, it really depends on what your use case is.
If you have a home prod and you really cannot mess with that, then maybe it's worth spinning up God forbid, I say this. Another Kubernetes cluster for you to test in first, but usually no, it doesn't require that much. I have a slightly different approach to you, Sean. My approach has been to keep my infrastructure relatively simple and then I effectively have the ability to just spin up a new VM from a template and run the IAC locally against the new VM. So I've got an app server template that contains all of the roles and applications and things that I need to run my containerized applications. There is no state in that playbook. So if I want to go and test that playbook, what I actually do is I've got an Ubuntu VM template in my Proxima's cluster. I create a VM from that template. I run my IAC against it and then I've been that VM off. And once I know that it's run successfully against that, there's a high likelihood it's
going to just work against my home production boxes. Are you using the same VM between changes? So you make a change to your role, you apply it, it doesn't work, you make another change, you apply the same VM or you throwing away the VM and creating new one every time. No, I throw away the VM and create new one every time. I mean, it's a template, right? And it's using differencing disks. So creating a new VM is just right click, create from template and it's maybe five seconds before a new Ubuntu machine is booted even on the modest hardware that I've got. Interestingly, I've got a completely different perspective when I start to see my coming. So I think I think firstly shot that was a great shout out on branching. I think if you're getting comfortable to get and you don't want to miss a history, if you rather than think about commits as a checkpoint, think about your branch as like a sandbox where you stuff all your commits in there, do what you need to do, get it working and then and then merge something that's much more clean so you can reason about your history pretty well. I kind of think of all of this like layers.
I'm always thinking about like developer productivity. And so what you describe when you say, I'm going to push something to see eye and then see something break is to me really expensive. So what are the like cheaper things you can do faster on your local machine? And one of the things I would suggest is looking into like pre-cumid hooks. So before you even do a git commit, like do things like lint everything, whatever I'm not sure what I see you're using, but like things like cube linter, yaml lint, cube conform is another one. I'm sure whatever I see you're using, there's a linter for that. You could also like dry run against real estate. So like I think flux diff might be able to do something to like dry run against your actual cluster. The other thing you could do if you didn't want to spin up another clusters like use kind, I did experiment with that a little bit to like so kind is like a local Kubernetes cluster. So you could spin up like an emulated environment locally. These are all really cheap ways to get feedback. So imagine you've got your branch, you've got your commits there, you're testing your sandbox and you're testing your sandbox locally.
So you're building up confidence that this thing when you push the CI is less likely to be broken. So that expensive thing is for when surprising things are going to break, not static checks effectively. So I kind of think about layers and I try to challenge myself to be like if something is broken in CI, how can I shift that closer to my local to validate that sooner. So I'm looking into something called dagger to do that right now. But again, I'm no expert in dagger. I'm just learning about it. I'll just really shift and closer to local to where you're making your changes. And I think you're right. And there is a cost there not only in terms of developer productivity, right? Because I don't want all of the checks in by CI pipeline to run and then for it to try and deploy and then see it fail and then have to grab through the logs in my CI web UI. So that's expensive in terms of time. But also if you're doing this in the real world outside of your home, there's quite often a cost associated with that. I'm taking up time on a runner that someone could be using to build code or I've got a Python function that spins up in a lambda somewhere that then means that there is an actual
pounds and pants cost associated with every commit that hits the CI platform. Say there's multiple different cost facets there. So I think in a production environment, everything you've said makes complete scenes. And there's a little bit of attention there in the sense that your CI CD is often a perfect replica of what you want to deploy out to production. And so you want everything tested there. But if you can find the problems as part of your commit workflow or whatever your dev workflow and save that CI run, that makes perfect scenes, right? In a home lab environment, I'm just sort of thinking through my home lab and I'm trying to think about this like CI CD seems an awfully flattering way of talking about testing what for me is extremely simple infrastructure. So what would be an example you can think of where the sort of CI CD here makes scenes?
I don't think I'm my person asked that question considering I have a test Kubernetes cluster that also mirrors my production one and they're both run by our CD and I have branching strategies around them. So probably not the person to ask around that. But I do understand because like what about Ansible? Ansible is a great example. I can test I see really easy to Kubernetes, but with Ansible, if I'm just saying, oh, I want to add an extra step to this role, then how do I test that in CI CD? I have to spin up a whole entire VM programmatically and then run it against it and then how do I even validate that? That's a lot for home lab. Well, yeah, near your home lab environment, I think you're much more likely to do something like I'm doing, right? Which is spin up a new VM from a template or test it or clone your existing production VM, run your changes there and then test it. I mean, I don't have a CI CD platform that I'm running at home, right? Even when I'm deploying stuff for real, I'm just running Ansible Playbook, dashiinfantry.innie from my machine.
Well, it's from a jump box, but yeah, it's a difference. So it's more a case, I think, of working out the steps that you take to get there in a home lab versus the outcome actually being different, because although I'm not spinning up a new machine with Terraform running the Ansible against it, running a bunch of validation hooks to look at the state and then tearing it down as part of the CI CD run, I am doing a small part of that. I'm just automating the bit that is the bit they actually want to automate because right click clone of VM is not that difficult and takes a really long time to template out. Yes, I hit this recently because there's existing Ansible roles for nebula out there on the internet, but the existing ones I found didn't meet all of my requirements. The CA had to be the same as the lighthouse, which for me is a security model I wasn't happy with. There were some missing pieces around groups, which is quite powerful in nebula, because
you can then use effectively network security groups for the way that you manage firewall rules. Neither of those were in the sort of leading Ansible roles that were out there on the internet. So I was looking to extend one of those Ansible roles on GitHub to be able to do those additional pieces I needed. And I was in that situation where I was looking at it and I thought, well, I know it works for me, but in order to actually test this, and to test that I haven't regressed the default behavior for anybody, I'm going to have to spin up a whole other nebula cluster that doesn't interfere with my existing kind of prod environment, which is my home lab right. And to do that effectively, I had a separate terraform that had to launch a new machine on the node that has a public IP, as well as some LXD machines on my server in order to have a reasonable setup to test that. So it is a bit harder, I think, once you're talking about Ansible, than it is with something
that's just inside a container or whatever. Yeah, because inside a container, right, your deployment unit is the namespace or whatever in general. I've got a bunch of pods that you see in this namespace, and that's my deployment unit. If you're talking about VMs, that's much harder to do, I think, because you are working at a much bigger level in terms of what you've got to replicate. And there's a bigger blast radius there, I suppose, as well, right. In your case, if you applied that change to your nebula cluster and it went wrong, you're in for a really hard time by picking that or restoring it from backups. Whereas if Sean goes and deploys a change, even straight to production on his Kubernetes cluster, he has the history of all the manifests and stuff for that and can revert back quite easily. Yeah, exactly. I've got a couple of machines that I only access over nebula, and it's pretty nasty if I lose access to them. So even things like updating the nebula that's running on those machines, you have to be very careful about the way that you restart that service so that it doesn't disconnect the
SSH tunnel in the process. Yeah, it's the same with the machines that I've got, and the access via tail scale. If I mess that up, it's a 20-mile drive with a keyboard mouse to climb up into an organ loft with a keyboard. That's not going to happen. I'm going to plug Boot C again because this is where that comes in really handy with your VMs where you can run your Docker file and then test it live in a pipeline relatively cheaply on any say, steep platform without any kind of special configurations. Validate that your files and your configurations are all there and then push to your registry if it's good. And you can do that before it lands on your action machine and messes up your nebula configurations like that. So it took me through that a bit. I'm sorry I'm running a bit slow. So I understand how your Boot C would give you the ability to effectively test a VM where that's otherwise quite hard and a lot of CI, CD pipelines because they expect containers.
But here the sort of challenge is a bit more network-wide. So Aaron, what are you actually wanting to test? Are you just trying to make sure that it can connect to your lighthouse or are you trying to see if it has a particular IP? What are you looking for? Yes. So in this case, it's a little bit specific. And the sense that you have, you want to check that you end up with a functioning network between those sort of three hosts where you have a lighthouse, you have two separate machines and you want to be able to test that the network security groups or the firewalling rules between them are doing what you expect. And that I haven't broken anything with the changes I've made to group membership or the changes that I've made around where the CA lives in the Nebula Network because in Nebula you do all the certificate management yourself effectively. And so that's very powerful because you can do that effectively offline and then nothing that is exposed to the internet has the CA keys that would cause a problem.
The downside of that is that when you start automating it with infrastructures code, then you need to effectively have that somewhere where it can do the signing of those certificates and rotate the certificates in a sensible place. The existing Nebula roles that I was looking at, they assumed that was on the lighthouse. Now in my case, and I expect many other cases, the lighthouse is the most exposed machine on the network because it's the one that's got the public IP. So that's quite a dangerous place to leave your CA key in my opinion anyway. And so I much prefer to have my CA key on the same thing that's running my Ansible Playbooks. And there you have to just reject things a bit to delegate that task to the right machine. But I don't want to break anyone who's done it the other way. Sure. So how about this? What if you add your CA as a CI secret? And then when you build your image for these machines, you inject that VSCI into your actual image itself. And then before you push it to the registry, after you've built the image, right, then
you actually spin up the instance itself, which has a secret and everything and the configurations already inside of it. And there you validate, do I connect to my lighthouse correctly? RICAs, correct to configure, am I getting any kind of network or SSL or handshake issues? You could test all that there in the pipeline. It'd be a live image. Yes, I think in this particular case, it's not going to work perfectly just because the CA is separate to the search of the machine. But yeah, I see what you mean. So you end up with a machine that is very much the thing that you would run in production that you can test, which is the helpful piece. OK, that makes sense. Glad someone understood. I'm glad I came through. Let's see where we can go with this because I'm relatively new to answering out my environments. And so I'm keen to hear from all of you who have used Ansible a lot more than me to understand what the sort of gotchas are and best practices that you've found using that
in production or even home lab production over time. What should I be doing or what should Andrew be doing when you're first using Ansible to change all those pieces that you've done manually in the past or you've scripted manually in the past when you convert over to doing Ansible for that? So I think the first step for me is think logically through the different roles that your infrastructure has and what they have in common. And this is something that I didn't do at first when I first started looking at Ansible where I just had these huge monolithic playbooks, which would do everything from run, apt update all the way through to second permission of files, installing applications, copying configs over. None of the configs were particularly well templated out either. So there was a lot of duplication of config. So thinking logically through that right is there a set of packages that I want all of my machines to have. When he every machine has a top on because that doesn't come out of the box, every machine
has W get on, every machine has tail scale on, which means every machine needs to have the tail scale repurve and signing key and everything else. So that's my baseline right. Every server I touch, I wanted to have those tools on screen is another one. Then what are the logical building blocks for servers on top of that? Are there a set of mounts that all of my machines have? So then I've got a mounts role. And then what are the different roles of servers I have for me? I've got reverse proxies. I've got application servers. And I've got tail scales, exit nodes or subnet readers. Those are three separate things again. But what I have in the tail scale subnet router playbook that also contains all of the code to install tail scale, install H top copy my SSH keys over. There is a separate roles for all of these things. I think thinking logically through what that looks like is a good first step. And then once you've got that, learn a little bit of ginger template for things like your
engine X configs because that's really powerful. It then means you can do things like passing an array of websites. Say if I were making a server for Joe, I might pass in 2.5 out of pins.com, late. Linux.com, Linux after dark.net as my variables for the websites, I then consume those in my template and it creates three files in sites available. That kind of thing is really powerful. It's going to save you a bunch of work in the long term. Say those are my kind of two top tips for someone getting started with this, I think. It may be tricky when you start out to be thinking in this way because you probably got to put your shell scripts that are running stuff. But if you can have in your mind the aim of like being item potent. So rather than just wrapping up your existing shell scripts into Ansible Playbooks trying to turn those into modules, it's a bit of a faff. Like you're going to have to change your world a little bit, but it will pay back. You're getting closer to the idea of what IAC is kind of trying to be, which is this like
on your data pipeline, it runs something goes awry, you rerun it and nothing's changed. Like it's, it's starting to declare state. You're declaring the state of your system. You're not running stuff that's making changes. And one of the tips I have as well is like if you want to be confident that what you're doing in Ansible isn't changing, I think it's just run the Playbook multiple times. And if you see changes zero, then it means that you've got an item potent playbook, which I've found just has unerqued pain that would have been applied to live systems much faster than shell scripts, which will probably only run on the live system and will then break in surprising and wonderful ways. So I wish I'd invested in item potency sooner, but it's definitely worth having in your mind as a long term aim. You can do things as well like just something called molecule, which are a member using an Ansible, which lets you like test your Playbooks like in CI effectively. And I think that depended on being quite item potent as well on how you, how you define your Playbooks. Yeah. So to give an example of that, I came across recently, I mentioned earlier in the show
about spinning the VMs up from a template. And of course, one of the things that I want to do when I spin the VMs up from a template is to change the host keys, change the machine ID, clean up things like cloud in its state that has been written to the disk. It would be really easy for me to write a Playbook that just goes and creates new host keys, empty cell, edsie machine ID, cleaned, valid, cloud data out. But then of course, the second time I run the Playbook, it's going to go, oh, you have something in edsie machine ID, I'm going to wipe that out and change my machine ID again, or it's going to create me a whole new set of host keys, which I don't necessarily want. So you've got to think about how you make things like that item potent. For me, it was something as simple as writing out a file saying templatized equals done or something. And then I just read that file in. I know that my Ansible is run, it's done the template cleanup stuff. And therefore it skips over all of those other roles.
But that's the sort of thing that can trip you up really easily. And cause you a real headache if you suddenly go and run this on your production infrastructure and your SSH host keys change and things like that, or your cloud in it data suddenly gets written back to the disk and the machine does a whole bunch of stuff that you weren't expecting it to do at the next review. I really like that simple example actually. This does not need to be fancy. I've suggested yes, go use modules everywhere, but actually just as simple as like a file to check something to make sure it doesn't rerun. So it's a great show. Yeah, it's like this file exists on disk. Therefore, I know I've completed this part of the playbook. Now, is that perfect? No, because it might fail those and then still get to the next step where it writes the file. But then if it fails, those I've got bigger problems because I think I've a bunch of machines with the same machine ID and the same SSH host ID, which is a bigger problem to have. I'm going to add to not sleep on the handlers because in a shell script, you have like, oh, I changed this at seek configuration and now I'm going to do like, you know, system D reload.
And then reason why service. But if you have to keep scripting that out and you translate it one for one to a playbook, you're going to be causing unnecessary downtime because if you don't change the configuration, then you don't need to restart your service and everything and have some downtime. What you can do instead is set those up as a handlers for the file that you're changing at sea and then that says, only if I ever make a change to this file at sea, do I want to do these things, right? And that would be reload and then restart the service. Handlers are a great, I would say underserved tool. Yeah. And now what I would say is if you have never written a line of answer in your life, starting as out as if it is a shell script is not the worst way to start in order to get used to the syntax. I used to have a bunch of shell scripts that would go and set up machines in the pre-ansible days, right? So I would regularly have to set up HAProxy service to do like balancing. And I had a bunch of shell scripts which would do things like run an app update, do app
install HAProxy, go and get pull all of the HAProxy configs from a get repo, put those in, restart the HAProxy service, moving that directly to an answerable playbook results in a really long, not id-impotent, really inefficient answerable playbook. But what it does mean is you get used to how answerable works. And over time you can say, okay, well, I've got a stage here, which says service HAProxy state restarted. I could probably move that to handler because I don't need to restart HAProxy every single time I run this playbook. I only need to restart HAProxy if the config files have changed. So I think there is value in if you have a bunch of shell scripts to configure your infrastructure, port those straight over to Ansible, get used to the way the answerable works and then improve it from there. And that might be a multi-year project for you because when new LTS comes out or a new version of Fedora comes out and you're going to rebuild your machine and you're going to want to tweak and change something every time or every time you run an update and you realize
it's overwritten all of my bloody HAProxy configs again. That's going to be annoying, right? But you will learn to check and validate this stuff and perhaps compare the dates on disk with the dates that I've just pulled from the get repo. Are they newer? Probably bomb out and fail and tell me. But all of these things that you would only learn by getting started somewhere. Yes, I remember the first time I started looking at Ansible and I was noticing that everything had effectively attest to see if what you had tried to run had succeeded. And it was an interesting way to start thinking about what were historically for me, bash scripts and Python scripts and things that just procedurally made changes to think, okay, how would I actually know if this change that I intended to make had successfully been made and it is a bit of a shift in thinking on some of the stuff? Yeah, and I don't have to validate that that change has been made, right? So in my example of pulling down some HAProxy configs, I would have to then run an HAProxy
test config, I can't remember what the command was now, but I'd have to test that config manually in my shell script before I restarted the service, whereas in Ansible, I could write something to handle that for me or it would just do it automatically before restarting the service. So I think you're right. There is a way of thinking that you only start to get once you've started deploying stuff in the kind of IAC way. I also found that I learnt an awful lot from looking at existing roles that people had put out there that obviously were people who knew an awful lot more about Ansible than I did. And so when I'm looking at the nebula role or whatever, then I see, oh, I understand you need to have these things in a good role and that helped me to understand what broadly was expected in the Ansible world. And there's so many, so many roles out there, finially everything. And sometimes I've found you need to change the way you're doing things a little bit,
but in some cases that's actually better than the way you were doing it because in my case at least there's reasons I've done things that aren't particularly sound. They were just how I did them at the time. And when I go back and think a bit more holistically about how I want to set up my infrastructure, then that maps a little bit more to how these roles have been set up and these best practice ones that are out there. Yeah, I think now I sort of sit back and say, what do I want this server to do? And then I go back and see if there's an existing role for it. And if there's something from upstream, even better, because they're going to maintain that and hopefully update it every time they change something. I don't think I've ever seen an upstream role. Are those a thing? Those are definitely a thing. I'm with it to my favourite projects. That's so cool. I don't think I've on off the top of my head now, but I've certainly had software vendors provide a role for I think that they wanted to deploy in the enterprise space. Is almost certainly a thing because I think I and I can remember because when it broke it was very frustrating.
But yeah, I think you're earning it as a thing. All right. Well, that's given me and maybe also Andrew and some of our listeners some things to go away and look at and our answer will repose. So thank you very much, everybody. We will be back in two weeks. If you've got any questions, comments or your own answerable tips, please send those in to show at hybridcloudshow.com. Until then, I've been Aaron. I've been Gary. I've been Sean. And I've been Shane. See you later.
More episodes
More from Late Night Linux Family All Episodes

Late Night Linux – Episode 405
Late Night Linux Family All Episodes

Linux After Dark – Episode 131
Late Night Linux Family All Episodes

2.5 Admins 318: End of Exchange
Late Night Linux Family All Episodes

Late Night Linux – Episode 404
Late Night Linux Family All Episodes