
Get every episode summarized
Each time Python Bytes publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
Python Bytes is made possible by:
“This is episode 496 recorded Tuesday September 15th on MicroKinity. If you want observability for your apps and your AI agents, Logfire is the business. I will be telling you more about them later. Find the link at the top of the show notes.”From the transcript
- Pandas Should Go Extinct
- Pydantic-pint puts real-world units in your Pydantic models
- How Libraries Run Rust Inside Python (With PyO3)
- AWS acquires DuckLabs
- Extras
- Joke
Sponsored by Logfire from Pydantic: pythonbytes.fm/logfire
Connect with the hosts
- Michael: Mastodon / BlueSky / X / LinkedIn
- Calvin: Mastodon / BlueSky / X / LinkedIn
- Show: Mastodon / BlueSky / X
Join us on YouTube at pythonbytes.fm/live to be part of the audience. Usually Tuesday at 7am PT. Older video versions available there too.
Finally, if you want an artisanal digest of every week of the show notes in email form? Add your name and email to our friends of the show list, we'll never share it.
Calvin #1: Pandas Should Go Extinct
- Pandas' slowness pushes teams toward "Big Data" tools (Spark, Databricks) they don't actually need — most workloads never hit true Big Data scale
- Amazon Redshift telemetry: ~95% of tables are under 100GB, ~87% of queries touch 80GB or less — that's "Medium Data," not Big Data
- Polars and DuckDB fill that gap: single-machine, fast, no cluster required
- 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory
- On a real-world NYC taxi dataset (3GB parquet), pure DuckDB ran 2x faster than pure Pandas while using a fraction of the RAM
- Bonus: Apache Arrow lets you pass data between Pandas/Polars/DuckDB with zero copying, so trying them out doesn't mean a full rewrite
Michael #2: Pydantic-pint puts real-world units in your Pydantic models
Pydantic-pint bridges Pydantic and Pint so models can validate physical quantities like 4m or 12 meters instead of bare floats. Fields annotated with PydanticPintQuantity parse user input, convert between compatible units, and serialize quantities back out as strings. That closes a real gap for anything consuming API payloads, config files, or sensor data with measurements, letting you enforce units at the validation boundary instead of hoping every caller remembered them.
- via PyCoder's Weekly newsletter
- Unit mix-ups have literally crashed spacecraft; now your Pydantic models can refuse them at the door.
- Annotate a field as Annotated[Quantity, PydanticPintQuantity('km')] and inputs like 12 meters arrive auto-converted to kilometers
- Validation covers string, numeric, and quantity inputs, and model_dump_json serializes quantities as readable unit strings
- Installable from PyPI as pydantic-pint, MIT licensed, with docs at pydantic-pint.readthedocs.io
- Early-stage solo project at version 0.4, so API stability and maintenance are open questions worth discussing
Calvin #3: How Libraries Run Rust Inside Python (With PyO3)
- Pydantic v2's validation core (pydantic-core) is Rust under the hood, built with PyO3 — this post shows how that bridge actually works via a small hand-built JSON parser
- Four steps to get Rust into Python: write a normal Rust module, annotate with PyO3 macros (#[pyfunction], #[pymodule]), compile/install with maturin, then just import it
- The parser builds a Rust tree first — Python never touches it until the boundary crossing
- Key insight: converting the Rust result into Python objects (.into_pyobject) is often the expensive part, not the parsing — 100,000 JSON values means ~100,000 Python objects built after parsing's already done
- Errors cross the boundary too: Rust's typed errors convert into real Python exceptions (ValueError, FileNotFoundError) via From/?, so callers get clean Python semantics
- Takeaway for anyone porting Rust in: if you're returning a scalar, don't sweat it; if you're returning a big structure, profile the boundary — that's the real cost, not the algorithm
Michael #4: AWS acquires DuckLabs
Thank you Dylan McConnell.
What does this mean for the DuckDB ecosystem?
DuckDB is the open-source in-process analytical SQL engine. MIT licensed. The IP is not owned by any company - it's held by the nonprofit DuckDB Foundation, which was created when the team spun out of CWI Amsterdam. Peter Boncz, the CWI representative on the Foundation board, describes it as the entity that holds all IP of open-source DuckDB.
DuckLabs (ducklabs.com) is the company, formerly branded DuckDB Labs. Founded a little over five years ago by Hannes Mühleisen and Mark Raasveldt to give the DuckDB team a stable long-term home, bootstrapped deliberately instead of taking VC, grown to 30+ people in Amsterdam, funded by support and feature-prioritization contracts. It employs the core devs. It does not own DuckDB.
DuckLake is one of three projects DuckLabs builds, what they call the Duck Stack: DuckDB, DuckLake, and Quack. DuckLake is the lakehouse format that puts catalog metadata in a SQL database instead of in files on object storage. Quack is newer - an RPC-style protocol that turns DuckDB into a client-server system where both ends are DuckDB instances, slated to stabilize in DuckDB v2.0 in September 2026.
MotherDuck is a separate Seattle company, Jordan Tigani's, selling serverless hosted DuckDB. It was started in partnership with DuckDB Labs and has worked closely with Hannes and Mark for four years. It contracted DuckLabs for engineering work and contributes heavily upstream - three of its engineers are among the top 10 outside contributors to DuckDB. It also sells its own DuckLake offering. Customer and collaborator, never owner.
What the AWS post changes. Amazon bought the company, not the project. DuckLabs joined AWS effective September 1, with the process concluding August 31, 2026. Hannes and Mark keep leading the team and the project's technical direction, the team stays in Amsterdam, and DuckDB stays MIT under the Foundation. AWS gets the people and a direct line to the roadmap. The license protects your code, not your priorities.
Three second-order effects worth tracking:
The Foundation board is the real question. It has three directors: Mühleisen, Raasveldt, and Boncz. Two now work for AWS. Commentary on the deal has focused on exactly this - the license protects the code, not the roadmap. The announced counterweight is governance: a technical advisory board on the Foundation, and opening the extension stack so extensions signed by other developers can run in DuckDB.
MotherDuck immediately moved into the business DuckLabs vacated. It now sells DuckDB enterprise support, which it had avoided because it didn't want to compete with DuckLabs' business model, and says it has explicit blessing from Hannes and Mark now that they're joining Amazon. It also bought Tower.dev the day before the AWS announcement.
Everyone expects an AWS DuckDB service. Tigani says Amazon will likely release one eventually, and welcomes the competition, citing Redshift's failure to slow Snowflake on AWS. The groundwork is already visible: Amazon Quick uses DuckDB to query S3 Tables and has processed over 2.5B queries with it since launching in October 2025.
The DuckLake angle is the one to watch. AWS is heavily committed to Iceberg through S3 Tables, and it just acquired the team behind a competing lakehouse format. The stated plan is to use DuckDB, DuckLake, and Quack together to power a new generation of data services, but which format wins internal priority is unannounced.
Extras
Calvin:
- astral-sh/uv 0.12.12: code-signed release binaries 🥳
Michael:
- My MacBook power supply rebooted to install updates (?!?)
- The Story of VS Code | Official Documentary
- Amazon/AWS acquires DuckLabs (see recent episode on DuckLake)
Joke: We’re agentic now
Get every episode summarized
Each time Python Bytes publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
375 searchable segments. Every word is indexed and playable.
Full transcript
Python Bytes — #496 A lake house in Seattle. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hello and welcome to Python Bites where we deliver Python news and headlines directly to your earbuds. This is episode 496 recorded Tuesday September 15th on MicroKinity. And I'm Calvin Hinterchbacher. This episode is brought to you by Logfire from Pydantic. If you want observability for your apps and your AI agents, Logfire is the business. I will be telling you more about them later. Find the link at the top of the show notes. Follow us on the socials, all the various things you can think of are there on the episode page as well. Inside of for the newsletter, I just sent out the most recent one a couple of days ago. It was a little bit late, sorry folks, but really cool stuff that we add, like extra information that doesn't even appear in the show that helps you get a little more out of the show. Yep, I love all the context it adds. Either too. I'm like, well, that's pretty good. We found some good stuff here. I would say that this newsletter that we're right in here, it shouldn't go extinct, but something might need to go extinct.
That's good on you. Yeah. So I found, this is a blog post from what's Eddie's last name. I was down here at the bottom, I was cap your Eddie Atkinson. He gave a talk at the most recent latency conference. So it's actually a talk from last year, but I think you kind of brought it back a little evergreen to it into a blog post last week about pandas that should go extinct. And we're not talking about the cute little fluffy things that are used for international diplomacy, but the Python data frame library. Michael, how many times have you thought you had big data only to find out you were ready to defintestrate your laptop because pandas was the problem? You know what? It's happened. It's happened. I'm going to get my dictionary real quick and then I'm going to know that that happened. I think the issue is a lot of folks really don't have truly big data problems. I mean, we've done some big data projects in the past, which were 10,000 tables, petabytes of data. That's truly big data.
Most folks probably lie in the medium sized data, but pandas definitely cops out. I mean, he does some interesting benchmarks in here. Gives a couple good code examples. Actually shows a really interesting post from Amazon Redshift team where they were looking at the composition of many of the tables that are out there in the Redshift environment. If anybody is going to have a good view on what the size of data is and what big data could be, they're probably the ones to look at that. But if you look at this chart, they basically say on a continuum of data size, most folks start over here in Excel. You've got under a gigabyte around a gigabyte data. About that point in time, Excel falls over. It's probably time to pick up another tool to handle that. And a lot of people reach for pandas because I think there's just a lot of built up inertia or momentum in the community around the pandas and data frames. And it's an easy UI. It's been taught in a lot of universities. So there's just not a lot of like need to kind of move out of that space because there's a lot of good code examples. A lot of blog posts have been produced. A lot of data science is based on pandas.
But they're really based on data frames. And there's more than one library out there to handle data frames and probably do it more efficient. So if you actually looked at the chart here, they're basically saying, when you get up into like the 10 gigabyte range for data sizes, pandas is probably still pretty good. But then there's a gap. It falls off somewhere between 10 and 100 gigabytes of data. And 100 gigabytes of data these days is not infathible. You can easily go find sample data sets that are in that realm and that range and that size. And so the next thing they reach for is typically a commercial tool like Databricks, Snowflake, Dask, or some of these other things that are spark. So you're distributing the memory of that data set across many machines or maybe even across one very large machine, but doing it in a distributed manner. Most people probably don't need to go that far. Like most people probably are still sitting in the range where you can see on this chart that Polar's handles. Polar's can handle straight up into 100 gigabytes of data easily on single machine. And then DuckDB takes it to next up, which actually kind of fitting for this
episode. I think this will be this will be an interesting episode because there's a lot of information here about DuckDB later on in the show. But he gets all kind of goes on again and shows that basically the average size of a row and rev shifts is about a kilobyte. Every redshift cluster has like 10 machines in it. They're capable of guzzling eight gigabyte gigabytes per second for my three. But really an actuality 90 almost 95% of the tables in redshift contain fewer than 100 gigabytes of data. Most people are still in the range of just using a single machine with DuckDB or even just Polars, which is probably similar, simpler to maintain and manage, but it's above the reach of pandas. That's why the the post is kind of going on about pandas needing to go extinct. Another interesting bit again, kind of good code examples in here. When we get down into some of the tables for the performance, what strikes you here, Michael, on their memory usage, the pandas library, we're talking about 30. This is one of his examples. I can't remember which one, but four minutes basically for the duration of the processing, 38 gigabytes of RAM. If you get into Polars that gets
halved, 18 gigabyte RAM. And if you go into DuckDB to do the same operation, five seconds at 1.93 gigabytes of RAM. So even the nine to 20 times as much. Yeah. Yeah. I mean, an iPhone could do this operation against the 100 gigabytes of data. It's fine though, because you can just get more memory, memory cheap these days. Totally cheap. Totally, totally cheap. So I just think focusing with this one, the bed, pandas was probably a good way to start, but you'll notice in here, like the Polars notation or syntax really, really similar. I like the DuckDB syntax. I think he's got some examples on here where he reads in and does some more operations. He does a couple against the larger machines, then falls back into an older laptop, like a framework 13, to show that this is still useful as a developer tool. So I think here, Polars versus DuckDB kind of comes down to your workload, your experience and your preference. It's a good post. I really liked all the
code snippets he links over into the GitHub where you can actually try it yourself. It goes against the New York City taxi data set. It's got a ton of data in there. So it's fun to play with. And you can see that we want to wait minutes or do you want to wait seconds? And would you want to use all your memory for this? Or, and actually, the other thing that Polars and DuckDB did a much better was utilizing the CPU. Pure pandas in this case was using like a multi-core machine, 146 percent CPU, where if you go to pure DuckDB, there were over 800 percent CPUs. Obviously, eight to 10 cores are being fully utilized as opposed to basically one and one and a half cores. Yeah, that's awesome. I feel like this is a pretty data-heavy episode for the data science crew out there. Yeah, it wasn't on purpose, but we ended up that way. Yeah, I have a little bit of real-time follow-up for you, Calvin, for people who are going to ask the one last bit in here. So he does mention Apache Arrow. And if you've not played with it, it'll add him to switch back and forth between pandas, Polars, and DuckDB without having to reload or copy the data into RAM. So you could
actually do the same operation with each of the libraries without actually having to take the data back out of RAM. So check out this Arrow, which is really cool. It's kind of a little bonus side bit that was in the blog post. So data folks who got medium-sized data, this is going to be a godsend for you. Yeah, I think Arrow is the foundation of pandas too, if I remember correctly, and also a Polar, so that's pretty sweet. Yeah, my real-time follow-up here. Oh, if you were working with one of these and you want to switch to the other, I had Marco Garali talk Python a while ago to talk about narwhals. And narwhals is a facade, adaptive layer that speaks native Polars, but also talks pandas. So if you want to try, I'm like, oh, let's see what we're doing. You know, you could use this as an intermediate layer to kind of swap that out a little more easily than rewrite and everything. Yeah, I think people just need to drop, drop pandas. I mean, it was great. It was great 10 years ago. Yeah.
Yeah. I have some funny jokes, but let's carry on. Let's move on to the Pydantic pint. So, I think the pint is really interesting. Do you know pint? Are you familiar with pint? I've never used pint. So pint, we've covered that on the show back in the day. And pint is interesting because if you're, I mean, all you got to do is say, Marslander, sample return, whatever. And it's like the $100 million plus fail because somebody used feet and somebody used neighbors or something like that, right? And so pint lets you do math in Python with units attached, which is pretty cool. All right. So I can say instead of just having a distance, I have 42. I have 42 kilometers. And you can say like two miles to whatever and so on. So it basically mean forces you to work in units, right? A lot of times we don't do this as regular programmers, but if you do anything scientific, well, there you go. Right. So that's the background on pint. But what I'm going to talk about is actually not pint. It's called Pydantic pint.
Because Pydantic is an awesome library that lets you validate the inputs and parse them and everything whenever you read some sort of JSON, right? Like if it's fast API or just a JSON file or whatever. Geny database, SQL model, all those things, yeah? So Pydantic pint takes, pint takes this idea and adds units to your data validation libraries. So instead of saying I have a box that has a length and a width, I can say I have a box that has a length and a width that is a pint quantity. And the validation is to convert it to meters. So even if you parse something that says feet or centimeters or whatever, it will show up correctly. What do you think? Well, that's definitely handy. Yeah. It's the kind of thing that's like you don't have to, you're not going to use it a lot unless you're really in, you know, some kind of engineering something. Yeah. Oh my god, this is so good. It's like so perfect. This has got to solve so many, like they said, small mistakes that end up in huge damage. A hundred percent. So yeah, and to be able to validate with it too. Yeah, just automatically, right? Just all the Pydantic validations.
You just have a Pydantic based model and it parses over to whatever it is. And if you put, I don't know, leaders into the length, whoa, whoa, whoa, you can't convert leaders to meters. I don't know. And it kind of fits perfectly under the Pydantic scope of the data validation and serialization. Like, there's natural. Like, this should exist and they made it exist. Yeah, it's really cool. So you can have, say, a fast API endpoint that just automatically just takes units and automatically converts units. And yeah, it's a really nice one there. Speaking of really nice, now this transition here has, this has nothing to do with the sponsorship, the previous one. I just happened to do great stuff and open source code too. But Pydantic also happens to create log fire, which I told you about at the beginning. So let me go ahead and tell you about our sponsorship offer log fire, not Pydantic pint, which is not even from them, but it's based on Pydantic. So here's the deal. It's 2 a.m. your AI agent failed. Was it the model, the tool
called the database, just the general unreliability of, hey, I think a new model is coming. So the current one starts breaking periodically. So most observability tools, they can't tell you, because they only see part of your stack. Pydantic log fire sees all of it. One trace across your agents, LMS, APIs, and databases down to the infrastructure services, Kubernetes hosts, is built on open telemetry with SDKs for Python, TypeScript, and Rust. And it works with any OTEL compatible language, every prompt, token count, and cost right next to your vector searches, your query, you query everything with Postgres compatible SQL to understand what your app is doing, and you soak in your coding agent. They can use the same way, because it talks SQL really well. If you connect it to the MCP server, your agent can also figure out what is going on. So stop guessing, read the trace, Pydantic log fire. AI, it is still just engineering, even if it's weird engineering these days. So visit Python by.s.fm slash log fire today and sign up. Get 2 million records free every month, no credit card required. You can even click, and I really like this,
there's a little copy of this text on board with your agent, click that, and it gives you a prompt you can drop in the cloud code or code extra, whatever, and it automatically knows what to do to set up log fire. So thank you to Pydantic for supporting the show. Calvin, I know you're a big fan of the visibility into the token. Yeah, I was just curious now if copy the setup prompt isn't the new pipe to bash, like pipe some curl to bash. Yeah, this is the replacing that. That's what it is. That's what I was going to do. Yeah, no, no, I think it isn't. It's amazing. I have some stuff that I'm working on. I'm like, oh, this is like this idea is perfect. I love it so much. So yeah, recall, thanks to Pydantic for sponsoring the show. Yeah, thank you. And let's jump over to your topic next. He was right there on or in the other show. What's next? Well, speaking of Pydantic, this comes from Bob Bilderboss, a friend of the show. I know you've had him on numerous times for other events and things, but this one is about how to run how rust code becomes something you
can import. I think it's interesting that we can if people are complaining about performance, we're the first news article I had about getting rid of pandas and bringing back in with Polars and Duck TV was about performance. This is similarly veined. Like if I've got a very computationally intense data structure or function that's happening in my program, it'd be sure to be nice if I can maybe replace it out with a rust version of that, but have it act natively inside of my Python code. So this blog post from Bob goes over basically what Pydantic V2 does, which is a data validation library that most Python apps are using these days. It's actually a rust extension under the covers. It does the work. It's core, Pydantic cores all built with Pi 03, the same tool chain we're going to use here in this example. So it goes over some examples of basically you write a normal rust module, you annotate it with some specific macros, and then you'll actually be able to import that into your Python code fairly naturally. I think there's basically the rust parsers
incredibly fast. Python never touches anything until until the boundary crossing. So if you call for some data that all happens over in rust, the thing that I think the article covers it's really important is that if you hold back that data across the boundary, those rust results get turned into Python objects. And so you're going to want to think carefully about how you bring back parts of that because maybe you're only interested in a small piece of what is coming back and you don't need to populate. Like for example, 100,000 JSON values means that you're going to get 100,000 Python dictionary objects after the parsing is all done and everything gets passed back. You don't incur the penalty until you cross that threshold back in the Python land. Maybe you don't need all 100,000, but there's some other operation you can do to get down to just the pieces you need. So what's nice is errors cross the boundary too. So if rust gets runs into errors, those come back as Python exceptions. So it makes it easy to debug and figure out what's going on. So you get clean Python semantics while still leveraging rust. So basically for anyone porting rust, if you're
returning a scalar, don't sweat it. If you're returning a big structure profile the boundary. And if that's the, you know, if it's, if that's a real cost, you want to switch over and maybe do more of the algorithm on the rust side. So much like you can use C or other languages and I don't know if you can use Ruby to do this kind of thing. I don't know if you can or not, but we can definitely use rust. And I'm kind of excited about that. I know you've been doing some coursework on it and sounds like Bob has also made some learning materials to lead folks through. It just feels like that there's a really nice friendship between the rust communities and the Python communities and all the nice cities that have been put in place to allow us to use rust and almost natively over in the Python world. So thanks Bob for the awesome post about that. He's again code examples in here, kind of explains to the Python folks who have never touched rust, what the function signatures look like, which I appreciate because breaking it down and telling me what each of those pieces means that I can now pretty easily read some rust code and understand what's going on because it doesn't
look terribly, yeah, then look terribly foreign to me, but it's just different enough, but this goes over a good, a good usage of what each of those pieces mean for you. It's pretty surprisingly similar to Python, honestly. Yeah, yeah, and incredibly fast, but you think about things a little differently because of the way it manages memory and I think that's a big, that's the big difference shader. What was that cartoon? I was the guy's like, I gladly pay you on Tuesday. It was Popeye. Popeye, I was Popeye, right? Yeah, it was a borrowed checker with a borrowed, it was wimpy, who will, who will glad you pay you Tuesday for a hamburger today. That's the difference of rust as you've got the borrowed checker always check. Yeah, yeah, always check again. So yeah, I don't have you've you've been doing a little more with rust and Python and and and teachings from folks these things. Yeah, yeah, a little bit, a little bit. I I have two follow-up here. So you talked about DuckDB, but the question is you have a lake house. You're right, a lake, I mean, a lake house. So we've heard of data lakes, which is a place you just kind of dump a ton of like you'd
insane. This is like back to your big data thing, DuckDB thing. You just dump a bunch of data into this data lake and you figure it out. Well, that's grown up a little bit and now there's this thing called DuckDB, but it's an implementation and example of what's called an open lake format. Who knew? Do you know? I did like that. I know, I know, I heard that. We've done lake house implementations. I didn't know there was an open lake format now. So the story is what if what if we could use s3 to scale our data access, right? s3 scales pretty large, right? If you can read stuff off the file system instead of out of memory, you can scale that tremendously large. In the open lake story is well, if you put file formats in s3 that everything could read like maybe JSON files that tell you what the files mean, you could read them first. Here's where the data lives in each piece and then parquet files or zipped CSV. I don't know, take your pick, right? It could be whatever. So I just had the folks from DuckLake on, which is a DuckDB implementation story of this open lake format. Well, then I get this message here saying,
guess what? AWS, I'm sure you know this is it. I did see this one come by. Yeah. In DuckLabs, just offered basically the H1 is bad. Let me read, let me read the first sentence. Today we are announcing that Amazon has signed an agreement to acquire DuckLabs, the Amsterdam-based company behind the open source analytical database DuckDB. I'll put a link to the announcement and I thought, well, what does this even mean? And I didn't, I wasn't entirely sure. So I, I like, I wait and did some looking here. I'm like, there's, there's actually a lot of pieces in play. So let me lay it out and I'll tell you what part AWS acquired, what part didn't. Okay. So first of all, thanks to Dylan McConnell who sent us in. What does this mean for the DuckDB ecosystem? So first of all, DuckDB, which you gave a shout out before twice really, is the open source, in-pros analytical SQL engine, MIT licensed. It's like SQLite, but for columnar data, which if you're in way more and way, way more. Yeah. Anything you pointed out becomes SQL queryable. It's amazing.
Yeah. Yeah. Yeah. So you can say pointed out a panda's data frame and then do SQL queries. Yeah. That's your panic. Like, there's a bunch of plugins. It's a, it's far beyond just a database, but it's an in-process sort of data processing engine, much like that. So the IP of this is not owned by any company. It's held in a DuckDB foundation. Oh, good. Thank you. That sounds good. That is good. It was, it was, it was not a CWI Amsterdam from the folks who are mentioned that article. Good. But we'll come back to it. Then there's Duck Labs and the story is, Amazon, AWS is a acquired Duck Labs. This is the company formally branded DuckDB Labs. Founded over five years ago by Hanna's, Willisson and Mark Rossvelts to give DuckDB a stable home, bootstrapped, agree to 30 people. Now they can go chill on their island, which congrats to them. Because DuckDB really has taken over, right? Then there's Duck Lake, which I mentioned earlier. It's one of three projects by Duck Lab. And there's DuckDB. There's Duck Lake. And then there's an
API for working with this called Quack. Check out the Duck By Thought episode. But it's that open open Lake house. All the waterfowl puns are great. It is. And we actually on the podcast had a phone conversation about like, do you need a more serious name like RPC for your data lake? No, they're like, we're calling it Quack. Come on now. And then also we have Mother Duck, which I think I thought Mother Duck was online version of DuckDB. But no, that is a separate Seattle company selling serverless hosted DuckDB. And originally it was started in partnership with DuckDB Labs and they worked closely with Hanna's and Mark for years and even contracted Duck Labs for some of the engineering. So now, what does this all mean? So the foundation owning DuckDB is awesome. But it has three directors that people who own Duck Labs. So there's a bit of a how much independence is it really going to have? There are some other folks, other other governance and so on there. But
you know, it's cool. There's a foundation. It's not super independent of DuckDB at the moment. Maybe it will be though after this. Mother Duck immediately moved into the business that Duck vacated. They now sell enterprise support for DuckDB and so on. Everyone expects an AWS DuckDB service, probably a Duck Lake as well. It already used S3, right? But maybe just a little more formal. Yeah. So there's S3 query and some adjacent like technologies that sound like this may be this will augment or replace. What's really nice about Duck Labs is it runs a local DuckDB or a local Postgres server. And a lot of the chatty API that would come from an open table format and the metadata now all happen the database and then it just fetches and reads the files. Yeah. That's pretty cool. Yeah. We can't beat physics if we can keep the data the bits where they're at physically and bring the compute to it. That's the win. Yeah. Yeah. So the Duck Lake angle actually is probably the most interesting one because AWS heavily committed to iceberg S3 tables. Yeah. Which is a competitor, at least a competitor competing concept to a Duck Lake.
So yeah, check it out. I think I think AWS just got better. We'll see what that means for the rest of the world. What do you think? I mean, you're on the inside of this a little bit. Yeah, a little bit, but you know, they've we've had mixed reviews on their handling of open source. But luckily that they don't have any control over the open source other than they've just bought out the founders who are on the board of the open source foundation. It's still separate enough of an entity. I don't think there's a conflict here. I mean, it's going to be good for the project. The project is already incredible. Like the DuckDB stuff is it's just like if someone had thought about SQLite inside, I need to grow all these other features that handle all kinds of crazy data and do give me native like JSON access and functions. And it's a really great platform for building cool little utilities or talking to giants chunks of data as we saw in the first episode. Our first article. Yeah. Yeah. Yeah. And use barely any memory. I mean, that's why this will run and work. Yeah. The DuckDB part is really interesting on that. And then the Duck Lake is like insane. You know, you could have terabytes of parquet files all broken in little. No big deal. No big deal. In BD. In BD. Yeah. Well, how about some extras? Well, I will continue on my
beating the UV drum. The latest release every week. We got something new. The latest release from the UV folks. We get code signing on Mac and Windows. So the binaries are now officially signed, code signed. Again, I think this is all coming together, ensuring we can secure the software supply chain part of this. So I'm excited to see that. That really is too. I was hoping I wouldn't see a UV thing this week, but sure enough, it popped up in my feed. And I was like, I have to mention it because they just keep making everything better and better and better. So UV is now code signed. So you can trust that it came from the right source on your own machine. If you're on Mac, yeah, that's excellent. Excellent. Yeah. We see it still going. Oh my gosh, the code signing is such a pain these days. It used to be you could just build an exe or dot app and you could just go here, try my app. And now dangerous is that. I know that's how the world used to be though. Well, we used to have what are log in with no password to remote machines. What could go right?
It's fine. It's fine. Trust. We've got to have a lot of trust. Why would somebody do something mean to computers? I don't know. I don't know. I remember in Windows 95, we had a bunch of them at a university. I worked at blog straight into Ethernet and Ethernet. Everyone got its own IP address. And guess what? I think got taken over pretty quickly. The university I was at. They all had public IP addresses in the labs too. Yeah, I didn't go well. No, I didn't. I go well. Speaking of things that need to be patched and updated, check this out. So there's two things that involve research. But I have a MacBook Pro in five pro. Very nice. Love it. I got it this summer or earlier, maybe spring on it or whenever I got it. And it came with a power supply. It was power bricks. The power brick had to reboot the other day to update itself. I was sitting there working and I saw this article come out, come by, say Apple releases a firmware update for the 141 USB C power adapter, which is the one that runs the MacBook. It's almost that that's almost actual size right there. It's so big.
Yeah, it's a beefy boy. Yeah, I think this is even a little small. This big old picture of it. It's heavy. But I read this article and I was saying they're working. In my Mac, when it comes off of power, it dims the monitor. Yeah. So I'm just doing this. And I was an hour, two later, I was sitting there and everything goes dim for a second. The little power disconnects and then two, two, three, five seconds later, something like that. Power comes back, brightness comes back. I'm like, I just my power brick just rebooted. What in the world is going on here? That's crazy. Weird world. I replace all my Mac power bricks. I've got not that I'm trying to be an ad for anchoring like that, but the anchor prime has a 160 watt, like eity bitty little power brick that because I travel quite a bit, it has four USB C ports on it. And they can all deliver up, you know, a combined sum of 160 watts. So I can full bore charge my MacBook Pro M4 Max and my iPad and my phone all at the same time. That's beautiful. And it's smaller than that brick.
Yeah, I'm also a fan of the anchor stuff. This one, since I had it anyway, I just plugged it into the wall part of the house. And just if I'm in that part of the house, I just grabbed that that chord. But yeah, normally if I travel, I have an anchor that's actually a power brick, little battery. And it has two USB things and it'll do not quite as high as yours, but pretty high. And it's super nice because it's also a power brick. Right. So if I need to charge up the battery charge, and charge the MacBook, but then just get it and just plug it into the wall and then it just becomes a power thing. I'm also a fan of this anchor stuff. Okay, a couple more extras really quick here. We've got, you know, play. So the story of VS code, the official documentary is out. Have you watched this? No, I'm not watching this. It's an hour and 38 minutes and I'm here for it. Okay. All right. It's got a lot of people that maybe you didn't see coming like Eric Gamma, for example, you know, thinking back to the gang of four patterns and all that kind of stuff. Because he was apparently involved in the early days. So yeah, cool. We had Colt repo do the Python documentary. We had them do the JetBrains document or IntelliJ documentary in his the VS code.
I'm really loving these like high quality production. I mean, these are these are nice little nice videos. Yeah. A lot of people behind it. I mean, there's such there's an audience for all these things I'm here for it too. I love the fact that the underdogs can feel and they are important for us. And we can now hear more of the story about how some of these things came about. Yeah, it's really interesting. I mean, VS code has taken over so much. And then yeah, anyway, the origins are way way more less ambitious. Let's say it's cool to check out. So it also has over a half million views. So there is an audience for this apparently. There's absolutely audience. Okay. Think see me going to things to make you reboot yesterday last night. I guess today Mac OS Golden Gate iOS Golden Gate watch OS 27 Golden Gate. All those things came out. So did you upgrade? I did. Why wouldn't I mean, I'm like, let's go. I'm not yet afraid of this. I'm usually of that opinion, but lately I may wait a month or till a dot one or O dot one to come out. I spent one day working with it and it's so far it's okay. Okay. I'm going to upgrade then on
your on your full recommendation. Well, I've not upgraded my MacBook or my streaming computer. I only recorded my main desktop. So we'll see. So we'll see. Yeah. Honestly, the one thing to be a little careful about is developers is the Reseta 2. Yeah. Like that's going away. So your ability to run Intel compiled stuff. You might think, Michael, why would I run and come and tell compile stuff like, ooh, Docker, certain Docker things only have Intel versions. So that's going to be a mega. I mean, that's pretty rare. There's people of cross compiled most of the stuff. Because one of the last time an Intel Mac was released. Yeah, but if you let's suppose I'm deploying to an x86 server. Yeah. And I want to test something. I think it's gotten a lot better. It definitely has gotten better. But it used to be the things I don't know. It used to be certain stuff would only work in an x86 version. Oh, I remember this. Yeah. But that was like three, four or five years ago when I was doing really dealing with that. Actually, it was when I was doing with like Databricks and trying to coordinate that stuff. Nice. So this one still has it, but the
one after it, whatever that's called won't. So this is like your last safe upgrade if you're worried about the reverse editing. So did Apple actually deliver some AI features this time? Well, I'll tell you what the new Siri caught me off guard. I'm like, oh, yeah, I did. I did actually upgrade the phone. I guess it has the new Siri because it sounded I asked it something like set a timer and it said something completely definitely I'm used to and it sounded better. I'm like, oh, wait, I haven't had a chance to test it though. It did set the timer like a champ. Let me tell you. Well done. Way to go. All right. Let's let's talk a joke. Okay. Speaking of, you know, the new series supposed to be agentic. So the joke is we're agentic now. We're an agentic startup. You ready? This is how you there's certain things you've got to position yourself. I was just watching an ad because I started watching football yesterday and normally and exclude ads are excluded from my life, but apparently not a football American football and there's some ad for zoom that zoom is an AI company. They're not about meetings anymore. Nope. They can they can they can book that thing so that you're you're dry clean gets picked up. They can do they can do a slide
chair like what? Okay. So everyone's got to be some kind of AI thing now. So here's the joke. I changed all of our loading dot dot dot states to thinking dot dot dot. We're an agentic startup now. Perfect. I'm going to get right on that. Yeah. Yeah. Get right in there like you can there's so much VC money to be had from this. Go for it. Just bombardulating. No, I'm thinking. Oh, I hate that so much about cloud code. It drives me crazy that it's got all these random little words. Yeah. The reason I don't like it is I don't it feels like it's made for someone with ADHD which is can't possibly let it just be for like five seconds. And so if I'm like doing something else and I look over and like word started to like oh, maybe it's no, it's not done. It's like maybe it's done. Oh, no, no, it's just still like randomly. Like could it just have the little icon go? No, no, no, no. It's thinking it's convoy lady and it's wording. I don't know what is it doing? It's why you need her to exactly. That's that's that's the story from another episode. All right. Sounds good. All right. Well, thanks as always for being here. Calvin and thank you
everyone for listening. Talk to you soon. Yeah, bye.
More episodes
