
How etcd Solved Its Knowledge Drain with Deterministic Testing
About this episode
The etcd project — a distributed key-value store older than Kubernetes — recently faced significant challenges due to maintainer turnover and the resulting loss of unwritten institutional knowledge. Lead maintainer Marek Siarkowicz explained that as longtime contributors left, crucial expertise about testing procedures and correctness guarantees disappeared. This gap led to a problematic release that introduced critical reliability issues, including potential data inconsistencies after crashes.
To rebuild confidence in etcd’s correctness, the new maintainer team introduced “robustness testing,” creating a framework inspired by Jepsen to validate both basic and distributed-system behavior. Their goal was to ensure linearizability, the “Holy Grail” of distributed systems, which required developing custom failure-injection tools and teaching the community how to debug complex scenarios.
The team later partnered with Antithesis to apply deterministic simulation testing, enabling fully reproducible execution paths and easier detection of subtle race conditions. This approach helped codify implicit knowledge into explicit properties and assertions. Siarkowicz emphasized that such rigorous testing is essential for safeguarding the sensitive “core” of large open source projects, ensuring correctness even as maintainers change.
Learn more from The New Stack about the etcd project
Tutorial: Install a Highly Available K3s Cluster at the Edge
Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Get every episode summarized
Each time The New Stack Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The New Stack Podcast

How Microsoft is governing thousands of Kubernetes clusters without manual inter...
The New Stack Podcast

Why long-running AI agents break on HTTP and how Ably is fixing it
The New Stack Podcast

Why the Linux Foundation adopted MCP, with Jim Zemlin and Mazin Gilbert
The New Stack Podcast

Fresh data has us asking, does AI demand Kubernetes?
The New Stack Podcast