Skip to content
TrackPodcasts
technologyOct 1, 20265:42

Coping is the Problem

Get every episode summarized

Each time Programming Tech Brief By HackerNoon publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

About this episode

“This audio is presented by Hacker Noon, where anyone can learn anything about any technology. One of the principles that has stuck with me ever since I read the system's BibleBee John Gaul is. Greater than the problem isn't a problem.”From the transcript

This story was originally published on HackerNoon at: https://hackernoon.com/coping-is-the-problem.
Retries won't save you. Some failures can't be prevented. Robust systems need to cope.
Check more stories related to programming at: https://hackernoon.com/c/programming. You can also check exclusive content about #system-design, #systems-thinking, #software-architecture, #designing-systems, #coping, #optimization, #risk-mitigation, #system-design-tips, and more.

This story was written by: @ryanab. Learn more about this writer by checking @ryanab's about page, and for more stories, please visit hackernoon.com.

Robust systems continue to function when components fail. “The problem isn't the problem, coping is the problem” reminds us to design the fallback before trying to reduce the failure rate. Once the system can cope, mitigation becomes an optimisation rather than a necessity.

Know when Ryan Abner turns up

Follow Ryan Abner and once a week we email you every new episode they appeared on — including guest spots the show notes never mention, because we read the transcript.

Follow Ryan Abner

Free. Pick your own day and time.

Hosts & guests

Transcript ready

95 searchable segments. Every word is indexed and playable.

Coping is the Problem

Programming Tech Brief By HackerNoon

0:00
5:42

Full transcript

Programming Tech Brief By HackerNoon — Coping is the Problem. Machine-transcribed; use the interactive transcript above to jump the player to any line.

This audio is presented by Hacker Noon, where anyone can learn anything about any technology. Coping is the problem, by Ryan. One of the principles that has stuck with me ever since I read the system's BibleBee John Gaul is. Greater than the problem isn't a problem. Coping is the problem. This principle encourages designing systems that continue functioning even when a component of the system fails. By way of example, let's consider HTT Prequests. Even in HTTP request fails, our first instinct is to add a retry. Butoting retries doesn't actually fix the problem. The second request can fail. And the third, by adding retries, we've hopefully made the problem occur less frequently. But there is absolutely nothing we can do to ensure an HTTP request will always succeed. So how do we cope? To decide how to cope, we have to take a broader look at the system. At sextend the example, the HTTP request is for push notifications between two servers. Server A is the sender, be the recipient.

There are all sorts of reasons a might not be able to deliver to be in order to cope in this situation. A could implement a new endpoint allowing B to pull notifications. Now, it gets to decide how much effort to spend mitigating delivery failures, or whether further mitigation is worthwhile at all. Replacing the delivery problem becomes an optimization rather than a necessity. In general, when designing systems, you should plan how the system will cope with a problem before planning how the system will mitigate the problem. For the purpose of this article, I will define mitigation as reducing the frequency, but not eliminating the problem. Other examples of mitigation could be increasing timeouts because processing is taking too long. Or increasing cache sizes because responses are too slow. This isn't to say that mitigation is pointless. Adding mitigation for an existing behavior will almost always be quicker than adding a fallback to cope. In the middle of an incident, adding retries might be the difference between 100 angry emails and 10. Coping provides the fallback, while mitigation tries to reduce how often we need it.

By designing the fallback first, we can decide how important the mitigation is. This principle was very important when I was recently building a personal agento replace a spreadsheet. Natural language can be a fickle beast, with a lack of context or precision leading to misunderstanding. Without adding the ability tocop, I'd still be playing whack a mole, tweaking prompts trying to get 100% of the evals to pass. The purpose of the spreadsheet the agent replaced was to keep an account of expenses between two households. We will often pick things up for each other from the shops. Then send a message and slack saying how much the expense was. To which the response almost always was, have you added that to the spreadsheet? Periodically, we would check the spreadsheet and settle the difference. I kept the design for the agent simple. It received single messages in slack and decides which tool, if any, to call. No back and forth, I wanted to keep thin and deterministic part to a minimum, deterministic code is predictable, faster, and cheaper.

In this system, the main tool I expect to be called is. In the main problem I ran into was classification between shared and individual expenses, where a shared expense would be split between the households. I was tempted to solve this problem by crafting a system prompt and tool descriptions so beautiful the LLM gods would weep. Alas, I am only mortal, so I added the tool. To be clear, there is no issue with adding a new tool, but an increase in complexity has to be earned. When an expense is mentioned in slack, it gets recorded, and the slack channel is notified. If the user notices that the details of the notification were wrong, they can send a follow-up message correcting the mistake. I will admit, I did burn a few hours attempting to create that beautiful prompt. I knew I'd still need the edit tool because no prompt was going to overcome the inherent non-determinism, but it was still an interesting exercise. The few hours I did burn on tweaking the prompt were spent trying to add context while keeping the prompt short and general. Keeping it short is important for keeping cost down, and, making sure the prompt is kept

general attempts to not overfit to the evals. I even removed a few evals which I decided were reasonable for the LLM to be getting wrong. After all, I was going to add editing to Cope. Once the edit tool was in place, it took a lot of pressure off the add tool. The add tool no longer had to be perfect, it was now good enough. Tuning the prompt became an optimization, rather than a requirement. I can also add further context to improve error rate if we notice a set of expenses being classified incorrectly. Hoping by editing was fine for our low-stakes expenses ledger. The main cost of a bad classification is the friction of having to correct it. In a higher-stakes environment where incorrect tool calls are more expensive, we would want a different way to Cope. For example, we could send a notification, but delay the execution, giving the user time to cancel. Or, at still higher stakes, require explicit approval through deterministic code before executing risky tool calls. Before closing, I also want to call out that sometimes logging the failure, or have been

doing nothing and moving on, is the right fallback. As long as that's at a liberate decision, everything is a trade-off, and in some cases the added complexity of doing more is not worth it. Design the fallback first, take the pressure off mitigation and let it become an optimization problem. Some problems can't be prevented. All we can do is decide how the system will cope. Thank you for listening to this Hackernoon story, read by Artificial Intelligence. Visit Hackernoon.com to read, write, learn and publish.

More episodes

More from Programming Tech Brief By HackerNoon

View all episodes →