
About this episode
Get every episode summarized
Each time The IDEMS Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
Transcript ready
177 searchable segments. Every word is indexed and playable.
Full transcript
The IDEMS Podcast — 295 – Putting Climate Data Rescue into Practice. Machine-transcribed; use the interactive transcript above to jump the player to any line.
Hi everyone, welcome to the items podcast. My name is James Musyoka, a climate data scientist at items international and I'm here with David Stan founding director. Hi David, hi James. We're going to carry on from our previous episodes where we were talking about data rescue to say a little bit more about how this actually plays out in practice in a couple of examples. Specifically, I believe Zambia and Zimbabwe. Yeah, there's been a lot of work there related to the epics application. Yes, and you've been going in quite regularly to both countries to work with the met offices and in fact you're going back to Zimbabwe, I believe next week. Yeah. And in both cases, there have been previous data rescue projects that were huge and where that data rescue led to a
store of the rescue data that was never being used by the met offices itself. And really most people in the mech office didn't even have access. Exactly. And this was really astonishing to us that they have these huge loads of data which are within the institutions, but most of the stuff don't know that this data exists. And for those who know they don't even bother looking at them. Let's put this into perspective. In Zambia, there are about 40 main stations which are maintained and which are giving data regularly because of the needs of ongoing data on a regular basis. And on top of those 40 stations, there were over a thousand stations that had been rescued. Yes. This is incredible. It is, it is, it is very interesting. Yeah, very incredible.
Everybody talks about data storage in Zambia is a big country. So 40 stations in Zambia is not that many. But suddenly a thousand stations and all rescued or digitized. This is huge. It is, it improves the coverage of met data across the country. The country is huge. The 40, the main 40 stations are just very sparsely in all. And it is very important to have a lot of data that is being distributed. But when you suddenly put in the 1000, then you get a lot more stations covering the country, which is interesting. And if I'm not mistaken, it's been about a decade since those data were rescued. Yeah, that was 2015 on the data were rescued. Yeah, the data is just sitting there separate from the rest. And to our knowledge, they have never been used in a concrete application. Not one. Actually, I think they have never even been opened, so to speak. They've not been processed from the formatting in which they were archived.
So it's not just being used. They've never been looked into. So this is something where we come along with the picture, which Zambia has been incredibly engaged with. And they've really met offices done brilliantly, not only at the headquarters, but also the provinces. Everybody's getting engaged and involved. Yeah, the model in Zambia is a very interesting one, which I don't think we've seen in any country. I think the people that decentralized offices are involved in this application. Not just the people at HQ, which is the case in many countries. So yeah, that model is very, very interesting. And so when we have this at the moment, trying to work with them to be able to get the stations, even just the 40 stations, so that they can be available for E-Pixer is hard work. And we should not underestimate how much work that is.
It's been like what, four years since we started working there. Yeah, it's taken this long to get the data from those 40 stations in good shape. And actually, they still are a lot more to do. So yeah, that's the magnitude. Let's be clear that some of that work. Zambia was the first country amongst the first countries to be implementing E-Pixer. And so some of that work was to set up the systems. Exactly. Didn't have any use case beforehand. They were the trailblazers, so to speak. So some of that work would not be there in the future. It is better known how to set this up. But a lot of that work is something where it is just data rescue clinics, even with these main stations, is what has been needed. You have gone in, others have gone in and actually worked with them to identify missing values and usual values to check them in the paper archives to confirm to improve the quality of the data.
And also to close gaps, which are very prevalent in these datasets. And to understand sometimes those gaps cannot be closed, but just to understand why those gaps exist is something that has formed a major part of this data rescue work. Absolutely. And let's be clear that data rescue work was simply on those 40 main stations that have been actively reported on regularly by the matter of this. This is part of what they use for forecasting. They are their main stations. Yeah. And so even for the main stations, it was a lot of work to try and get them to the quality that's needed for E-pixer. And let's be clear what that quality is. It is 30 years worth of historical data without major gaps and without real anomalies, what could say. Yeah. So what E-pixer need is long-term records, which are fairly complete and of high quality, basically. Yeah. Yeah. And that's actually a big ask.
Now, what's of course very interesting is that the discovery of the thousand plus extra stations rescued up to 2015. In theory, should become a relatively straightforward addition to that data store, but that's not been easy either. No, it's not. It's not been like as simple as that. So in terms of numbers, you would think that they quickly add or improve the number of stations that can be available for E-pixer, because at the moment the baseline is 40. But when we looked at the data, it was not trivial. It's not obvious that all those stations have data that is a required level for E-pixer application. Okay, but even if just 10% at the required level, that would be 100 stations. That would triple the number of stations, which are there.
You'd go from having 40 stations to let's say 160, you know, 120, at least. Yeah, that's true. But that's not easy. And one of the things is that the data that was rescued, it's not, there would still be work needed. I think to get the data up to that states, those requirements that are required basically. But yeah, as you say, even 10% of that, if they can be useful for E-pixer, that would be a great improvement to the number of stations that are available, or even the amount of climate information that is available for farmers to use to make decisions. So this is something where there's a real problem here, there's been a huge amount of work that has been done. Compared to the huge amount of work, the additional work needed is relatively small, but it's still a non-trivial amount of work. It's still a big amount of work to be able to make them usable.
And so this is the challenge that we're facing and that we're seeing in many contexts, that when you have these big projects that happen and they leave an artifact behind, but there is no ongoing effort to make that artifact used and useful, then you're losing the big investment that was made. Exactly. And this is something where we talk about now rescuing data rescue projects. Yeah, exactly. For example, for this particular case in Zambia, if we were to use those stations, the immediate task would be to get that data up to date. Because since the rescue stopped in 2015, there's no data for the last 10 years. And for Epexa application, and I believe many applications, the reason data is required, not just the old data, but the data has to be as reason to us now when we need it and when we want to use it. Let's be clear as to why with climate change data, which does not have the latest 10 years, maybe a misrepresentation of that site's current climate only by having the historical and the current one can actually be sure of not only what the current climate it.
But also what the climate trends are and how things have changed in the past. Yeah, that understanding is very important for these applications. So in that case, I think getting that data up to date is in itself a huge data rescue task. So that these datasets can at least make those requirements to be used. And this brings us to another aspect of the data rescue that we did not discuss in the last episode, which is that it is one thing to actually rescue historical data. And it's another thing to put in place the structures so that the new data that comes in doesn't need rescuing in the future. And that's such an important second piece of the puzzle. Yes, it is. And this is one of the reasons that we feel it's so important to get the right application, even if you have a single use of the data that can provide the incentives to then keep the data up to date as you get it.
So it's a wonderful example of this where actually as an application, the war data is not shared that remains the sovereign property of the Met Office. But the useful analysis of that data for farmers is publicly visible. You actually have all the incentives to recognize when a data source is up to date or not and for it to stay up to date. Yeah, and every year for a picture, there is needs to update the summaries with the most current, you know, the previous season. I hope that will be an incentive to keep the data being rescued every year within the Met Services. For example, in Zambia where a picture is already in operation. Well, and this is where within Zambia, the responsibility for that updating has agreed to be shifted from the head offices to the provinces.
And they have been given the tools to be able to do the data entry in such a way that as the data comes in, it is able to be updated as part of a continuous process. So this is all part of the same trainings that we've been doing over the last four years. And I don't believe we're yet there in terms of those structures. We know that there's maybe a couple of provinces that are really up to speed and more work is needed in many others. And we also know that this historical rescue data is not yet part of those systems. And so there's a lot of extra work which would be needed to really make Zambia a exemplar country of how this works. But from where they were to where they are has been an incredible, I would argue transformation. Yeah, exactly. I agree. But yeah, it's been hard work and it's been a lot of, I think as we said in the previous episode, this is like a new area for them, basically.
So these are new skills they're learning. I think they're getting to realize they need these skills. And so I think even with the time that we have been in Zambia, although we are not yet there, I think what has been achieved so far is really encouraging. And if this can be sustained, then I think we could get Zambia to be an exemplar country, as you say. But it's fragile. I mean, this is one of the things which I think is worth taking note of that there are individuals within the Zambian Met Office without whom there would be institutional loss of knowledge. And it would be hard to continue. So it's something where even with this work and with what's happening institutionalizing this, making this work beyond individuals is difficult. This is with a long term engagement and we should really recognize GIZ for that support and that long term support to try to build something which will sustain.
And it's hard. It is. It is. This last year has been, I think working in Zambia has been sustainability has been something that we have been focusing on. It is hard. I think what we are trying to do as you say is to institutionalize these changes. And I think what we are trying to, the approach we're using at the moment is to include these activities or these tasks as part of day to day activities for the staff in Zambia. I think that would be the easiest way to institutionalize it or the first step basically so that they do this every year as part of their day to day activities. And that remains to be seen how successful it will be. That is something I just wanted to mention that it's the approach that we're using there. And let's be clear that this is both at headquarters but particularly in the provinces, which is what again I feel is so exciting and unique about this approach that the headquarters are important.
But it is really the provinces that I would argue are taking the lead and they're really at the forefront of this. I think they have a good division of labor or tasks if you want. So my understanding is the HQ is in charge of coordinating the whole thing but the main work of improving the quality of the data, asking the data is now with the provinces. So each provincial team is in charge of the data for the stations within their provinces. So they're taking that responsibility and they are coordinating with the HQ people who are responsible for providing them with information as to where the gaps are what kind of improvements need to be done. And then they get to focus the provincial team is now focusing on addressing those issues that have been raised in the reports that are coming from HQ. So to me, I think it sounds like a very good model there they have with the provinces taking responsibility for the data. I think it has led to a lot of improvements in the data for each of these provinces.
Let's move from Zambia to Zimbabwe. And what's interesting there is that this is a project which doesn't have the same history. Our initial engagement in Zimbabwe was actually to try and help with the integration of data systems between the Met Office and the hydromat groups. But has now transformed into this effort to get them set up for E-Pixer. Yes, it is in Marbwe's case is a bit different, as you say, it started from the angle of data integration from different sources. And I think that integration will set up a supply castle to the data being used for E-Pixer. And I think it's only that when we go to the E-Pixer stage where we want to use the data for E-Pixer that we realize actually the data is not in a good shape. So a couple of things didn't happen during the data integration.
And so that meant that the data was not ready for E-Pixer. So that's where I think the gap in Zimbabwe regarding the data for E-Pixer, I think that is how it came up to be seen or to be realized. And I think this is one of our our moments where we realized that the work that was being done, which was all well intentioned, trying to bring together the data from different sources so that they could be accessible in the same way. This is a standard data rescue project. And yet what we realized was it was only when we actually started using it or trying to use it for the application, that the limitations in the data quality became apparent and needed to be acted on and the real data rescue started. Yeah, exactly. It's a bit different, but it's similar in many ways to Zimbabwe. The case in Zimbabwe is much similar to the case in Zimbabwe in terms of the data preparedness.
It's only that I think they approach to realizing this gap in Zimbabwe was a bit different to Zimbabwe. As you say, we had an aha moment where we expected the data to be ready, but then it wasn't. In Zimbabwe we have about 48 main stations, which the metallurgical service is in charge of our operations. So those are the stations that are marked by their own staff. So these are the equivalent to the 14 in Zimbabwe. The parallel, I think in terms of what we found in Zimbabwe in terms of the data that has been rescued, but it's not being used is. I think in Zimbabwe is a very complicated case compared to Zambia, because first of all, they've had a lot of databases resulting from using different clips of versions. And I think those are still stored separately, but I think at some point they came to using the official clips of the versions. So the databases in their own clips of the versions were then matched into this official clips of the version. I think that's operation brought along a lot of challenges with the data quality.
As we saw during the we had an audit just to check whether the data was ready for a picture. And what we found is that actually as a result of the way the data sets were put together integrated into the official clips of what that led to is for some stations. There were a lot of duplicates, which is a very interesting problem that we found in Zimbabwe that was not in Zambia, for example. So we had a very huge amount of duplicates duplicates in the sense that the daily records exist. We expect a single record for a day for a given element, but here we had cases where we had two or even more observations for the same day. And in some cases, they were very different. They are not the same. So it's a very hard problem to solve because then it involves going back to the records just to check which of those is the correct one. If they were the same, then it would be easy to just say, OK, it's just a repetition of the same value. But in this case, it's just different values on the same day. So the challenges there, the usual challenges of gaps and quality of data are there.
But one of the differences is the duplicates problem, which we found in the main station records. So what you're highlighting here is just a level of complexity. The reason the first task was to bring together these different data sources was because this was a known problem. What I think it uncovered were the unknown problems that have or what the E-pixel approach and covered was the unknown problems. Where by the work to bring these together was not enough to resolve the issues that were needed for the data rescue to be able to make it usable for E-pixel. What I'd like to come back to is that you've actually got through most of this now with them for at least 14 stations. But in doing so, you've also uncovered that there was this massive data rescue project in Zimbabwe as well that was again hundreds of stations.
So again, there were over 600 stations that had been rescued and not were there in ways where they've not been used. Yeah, they were just like Zimbabwe, they were just archived somewhere not being used. Yeah, and it's exactly the same fundamental issue that we found that the data rescue happened. The end goal was to get it to a climate data management system. It was putting Climbsoft, but it was not ever used for an application. Actually, I think we were the first people to look into the data to open it up and see how big it is. And again, this was how many years ago? Apparently, this was 2021 when it was been done. So five years ago, and this is something where again, the idea that such a large investment has been made to do that data rescue.
A small investment that's needed to make that useful is again, something which is not on people's radar even as something which is needed. It was just hidden. It was, yeah, it was hidden. I think there was, it was mentioned some time back to us. But yeah, before that, there was not any effort to look at it and to take advantage of this huge data resource that is available there. Just to put things into perspective, I think the way we've used it, because I think in Zimbabwe right now, the focus or the immediate focus is just to get to implement a picture in the three provinces where they have 14 stations. And so what we've done currently is just to use that huge resource of data to try and improve the quality of the 14 main stations because apparently the 600 includes them 48 main stations as well. And so this is a bit different from what we had in Zambia. And so in this case, we've been able to use that and taps data rescued over five years ago to try and improve the quality of the 14 stations that are in the data.
Wonderful, absolutely. Okay, this is probably something we should wrap up and maybe have a follow up later because the thing we haven't really mentioned is what we really envisage is that as part of the application process, a part of getting it ready for a picture, what we're developing is a way of accrediting the data. And this comes back to your Zambia example and really what we learned from them is that really part of the complexity of what we're dealing with is its multiple actors. So one scenario we have there is that it is the data where responsibility is taken by the provinces, but the head office lays an accreditation role to say yes, this data has reached this quality setting up those systems is something which is needed for the application to get to the stage where you can say yes, this station can now be used for epics.
And this is something which is really what we see as the narrow part of the hourglass when we talked about our hourglass model in the previous episode where that accreditation is all happening in the narrow part of the hourglass so that all applications can use the accredited data. And this is the piece that I think we're really excited about it's a lot of work to get to something where we have real demonstrators for this, but it really comes in and we'll discuss this in another episode to the fact that even publications need to change that what we're looking at here is these accreditation processes need to lead to publications, but these are not research publications. They are the equivalent of what in education would be considered practice studies people actually using something which is known in their context and being able to put in place systems to have those sorts of evidence focused publications rather than just research publications is something which we're just starting to discuss with our partners and really fear could be transformative in this context.
Exactly it's a new idea, but I think on a positive not the meta services are beginning to understand why it's important. Yes, well it is coming from them as much as it's coming from us and based on the needs we're observing in these contexts. Exactly, yeah. Anyway, this has been a really good discussion. I think we should have it there and I think on an extra episode should be my credit test sounds good. Thank you.
More episodes
More from The IDEMS Podcast

296 – Navier–Stokes and What AI Means for Mathematics
The IDEMS Podcast

294 – Reflections on Climate Data Rescue
The IDEMS Podcast

293 – The Problem of Zeros in Data
The IDEMS Podcast

292 – Integrating STACK Assessment into Open Statistics Textbooks
The IDEMS Podcast