
Duck Tales: DuckDuckGo will find and remove your personal information — like your name and email — from sites that store and sell it (Ep.44)
Om avsnittet
In this episode, Konrad (Engineering) and Shilpa (Engineering) discuss Personal Information Removal: what is a data broker, how we remove your records without ever sending your data to our servers, and the ‘whack-a-mole’ of constantly changing data broker sites.
If you have feedback on Duck Tales, or episode ideas, email us at [email protected].
Learn more about the DuckDuckGo Subscription here.
Disclaimers: (1) The audio, video (above), and transcript (below) are only lightly edited and may contain minor inaccuracies or transcription errors. (2) This website is operated by Substack. This is their privacy policy.
Konrad: Hello everyone and welcome to Duck Tales. We will look under the hood of DuckDuckGo, talk about stories, technology, and people who are building the privacy tools that you all use. And in each of the episodes, you’ll hear from different employees about our vision, updates to our products, how we do engineering, or how we approach AI. So today I’m your host. My name is Konrad Zwinell. And I’m working at DuckDuckGo for seven years now. I work on the front end team at the very beginning. Then I run privacy engineering team. And now I’m back on the front end team working with Shilpa, who is my guest today. And Shilpa is running a personal information removal tool team. And she will tell us more about it today. Shilpa, why don’t you introduce yourself?
Shilpa: Yeah, hi. I’m Shilpa. I have been in DuckDuckGo for, I think, a year and a half now. I joined in on the Apple team working on iOS and macOS. And for the last few months, I’ve been driving the team that does works on the personal information and mobile product.
Konrad: Nice. OK, so first question that people may have after hearing this intro is, what exactly is personal information removal? Could you tell us, high level Shilpa, what’s the product and what it does? What’s the idea behind it?
Shilpa: Yeah, so in the US, this is predominantly a US-based product. In the US, there’s entities called data brokers that basically harvest your personal information from various different sources, your online activity, social media, what have you. And they collect that information, and they basically buy and sell it as a commodity for advertisers and different entities. And our product basically goes and looks at all of these data brokers, looks for your personal information. If it’s there, then it submits a request to remove it. And if the data broker is complying, they will go ahead and remove it and we will be able to verify the removals for you as well.
Konrad: Nice. So basically reducing your footprint on the web, reducing availability to your personal information, such as address, phone number in some cases, relatives, if I remember correctly.
Shilpa: Yeah.
Konrad: So they collect quite a bit of information about you and make it really easy for advertisers, but also other people to look it up. For FV,
Shilpa: Yeah.
Konrad: yeah, so that’s a bit scary. How does it work from the user perspective? Let’s say you want to use PIR. You want to try it out. How is the experience looking? Could you share with folks how it looks UI-wise and what you have to do to try it out?
Shilpa: Yeah, so today, actually in some platforms, you can even go try it out for free, part of the experience on macOS or Windows, where you can go ahead and put in your personal information, some basic stuff, like your name and date of birth and city and you will be able to go see, we will go ahead and do a search on the various different data brokers that we support. And it should be able to show you all of the different brokers that have your information, the information that they do have. And the first time, if you see this actually, it’s a little bit, it’s actually quite a bit scary to see how much of your information is out there. And literally anyone can do a search and pull up a lot of your information about you, your family and stuff. So yeah, once we do the search and then we’ll go ahead and go to those, once we know that your records are there on those specific brokers, we go ahead and submit requests to opt you out.
Konrad: Thank
Shilpa: And then there may or may not be some steps in between all of which we will take care of for you. And then usually there’s a delay from a couple of days to a few weeks maybe. Then if the record is removed, we will go ahead and do a scan again and make sure that your record is removed. And then we’ll report to you that that record’s been flagged as removed. And not just that, we also do periodic scans to make sure we do like maintenance scans to make sure that your record is removed and it continues to be removed. Because the way that data brokers work is once they remove it, they’re not very intentional or deliberate about these records. They’re just collecting and gathering from whatever sources they’re getting. So if it’s removed, doesn’t mean it stays removed. Sometimes it goes, sometimes it comes, sometimes it’s merged from other sources. So it’s important to continue to keep, keep removing it.
Konrad: Yeah. And we keep on adding new brokers, right? Meaning, you keep the PIR running, we will add new brokers and we will opt you out out of those brokers. So it’s not only like insurance that your records don’t pop up on the same brokers, but as we find new brokers starting business and appearing on the web, we will try to opt you out from those too, right?
Shilpa: Yeah, yeah. It’s the whole world of data brokers is very fluid. There’s new ones coming up. There’s mergers happening. There’s some of them are going defunct. Some of them are temporarily defunct. They come back. So we’re all constantly keeping an eye out for how things are changing, fixing it and adjusting for the changes they’re making, as well as constantly looking out for new ones that are coming up, seeing how we can support removals on those sites.
Konrad: That’s right. And I guess that’s our biggest challenge, right? Like fighting with this ever-changing websites and brokers kind of even fighting back on purpose, trying to kind of avoid our product, like make it not work. Would you say it’s the main challenge? Do we have like some other challenges?
Shilpa: I would say that is pretty much the main challenge. They are deliberately or not constantly changing things on their side. And since they don’t have a very clean, clear request that we can make to them, just to say, here’s my record, go remove it. But it involves usually going through and scraping the website. It’s not a very clean process, which means that we have any minor changes on their end leads to us breaking things on our side. So we have to constantly keep an eye out and monitor, make changes, fix. So this constant whack-a-mole, I think, is one of the biggest challenges. It’s also a good engineering challenge. So we are always thinking of various different ways to innovate and try to be as flexible as possible and see how we can actually get ahead of them. That actually is like the main engineering challenge related to this product for us.
Konrad: Yeah, we both like a challenge. That’s why we are on this team, right? OK, let’s see. So we are both engineers.
Shilpa: Absolutely.
Konrad: So maybe let’s do some shop talk and talk how this work under the hood. So yeah, can you tell us a bit about that?
Shilpa: Yeah, so doing a little bit of engineering talk, normally you would imagine that, you know, like we take the personal information from the users and then we go and make a nice clean API call directly to data brokers and they give us a response and we can say either success or failure. But that is actually not the case. Data brokers do not make our lives easy. What we actually have to do is basically what you would do if you were to manually go ahead and remove it. You would go to their website, type in your information, do a search, see if the record is found, and then go to another form. Again, type in an opt out request, and in between there may be various different stages. Depending on the data brokers, you may have to solve some captchas. You may have to get a confirmation code from the email or click on a link in the email. So we do all of that in the entire lifecycle of search and remove a record. So yeah, and I think one unique thing to the way we do versus our competitors is that when we take your personal information, it stays on your device. We do not send it off to our servers where it stays forever and potentially vulnerable to various other attacks. So we’re very particular about keeping all of your personal information within your device. So we just make the request from your browser within your device and keep it in your device so that it’s a lot less liable to be breached in any other way from server side.
Konrad: Right, because that’s our approach. Because the competition existed when we first started building this product, but we didn’t just copy the existing model. We thought it through and decided that what they are doing, it’s not going to work for us because we don’t want to receive personal information from our users. We don’t want to process that. We don’t want to deal with it. And so we came up with this model where everything happens on device, which has some pros. It has some cons. But yeah, we are making it work. And yeah, as you said, it’s challenging. But it’s the right thing to do to keep your data private. OK, so we talked a bit about what’s under the hood. We talked about the challenges, how those brokers are fighting back and by also going offline and changing their websites all the time and changing those forms. And yeah, so that has been our focus for a bit. You started talking, you mentioned like a bit about like flexibility, but we also do monitoring. We do other things like how we are adapting to this like well, complex environment.
Shilpa: Yeah, so constantly we have various different kinds of monitoring set up. So we’re always keeping track of how many of the brokers are responding correctly to the scans. We’re seeing what percentage of the scan attempts are going through, what percentage of the opt out request submissions are going through. Anytime we see more than a reasonable amount of variation in these, we always have one person that is on call watching out for any of these anomalies, any anomalies come up, we go do investigation to see what fixes need to be done. And also, we’re also trying to see how to reduce the amount of reactivity in terms of fixing this. We are trying to see how can we incorporate as much flexibility into our products, so there’s less reacting and more, we’d be staying on top of things when in case things do break. Currently, major part of our roadmap is actually focused on how to be more flexible as a product. And we’re also exploring to see how we can like incorporate AI, it’s fairly well suited to monitor things and maybe even try to attempt to fix it even before a human takes a look at it. So we’re pretty excited to kind of start working on all of that in the near future.
Konrad: Cool. Yeah, what’s next for PIR? What’s on our road?
Shilpa: I think one of the biggest requests from our customers has been to add the feature, make it available on mobile. We’ve had it on desktop for a while now on macOS and Windows. And there’s been a lot of requests to add it on iOS and Android because as expected, not everyone wants to run this on their desktop or has access to one all the time. So yeah, we’re currently working on that. We are excited to have it released on mobile, which should be coming up in the near future. There are challenges associated with making this work on mobile, especially since we don’t just send everything off to a server that will take care of things for us. On mobile, we need to make sure we have the connectivity. We don’t drain the battery. We are using the resources reasonably and efficiently. And we may not get a whole bunch of time from either the iOS or the Android operating system to run things as much as often as we want. On a desktop, we don’t have to worry about Wi-Fi connectivity or battery as much, but on mobile, there’s all of these additional challenges that we have to take care of. The team has been working for the last year almost to figure out ways around this and it’s a very good challenge to be working on.
Konrad: Yeah, the word ‘challenge’ comes up a lot when we discuss PIR. But thankfully, we have some great people on board and the work is moving smoothly forward. So yeah, look out for PIR on mobile. OK, and if anyone after hearing this discussion wants to try it out, even without becoming a subscriber, because this is like a subscriber feature. You mentioned they could do freemium, right? So how one does that? Where do you find it in our?
Shilpa: Yeah, so in the browser, if you’re not a subscriber, you should be able to find it in Settings, or sometimes you might just get a pop-up in a new tab when you open asking you to go try a free scan. So do try that. And if you live in the US, it’s going to be quite shocking. I will warn you of that to see how much of your personal information is out there. So it’ll ask you for a few things about your personal information, three to four fields. If you fill that in, it’ll do a quick search across the number of data brokers we support vary time to time because of the various reasons I said how they’re coming up and down all the time. But we support a maximum of about 60 currently. So we go do a search across all of these brokers. And you’ll be able to see how much of your information and even your family’s information is associated with you. And if that seems like a reasonable thing for you to invest in, then you can go ahead and subscribe in order to go for, allow us to go make the request to submit a request to remove your records from those sites.
Konrad: Perfect. I think we covered most of this stuff. Any kind of a closing thoughts, anything that we should have mentioned?
Shilpa: I think the only thing I would say is, yeah, it is a little bit scary to go see your personal information everywhere, but rest assured that we are constantly scanning and constantly monitoring. And a lot of the users also often of these products often mention how the number of scam calls that they get on their phones also goes down. So it’s really a very worthwhile product to be investing in. Check it out, check out the free scan version and we’d love to hear from you as to how things are working. And we’re also making a lot of improvements to our UX to make it more user-friendly and to make it more clear as to what’s going on behind the scenes and what stage of the various processes that we mentioned. End to end, each of your records is in with each of the brokers. So we’re excited to make all of those user facing improvements in the future.
Konrad: Perfect. Thanks so much, Shilpa. That was great. I really appreciate this discussion. And yeah, that’s all folks. See you in the next episode, I guess.
Shilpa: Thank you.
This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit insideduckduckgo.substack.com
Fler avsnitt
Visa alla avsnitt av Inside DuckDuckGoInside DuckDuckGo med DuckDuckGo finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.