TU Delft Research
The benchmark methodology
The benchmark methodology was developed in collaboration with TU Delft. Providers are ranked based on their response to notifications. These notifications concern hacked servers, phishing sites, command-and-control servers, spam servers, and exploit kits on their networks. The methodology takes into account the characteristics of the provider. The abuse feeds come from various sources, such as the Shadowserver Foundation, Spamhaus DBL, and Google Safe Browsing.
Goal
Almost all online service providers deal with abuse. This also applies to the hosting market, where it takes the form of hacked servers, phishing sites, malicious redirects, command-and-control servers, spam servers, exploit kits, and more.
When you as a hosting company try to combat abuse, how do you know if you are effective?
If you have ten incidents in a year — say hacked servers or spamming customers — is that a lot? A little? How many incidents do other providers in the hosting market actually have? How can you compare those numbers between companies of different sizes and with different types of services? Without answers to these questions, it is difficult to know whether you are doing well.
Over the past years, a team at TU Delft has been working on an abuse benchmark for the hosting sector. We are now sharing the most recent version with the sector. That is not to say the benchmark is perfect or without limitations — there is still uncertainty and noise in the data. What we do know is that the benchmark contains valuable information that is useful for any hosting company that wants to learn how to combat abuse more effectively. We have tested it extensively and it turns out that the benchmark can predict to a high degree how many incidents will occur in a hosting network. The benchmark is not intended for naming and shaming, but to give each individual company insight into where they stand compared to other providers.
Model
The details of the benchmark and the tests we have conducted have been published in a peer-reviewed scientific article [1]. In a nutshell, it works as follows.
-
We define a hosting provider as the entity responsible for IP ranges with hosting services according to WHOIS data — we do not use Autonomous Systems (AS) as a starting point. In an earlier study, we found that there are on average seven providers per AS. In this first version of the benchmark, we have only included providers that are members of NBIP, ISPConnect, or DHPA and for whom we could retrieve the WHOIS information. That amounts to 129 providers.
-
We then take a number of abuse feeds [2]. Per feed, we count the number of incidents observed at each provider during the period January–August 2018.
-
Large providers have more incidents than small ones because they have more customers and more infrastructure. That does not mean they are less secure. The type of services also makes a difference. We therefore collect several characteristics of the providers, such as how many IP addresses they advertise, how many domains they host, and how much shared hosting they offer.
-
The abuse and provider data are fed into a statistical model. The model looks at the number of incidents, taking into account the size of the provider and, to some extent, the type of services in the network. It goes too far to explain exactly how the model works, but it is essentially the same as an IQ test or a standardized exam. Such a model estimates how good someone is at math based on their score distribution across the different questions. In the benchmark model, each feed is essentially a test question on which the provider scores a number of points. The model then estimates which underlying capability best explains that score distribution relative to the other “students.” It also provides a margin of uncertainty around that score — some scores are fairly robust, others have a wider margin.
-
The benchmark is a score that expresses where the provider stands within the total group of 129 providers according to the model. We express this score as a percentile. A provider with a score of 20 is in the 20th percentile, meaning that 20% of all providers have more abuse than this provider and 80% have less, taking into account size and type of services. A score of 20 is therefore a poor score in terms of abuse prevention — it places you among the worst 20% of the market. We use the following simple labels to communicate these results: a score between 1–20 is poor, between 20–80 is average, and between 80–100 is good.
-
Finally, in addition to the abuse benchmark, we have also calculated a vulnerability benchmark. This follows the same steps but uses vulnerability data instead of abuse data [3]. This benchmark is also expressed using the labels poor (1–20), average (20–80), and good (80–100). It is therefore possible for a provider to perform well in terms of abuse and poorly in terms of vulnerabilities.
The abuse and vulnerability data used in the benchmark covers January through August 2018. This data is only available to a limited extent in the AbuseIO environment of abuseplatform.nl. This environment does not yet contain all feeds and only includes the most recent data. The intention is that in future iterations of the benchmark, the data will be aligned with the platform — meaning the benchmark will be based on the incidents and vulnerabilities that the provider can see in their own AbuseIO account on abuseplatform.nl.
More information
If you have questions about the benchmark or the underlying data, please contact Carlos Gañán at TU Delft:
C.HernandezGanan@tudelft.nl
+31 15 27 82216
References
[1] Publication with a scientific description of the benchmark
Arman Noroozian, Michael Ciere, Maciej Korczynski, Samaneh Tajalizadehkhoob & Michel van Eeten (2017), Inferring Security Performance of Providers from Noisy and Heterogeneous Abuse Datasets, Workshop on the Economics of Information Security (WEIS2017), La Jolla, CA. http://weis2017.econinfosec.org/wp-content/uploads/sites/3/2017/05/WEIS_2017_paper_60.pdf
[2] List of abuse feeds used
Spamhaus DBL (split into C&C, phishing, malware, spam)
Shadowserver Compromised Website Report
Shadowserver Command and Control Report
APWG
PhishTank
Google Safe Browsing
[3] List of vulnerability feeds used
Shadowserver Drone Report
Shadowserver CHARGEN Report
Shadowserver Open Memcached Server Report
Shadowserver Open Resolvers Report
Shadowserver SSL FREAK Vulnerable Report
Shadowserver SSL POODLE Report