Lucrezia Berto and Marcelo Estrella, IP-IT and new technologies department.
What does web scraping mean? It is a series of techniques aimed at obtaining information from web pages in an automated way, through software, commonly defined as bots, which simulate the browsing of real users.
The information obtained can be processed and used for various purposes. A large number of entities use data extracted from third-party websites through scraping for their own purposes, and others even base their value proposition on the data and other information thus obtained.
But is it legal?, can web scraping always be carried out? and can the data obtained through this technique always be used? The fact that websites and their contents are published on the Internet and available to any Internet user does not mean that they are not protected and in the public domain. Likewise, when determining the legality of the extraction of the information and its subsequent use, it is necessary to take into account a series of risks that arise in the light of European and, in particular, Spanish regulations.
The risks
The first risk that these techniques may entail is the infringement of the intellectual property rights of the owners or licensors of the websites from which the data are extracted. In particular, in the event that the structure of the data contained in these web pages meets the requirements to be considered a “database”, according to the definition of Royal Legislative Decree 1/1996, of 12 April, approving the revised text of the Intellectual Property Law (hereinafter TRLPI), there could be an infringement of copyright and/or sui generis rights over said database [1]. It should be noted that not all websites and their content meet the requirements for obtaining the protection provided by the TRLPI for “databases”, and may lack originality and/or investment in obtaining, verifying or presenting their content, as stated by the judges in the case known as Ryanair vs Atrápalo [2].
Another risk to be taken into account is the violation of the legal terms of the scraped website, which are usually included in the “legal notice” or “terms and conditions”. Also, despite the possible protection granted by the TRLPI to the content of websites, owners can contractually prohibit or limit scraping.
In addition, there are cases where website owners publish content on their websites under licences that do not require the express acceptance of the person accessing or using the content in order to be binding. This is the case for open content licences, such as Creative Commons licences, or open database licences. Such licences may limit the use of the content, excluding, for example, use for commercial purposes, or they may oblige anyone who wants to publish the data collected through scraping to do so only in an open manner. Therefore, use or publication of the data in a manner contrary to the provisions of these licences would result in an infringement of the licences on the scraped content or databases.
In case the data being extracted are personal data, these activities could lead to a breach of the Personal Data Protection Regulation, the General Data Protection Regulation and the Organic Law on Data Protection and Digital Rights Guarantees. It is relevant to consider here, that all processing of personal data must have an appropriate legal basis, and all individuals whose personal data are processed (the “Data Subjects”) must be duly informed of, among other things, the purposes of the processing and who the responsible entity is. Failure to comply with the legal basis requirement, as well as failure to inform the Data Subjects, creates a breach and is subject to infringement by the relevant competent supervisory authority.
Finally, if web scraping is carried out on websites that provide similar services to those of the company that extracts the data, this may be considered unfair competition.
Conclusions
In order to mitigate the risks that web scraping entails, and to be able to carry out these techniques in an environment that generates greater certainty and legal security, a prior analysis on a case-by-case basis is recommended in order to determine the specific risks of scraping according to the area, the type of data and the characteristics of the web pages, among other variables.
In addition to the fundamental prior analysis, there are good practices that reduce the risks derived from these techniques, such as, for example, not evading captchas and not extracting when the website has explicitly denied access. From Across Legal we can provide the necessary and appropriate legal advice to be able to use web scraping and add value to your project or your company. We have the necessary expertise and the experience that supports us.
DISCLAIMER
The information contained herein is for informational purposes only and should not be considered legal advice. An attorney-client relationship is not formed with Across Legal until the parties enter into a legal services agreement.
________________________________________
[1] The Royal Legislative Decree 1/1996 foresees requirements that the database has to meet in order to obtain copyright and/or sui generis right protection. Therefore, only if the database subject to web scraping meets these requirements can there be an infringement of copyright.




