Is Web Scraping Legal in the EU? GDPR, Database Rights and TDM (2026)
Published 2026-10-11 · Updated 2026-10-11 · By the Scrapeshop team
Web scraping is legal in the European Union when the content is lawfully accessible, the rights holder has not opted out of text and data mining in machine-readable form, no substantial part of a protected database is extracted, and any personal data is processed under the GDPR. The EU is the only major jurisdiction with a statutory text-and-data-mining right that covers commercial use, and also the one with the strictest privacy enforcement against scrapers.
This guide covers the EU-level rules: the Copyright Directive, the Database Directive, the GDPR, the AI Act, and the Court of Justice decisions that interpret them. National rules add criminal offences and unfair-competition law; the Germany guide shows how one member state layers on top. General information current as of October 2026, not legal advice. The overview across all jurisdictions is Is Web Scraping Legal? What the Law Actually Says.
Quick answer
Scraping in the EU rests on four pillars. First, Article 4 of Directive 2019/790 permits reproductions of lawfully accessible works for text and data mining by anyone, unless the rights holder has expressly reserved its rights in a machine-readable way. Second, the sui generis database right in Directive 96/9/EC prohibits extracting or re-utilising a substantial part of a database that required substantial investment; the Court of Justice has read this narrowly enough that many scraped sources, including Ryanair’s flight data, do not qualify. Third, the GDPR applies in full to personal data regardless of where it was published, and four regulators have fined Clearview AI €20 million or more each for scraping faces. Fourth, the AI Act obliges general-purpose model providers to respect TDM opt-outs and bans untargeted facial-image scraping. Contract terms and national computer-misuse laws fill the remaining gaps.
Which EU laws govern web scraping?
| Instrument | What it covers | What it means for scrapers |
|---|---|---|
| Directive 2019/790 (DSM Directive), Arts. 3–4 | Text and data mining exceptions to copyright and the database right | Commercial TDM is permitted on lawfully accessible content unless the rights holder opts out in machine-readable form. Copies may be kept only as long as needed. |
| Directive 96/9/EC (Database Directive) | Copyright in original databases; sui generis right in databases with substantial investment | Do not extract or re-utilise a substantial part. Repeated, systematic extraction of insubstantial parts that together harm the maker also infringes (Art. 7(5)). |
| Regulation 2016/679 (GDPR) | Any processing of personal data of people in the EU, wherever the scraper sits | Need an Art. 6 lawful basis, Art. 14 transparency, data minimisation, retention limits, and a DPIA for large-scale scraping. |
| Regulation 2024/1689 (AI Act) | Obligations on AI providers; prohibited practices | Art. 5(1)(e) bans untargeted scraping of facial images for recognition databases. Art. 53 requires GPAI providers to respect Art. 4 opt-outs and publish training-data summaries. |
| Directive 2013/40/EU (attacks against information systems) | Minimum rules for national illegal-access offences | Implemented nationally (e.g. §202a StGB in Germany). Triggered by circumventing access controls, not by reading public pages. |
| Directive 2005/29/EC and national unfair competition law | Unfair commercial practices and obstruction of competitors | Scraping a competitor is not unfair per se; misleading consumers or circumventing technical measures can be. |
How does the EU text and data mining exception work?
Article 4 is the most scraper-friendly provision in any major legal system, with three conditions attached:
- Lawful access. The content must be accessible without circumventing a technical measure. Open pages qualify; paywalled or login-gated content qualifies only if you hold a licence or account that permits it.
- No machine-readable reservation.A rights holder can opt out “in an appropriate manner, such as machine-readable means in the case of content made publicly available online.” A robots.txt disallow for your crawler, a TDM-reservation protocol header or meta tag, and arguably a clearly worded clause in terms of service all count. Checking for reservations before mining is part of acting lawfully.
- Retention only as long as necessary. Copies made for mining must be deleted once the analysis is complete. The exception covers the analysis, not republication or redistribution of the works.
Article 3 grants research organisations and cultural heritage institutions a broader exception for scientific research with no opt-out. The Hamburg Regional Court applied Germany’s implementation of it to LAION’s image dataset in 2024.
When does the database right stop scraping?
The sui generis right protects the maker of a database who made a substantial investment in obtaining, verifying, or presenting its contents. Four Court of Justice decisions define its reach:
- British Horseracing Board v William Hill (C-203/02, 2004). Investment in creating data does not count, only investment in collecting and verifying existing data. A fixtures list or product catalogue generated by its owner may fall outside the right.
- Innoweb v Wegener (C-202/12, 2013). A dedicated meta search engine that queried a car-listings database in real time and presented the results re-utilised the whole database, even though each query returned only a few records. Aggregators that query a source live are more exposed than one-off extractions.
- Ryanair v PR Aviation (C-30/14, 2015).Ryanair’s flight database was protected by neither copyright nor the database right, so the mandatory lawful-user rights did not apply and Ryanair could restrict use by contract.
- CV-Online Latvia v Melons (C-762/19, 2021).A job-listing aggregator that indexed and linked to CV-Online’s adverts infringes only if its activity risks depriving the database maker of the income that would let it recoup its investment. The Court weighed the public interest in search and aggregation against the maker’s interest, narrowing the right.
Practical reading: scraping facts for analysis rarely infringes; building a competing product that lives off a continuously scraped copy of someone’s curated database does.
Can I scrape personal data in the EU?
Yes, within the GDPR. Names, photos, usernames, employer, job title, and anything else that identifies a person are personal data whether or not the person published them. The European Data Protection Board’s May 2024 ChatGPT taskforce report and its December 2024 Opinion 28/2024 on AI models both treat web scraping as processing that can rely on legitimate interests only with safeguards: excluding sensitive sites and categories, honouring robots.txt, limiting collected fields, and offering an easy objection route. The Dutch data protection authority’s 2024 guidance goes further, stating that scraping personal data is “almost always” unlawful for commercial purposes without a specific, narrowly defined interest.
Enforcement is real. Clearview AI was fined €20 million each by the Italian, Greek and French authorities in 2022 and €30.5 million by the Dutch authority in 2024 for scraping facial images. Article 82 gives individuals a private damages claim; the German Federal Court held in November 2024 that loss of control over scraped data is compensable on its own.
What does the AI Act add?
Two provisions touch scrapers directly. Article 5(1)(e) prohibits placing on the market or using AI systems that create or expand facial-recognition databases through untargeted scraping of facial images from the internet or CCTV. Article 53 requires providers of general-purpose AI models to put in place a policy to comply with EU copyright law, in particular to identify and respect Article 4 opt-outs using state-of-the-art technologies, and to publish a sufficiently detailed summary of training content. The 2025 General-Purpose AI Code of Practice spells out honouring robots.txt and emerging opt-out protocols as the expected standard.
Checklist for scraping EU websites
- Confirm lawful access: public pages only, no circumvention of logins, paywalls, or deliberate blocks.
- Check robots.txt, response headers, and page metadata for a TDM reservation. Honour it.
- Extract fields for analysis; do not mirror or republish a substantial part of a curated database.
- If any field is personal data, document legitimate interests, run a DPIA for scale, publish an Article 14 notice, and set retention limits before scraping.
- Never scrape faces for recognition; never scrape special-category data without an Article 9 condition.
- Check the member state’s computer-misuse and unfair-competition rules for the targets you scrape most.
Implementation details for rate limiting, identification, and robots.txt are in Web Scraping Best Practices.
Frequently asked questions
- Is web scraping illegal in the European Union?
- No. No EU regulation or directive prohibits web scraping. The 2019 Copyright Directive expressly permits text and data mining of lawfully accessible content, subject to a rights-holder opt-out. Limits come from the database right, the GDPR, contract law, and national computer-misuse offences.
- What is the EU text and data mining exception?
- Articles 3 and 4 of Directive 2019/790. Article 3 allows research organisations to mine content they can lawfully access, with no opt-out. Article 4 allows anyone, including companies, to mine lawfully accessible works unless the rights holder has expressly reserved its rights in a machine-readable way, such as robots.txt or metadata.
- Does the GDPR apply to scraping public data?
- Yes, whenever the data relates to an identifiable person. Public availability is not a lawful basis. Scrapers need a basis under Article 6, usually legitimate interests, must inform data subjects under Article 14, and must honour objections. Clearview AI has been fined over €90 million across four EU regulators for ignoring this.
- What is the sui generis database right?
- A right under Directive 96/9/EC that protects databases whose creation required substantial investment in obtaining, verifying or presenting their contents. Extracting or re-utilising a substantial part infringes, even if no individual entry is copyrighted. It lasts 15 years and is renewed by substantial updates.
- Does the EU AI Act regulate scraping?
- Indirectly. Article 53 requires providers of general-purpose AI models to adopt a copyright policy that identifies and respects Article 4 opt-outs, and to publish a summary of training data. Article 5 bans untargeted scraping of facial images to build facial-recognition databases.
- Can I scrape a website that prohibits scraping in its terms?
- If you accepted the terms, they bind you. For databases that qualify for copyright or the database right, Article 15 of the Database Directive voids contract terms that restrict a lawful user's normal use; Ryanair v PR Aviation held this protection does not extend to unprotected databases, where contract restrictions stand.