Is Scraping Facebook Groups Legal in 2026?
Meta's terms ban automated collection whether or not you're logged in. What it means for freelancers finding clients in Facebook groups, every claim sourced.
Somewhere in a local business group this morning, someone posted eleven words: “Can anyone recommend someone who actually builds decent websites?” Four people will see it in time to matter. If you are a web designer in that group, you are one of them — and if you are not watching the group at 9:14am, you are not.
The obvious fix is to stop watching by hand. So you search for a way to automate it, and within two clicks you are on the pricing page of a Facebook group scraper, reading a headline that says scraping public data is legal, backed by a court case, with a green tick beside it. What you are reading is a vendor page. Every page ranking for this question today is written by someone selling the tool the answer is about.
This guide is not that, though it is not disinterested either: we sell a Chrome extension in this space, and there is a section near the end that applies each of these rules to our own product before it applies them to anyone else’s. What follows is what the primary documents actually say — Meta’s terms as of 1 January 2025, the court file rather than the commentary about it, Reddit’s eight platform rules read in full, and the GDPR articles that apply whether or not any platform ever notices you. Every claim below links to the source it came from, and where a claim is contested, undated, or turned out to be wrong, it says so.
Is it legal to scrape Facebook groups?
Scraping a Facebook group with an automated tool breaches Meta’s Terms of Service, which since 1 January 2025 apply “regardless of whether such automated access or collection is undertaken while logged-in to a Facebook account”. Breaching terms is a contract matter, not automatically a crime. Storing personal data is a third, separate question that GDPR answers.
Almost all the confusion in this topic comes from three different questions wearing one word. Is it against the rules is a contract question, answered by Meta’s terms, and the answer is yes. Is it a crime is a statutory question, answered in the United States by the Computer Fraud and Abuse Act, and the answer after 2021 is usually no for public data. Am I allowed to keep what I collected is a data-protection question, answered by GDPR, and it is the one most guides skip entirely — even though for a European freelancer it is the only one with a regulator attached.
Those three answers do not move together. You can be in breach of Meta’s terms and commit no crime. You can commit no crime and still owe a stranger a notification under Article 14. The rest of this guide takes them one at a time.
The rules that actually apply, in one box
- Meta’s Terms of Service, effective 1 January 2025, prohibit accessing or collecting data by automated means without prior permission — expressly including while logged in. (Meta Terms of Service, §3.2)
- Meta’s Automated Data Collection Terms, effective 7 October 2024, define the covered mechanisms broadly: “web scrapers, bots, robots, spiders, crawlers, user-agents, and other automated or programmatic mechanisms”. (Automated Data Collection Terms)
- Accepting those terms is not permission. Permission “must be obtained separately through Meta’s formal authorization process”. (Automated Data Collection Terms)
- Meta’s spam standard restricts activity “either manually or automatically, at very high frequencies” — and separately at lower frequencies where signals of inauthenticity are present. (Meta Community Standards: Spam)
- Meta v. Bright Data (23 January 2024) held Meta’s then-current terms did not reach logged-off scraping of public data — and Meta expressly waived its right to appeal. (Order, N.D. Cal. Dkt. 181)
- Van Buren v. United States (2021) narrowed the CFAA to a “gates-up-or-down inquiry” — but expressly declined to decide whether contracts and policies count as gates. (Supreme Court slip opinion)
- LinkedIn’s User Agreement names “crawlers, browser plugins and add-ons” among the means you may not use “to scrape or copy the Services”. (LinkedIn User Agreement, §8.2)
- Reddit’s eight platform rules contain no self-promotion ratio at all. No 9:1, no 90/10, no percentage. (Reddit Rules)
- GDPR Article 6(1)(f) requires a documented three-part balancing test, and Article 14 requires notice for data you did not get from the person — public availability is not one of the four exemptions. (Art. 6 · Art. 14)
What Meta’s terms actually say — and the sentence most guides miss
The single most-cited fact in this entire topic is out of date, and the pages repeating it have not noticed.
In January 2024, Judge Edward M. Chen of the Northern District of California granted summary judgment to Bright Data against Meta. The reasoning is narrower than its reputation: Meta’s terms governed “your use” of Facebook and Instagram, and the court found that “Bright Data did not ‘use’ Facebook or Instagram when it engaged in logged-off scraping of public data”, because “mere visitors without a Facebook or Instagram account do not see the Terms and are not bound”. That is a ruling about Meta’s drafting, by one district judge, on one company’s conduct. It is not a holding that scraping is lawful. Meta then stipulated to dismiss its remaining claim and expressly waived its right to appeal — which is what you do when you would rather fix the contract than litigate it.
Meta fixed the contract. The Automated Data Collection Terms took effect on 7 October 2024, defining Automated Data Collection as the use of “automated or programmatic tools capable of navigating or indexing the surface-layer of the World Wide Web” and requiring express written permission obtained “separately through Meta’s formal authorization process”. Then, on 1 January 2025, the main Terms of Service took effect carrying the sentence that closes the gap the 2024 ruling opened:
“You may not access or collect data from our Products using automated means (including by engaging in Automated Data Collection as defined in the Automated Data Collection Terms) without our prior permission, or attempt to access data you do not have permission to access, regardless of whether such automated access or collection is undertaken while logged-in to a Facebook account.”
That final clause is not a gloss or an interpretation. It appears twice in the current terms — once for access and collection, once again for selling or licensing what you obtained. If you are reading a 2026 guide that argues from Bright Data without mentioning either the October 2024 or the January 2025 document, you are reading advice about a contract that no longer exists.
Breaking a website’s terms is not the same as breaking the law
Three layers govern this, they have different enforcers, and the consequences do not resemble each other.
| Layer | What governs it | Who enforces | What a breach means |
|---|---|---|---|
| Contract | Meta, LinkedIn, Reddit and X terms of service | The platform, unilaterally | Account and asset removal, revoked permission, forced deletion, civil suit |
| Criminal statute | The CFAA in the US; equivalents elsewhere | Prosecutors | Narrow after Van Buren — a “gates-up-or-down” test, not a policy-violation test |
| Data protection | GDPR Articles 6 and 14 in the EU and UK | Data protection regulators | Regulatory action against you, independent of what any platform thinks |
The criminal layer is the one that shrank. In Van Buren the Supreme Court held that a person “exceeds authorized access” only when they obtain information “located in particular areas of the computer — such as files, folders, or databases — that are off-limits to him”, and framed liability as “a gates-up-or-down inquiry”. The Court was blunt about why: “If the ‘exceeds authorized access’ clause criminalizes every violation of a computer-use policy, then millions of otherwise law-abiding citizens are criminals.”
But read footnote 8, which scraping guides invariably drop. The Court expressly declined to decide “whether this inquiry turns only on technological (or ‘code-based’) limitations on access, or instead also looks to limits contained in contracts or policies.” The question of whether a terms-of-service violation can itself be a gate is open. The Ninth Circuit in hiQ went further for public pages — “that computer has erected no gates to lift or lower in the first place” — but even that was a preliminary-injunction ruling about “serious questions”, not a merits judgment.
And here is the part that almost never travels with the hiQ headline: hiQ lost. Back in the district court in November 2022, the same Judge Chen found that “the relevant language of the User Agreement unambiguously prohibits hiQ’s scraping and unauthorized use of the scraped data”, and that hiQ had breached it through workers using fake accounts. The case ended in a consent judgment against hiQ reported at $500,000 with a permanent injunction requiring it to stop scraping and destroy the derived data — figures reported by law firms covering the settlement rather than read off the court’s own document, so treat them as secondary.
| Case | What it decided | What it did not decide |
|---|---|---|
| Van Buren (2021) | “Exceeds authorized access” means going into off-limits areas, not misusing access you have | Whether contracts or policies can themselves be “gates” — footnote 8 leaves it open |
| hiQ (9th Cir. 2022) | Public pages raise serious questions about invoking the CFAA at all | Anything about contract — and hiQ then lost on contract in 2022 |
| Meta v. Bright Data (2024) | Meta’s then-current terms did not bind a logged-off scraper with no account | Whether scraping is lawful; Meta rewrote the terms and waived appeal |
The pattern across all three: logged-off, no account, public-only, no fake identities is the defensible posture. An account, plus terms you accepted, plus authenticated access is where liability actually lives — which is precisely the position a freelancer scanning groups they belong to is in.
”Scraping a group” means three different things
People use one word for three acts with wildly different exposure, and conflating them is why the advice online is useless.
Reading posts in a group you have joined is what your browser does when you scroll. The terms do not restrict reading; they restrict automated access and collection. The question is never “did a human eye see it” — it is whether a programmatic mechanism did the fetching, and at what rate.
Copying posts out of a private group adds a fourth layer the legal analysis usually misses: the group’s own rules, and the admin who enforces them. Membership is not a licence. A private group is a room you were let into on terms, and being removed from it costs you the channel regardless of what any court in California thinks. This is the layer that actually ends most freelancers’ access, and it is the one no terms-of-service analysis will warn you about.
Exporting the member list or email addresses is the highest-exposure act in this area, and it is the feature most scraper products lead with. It is bulk personal data, collected without the person’s knowledge, for a purpose they never contemplated. There is no plausible reading of the Article 6(1)(f) balancing test in which a stranger reasonably expects their group membership to become a cold-email list. If you take one operational rule from this guide: do not export member lists. Everything else here is a matter of degree; that one is not.
Where does your setup sit? A six-rung ladder
The useful question is not “is scraping legal” but “which rung am I on”. Every rung below is more exposed than the one above it, and the jumps are not evenly sized.
Note what the ladder does not say. Rung 2 is not labelled “compliant”. A tool that programmatically reads pages is doing programmatic reading, and Meta’s clause is written broadly enough to reach it. What changes between rung 2 and rung 4 is not whether the clause is arguably engaged — it is frequency, volume, what gets stored, whose account is doing it, and whether anything is published without a human. Those are the variables enforcement actually keys on, which is the subject of the next section, and they are also the variables that decide your ban-risk exposure in practice.
What Meta actually enforces is frequency, not automation
Read Meta’s spam standard and the enforcement trigger is not the word “automated” at all. The first thing it says it does not allow is:
“Posting, sharing, engaging with content or creating accounts, groups, Pages, events or other assets, either manually or automatically, at very high frequencies.”
Either manually or automatically. Doing it by hand is not a defence, and using software is not the offence — the noun that carries the weight is frequency. The bullet immediately after it closes the obvious loophole: “We may place restrictions on accounts that are acting at lower frequencies when other indicators of spam (e.g. posting repetitive content) or signals of inauthenticity are present.”
That is the whole enforcement model in two sentences: volume, or a lower volume plus something that looks fake. It explains a pattern that otherwise makes no sense — why a careful person using a tool goes years untouched while an enthusiastic person doing everything manually gets restricted in a fortnight. We wrote up the behavioural side of this separately in the five things that actually get accounts flagged, and it is why every account in our own product carries a live safety score with visible daily limits rather than a “compliant” badge.
Meta’s enforcement powers, incidentally, are worth reading once in the original. It may act “at any time, including while we investigate you, with or without notice”, and its remedies include revoking permission, requiring you to delete collected data, and “terminating other agreements with you or your ability to use Meta Company Products”. There is no appeal ladder in that sentence.
Reddit, X and LinkedIn: the same question, three different answers
Facebook is not the only room, and the other three answer this question quite differently. If you are working these channels, the platform-specific guides go deeper — Reddit, LinkedIn and X each have their own economics — but here is the compliance picture in one view.
| Platform | What the terms say | The licensed route | The distinctive trap |
|---|---|---|---|
| Rule 2: participate authentically “in communities where you have a personal interest”; no ratio anywhere | Reddit’s paid Data API | The 9:1 rule people quote does not exist as a platform rule | |
| X | Automated access is metered and priced, not prohibited | Pay-per-use API: $0.005 per post read | A post containing a URL costs $0.200 to create — 13× a plain one |
| Bans the means “to scrape or copy the Services”, and bans bots for engagement | LinkedIn’s official APIs, partner-gated | Browser plugins are named explicitly in the list of prohibited means |
Reddit’s 9:1 rule is folklore, and it is worth being precise about why. We read all eight platform rules in full and searched for every form of the ratio — 9:1, 90/10, “one in ten”, any percentage. There is none. The enforceable standard in Rule 2 is qualitative: “Participate authentically in communities where you have a personal interest, and do not spam or engage in disruptive behaviors”. The closest official text is a moderator-advice article, How do I keep spam out of my community?, which says “some communities” abide by “the 10% rule” and ends with “It is ultimately up to you and your team to decide what works best for your community.” So the number Reddit actually publishes is 10%, not 9:1, and it is delegated to moderators rather than imposed. Individual subreddits then set real, enforceable versions of it — r/marketing’s karma and account-age gate is a live example — which is why reading the sidebar beats reciting a ratio.
X priced the problem instead of prohibiting it. Its published pay-per-use rates are $0.005 per post read, $0.015 to create a post, and $0.200 if that post contains a URL — a thirteen-fold surcharge on linking out, which tells you everything about what X thinks link-dropping is worth. Pay-per-use is capped at three million post reads per billing cycle. Before assuming scraping is the cheap option, price the licensed one: for most freelancers monitoring a handful of searches, the API cost is not the barrier people assume.
LinkedIn is the one that names browser extensions. Section 8.2 of the User Agreement, effective 3 November 2025, says you will not “Develop, support or use software, devices, scripts, robots or any other means or processes (such as crawlers, browser plugins and add-ons or any other technology) to scrape or copy the Services, including profiles and other data from the Services”. Two things are true about that sentence and both matter. Browser plugins are named — that is not a competitor’s spin, it is the text. And the prohibited purpose is “to scrape or copy the Services”: plugins appear inside a parenthetical list of example means, not as a standalone ban on extensions. A separate item in the same list bans using “bots or other unauthorized automated methods” to “send or redirect messages, create, comment on, like, share, or re-share posts, or otherwise drive inauthentic engagement” — which is the clause that makes auto-connect and auto-DM tools unambiguously non-compliant, and the reason we decided never to build automated DMs on any platform.
GDPR applies even if the platform never notices
Everything above is about what platforms permit. This section is about what a regulator requires, and it is independent — Meta’s indifference is not a defence, and Meta’s permission would not be one either.
If you save a post that identifies a person, you are processing personal data and you need a lawful basis. In practice that means Article 6(1)(f): processing “necessary for the purposes of the legitimate interests pursued by the controller… except where such interests are overridden by the interests or fundamental rights and freedoms of the data subject.” The ICO breaks that into three named tests — a purpose test, a necessity test and a balancing test — and the balancing test is where scraping usually dies. The ICO’s own phrasing: a person’s interests “are likely to override yours if they wouldn’t reasonably expect you to use their information”. Someone asking a group for a plumber reasonably expects a plumber to reply. They do not reasonably expect to appear in a CSV.
Then there is Article 14, which is the one that surprises people. Where personal data has not been obtained from the data subject, you must tell them you hold it. Its exemptions are an exhaustive list of four — the person already has the information, provision is impossible or requires disproportionate effort, disclosure is laid down in law, or professional secrecy applies. “It was publicly available” is not among them. Article 14 in fact runs the other way: 14(2)(f) obliges you to disclose “whether it came from publicly accessible sources”.
Two dated markers worth knowing. The European Data Protection Board adopted Guidelines 03/2026 on web scraping in the context of generative AI on 8 July 2026, open for feedback until 30 October 2026 — aimed at AI training rather than lead generation, but it states the principle cleanly: “The GDPR applies to web scraping when it includes personal data processing operations.” And the Irish Data Protection Commission’s €265 million fine on Meta, announced 28 November 2022, is routinely miscited in this debate: it was imposed on Meta, for failing to design against scraping under Articles 25(1) and 25(2), not on a scraper for scraping. It is evidence that regulators take the harm seriously. It is not a precedent about you.
None of this makes finding clients in public communities unlawful. It makes hoarding people unlawful-ish and answering people fine — which, conveniently, is also what works commercially. Our note on where your data lives sets out the same distinction from the other side of the product.
Why people buy scrapers anyway
Nobody buys a group scraper because they want a database. They buy it because outbound stopped working and volume feels like the only lever left.
The volume lever is closing. Google’s sender guidelines require SPF, DKIM and DMARC for anyone sending 5,000 or more messages a day to personal Gmail accounts, and tell senders to keep spam rates “below 0.10% and avoid ever reaching a spam rate of 0.30% or higher”. Microsoft matched it: since 5 May 2025, Outlook.com rejects non-compliant mail from high-volume domains outright with 550 5.7.515 — the widely-quoted “it goes to junk” line was superseded by an update on 29 April 2025 and is no longer what happens.
And the returns are thin at the top of that funnel. Belkins analysed 7,530,489 cold emails sent across its 2025 client campaigns and reported an average reply rate of 0.45% — theirs is a lead-generation agency, so flag it as vendor-published, and note carefully that they changed the denominator to replies divided by total emails sent, which is why the figure looks so much lower than the 3–8% numbers elsewhere. Those numbers are measuring a different thing; do not put them in one table.
Meanwhile the buyers have moved. Gartner’s research finds that 75% of B2B buyers prefer a rep-free sales experience. 6sense’s 2025 Buyer Experience Report, from nearly 4,000 buyer responses — their research, and they sell intent data, so weigh it accordingly — found that 94% of buyers ordered their shortlist by preference before engaging with any seller, and that the pre-contact favourite goes on to win 77% of the time. Their measure of how far into the journey buyers get before talking to a seller fell to 61% in 2025, from about 69% in the two prior years.
Put together, that is an argument against list-building rather than for it. A bigger list is not the constraint. Being present, visibly and early, in the thread where the shortlist forms is the constraint — which is what social listening is for, why where freelance clients actually come from looks the way it does, and why it is worth working out your real cost per lead before buying anything at all.
A seven-step audit of the setup you already have
If you are already running something, this is the order to check it in. It maps one-to-one onto the steps in this article’s structured data, so an assistant summarising this page gets the same list you do.
Step 1 — Write down the act, not the tool name. One sentence, no brand names: “a script opens group pages and copies every post” is a different sentence from “a browser tab reads posts my own session already renders”, and terms attach to the sentence.
Step 2 — Separate reading from storing. Which does yours do, which fields does it keep, and for how long? Nearly every hard question below is about the storing half.
Step 3 — Measure frequency, not automation. Meta restricts activity “either manually or automatically, at very high frequencies”. Count your real requests per hour per account and compare it to a fast human.
Step 4 — Find anything that acts without you. Auto-DM, auto-comment, auto-connect, scheduled posting. If any part of your stack sends while you are asleep, stop there — that is the finding, and it is the one platforms enforce hardest.
Step 5 — Check whose account is doing it. Your own logged-in session is one thing. A managed, rented, shared or purchased account is a different and more serious thing, and fake accounts are what sank hiQ.
Step 6 — Decide whether you are exporting a list. Member lists and email addresses are the top of the exposure ladder. If you are exporting them, do step 7 before you do it again.
Step 7 — Write the balancing test down. Purpose, necessity, balance, on one page, plus how you would answer an Article 14 request. It takes an afternoon and it is the entire obligation for a business this size.
Then keep the buying-intent keyword set sharp and the group choice deliberate. A precise phrase list in five good rooms beats bulk collection from fifty, and it is the version that survives an audit.
Where ClientRadar sits — and where it’s the wrong choice
We make a Chrome extension that watches Facebook groups, subreddits, X and LinkedIn for buying-intent posts, so treat this section as what it is: an interested party grading its own homework in public. Here is the grade.
ClientRadar sits on rung 2 of the ladder above. It reads through your own logged-in browser session rather than a server-side scraper or exported cookies, moves at a human pace with daily caps, cooldowns and quiet hours, and never posts, comments, connects or messages without your explicit tap — there are no auto-DMs and no managed accounts anywhere in the product. Leads, notes and pipeline stay in your own browser by default, and what leaves the device and when is documented rather than implied. Each matched post is scored 0–100 with the reasoning shown, and you approve every reply.
What we will not claim is that this puts us outside Meta’s clause. Section 3.2 says “regardless of whether… while logged-in”, and the Automated Data Collection Terms define the covered mechanisms broadly enough to reach any programmatic reader — including ours. Anyone in this category telling you their architecture makes them compliant is selling you a reading of the terms that the terms do not support. What differs between rung 2 and rung 4 is the enforcement surface and the data-protection exposure: frequency, volume, whose account, whether anything publishes without you, and what is retained. Those are the variables we cap, disclose and score, and they are the honest basis for the comparison — not a claim of permission we do not have.
The same discipline applies to LinkedIn. Its User Agreement names “browser plugins and add-ons” among prohibited means, and we are a browser plugin. The clause’s operative purpose is “to scrape or copy the Services”: we read what your session renders and never bulk-copy or republish LinkedIn’s data, which is a real distinction — and it is also one LinkedIn adjudicates, not us. That is why LinkedIn support is read-and-draft only, gated behind the top plan, and why we tell you the risk instead of a green tick.
Where it is the wrong choice, plainly. If you need bulk member lists or exported email addresses, we do not do that and will not add it. If you want to message people who never asked, we are the wrong product — try the comparisons and read the risk ladder first. If you want something posting while you sleep, that is rung 4 and above, and tools that post from managed accounts exist. If you just want free keyword alerts and will do the rest yourself, F5Bot is genuinely good and costs nothing — we say so on its comparison page too. And if none of this is your bottleneck, the free tier is a preview rather than a trial, and the pricing is public.
Methodology and sources
Every legal claim in this article was checked by fetching the primary document — the terms page, the policy, the signed court order, the regulation text — rather than a summary of it. Where a source blocked automated fetching, the page was rendered in a browser and the text read from the page itself; where a document could not be reached at all, the claim is either absent from this article or marked as secondary in the sentence that carries it. Vendor-published statistics are flagged as vendor-published where they appear, not in a footnote.
Three corrections made during research. First, the widely-repeated claim that Meta’s Automated Data Collection Terms took effect on 1 January 2025 is wrong: those terms are stamped effective 7 October 2024, and they contain no logged-in/logged-out language at all. The 1 January 2025 date and the “regardless of whether… while logged-in” clause belong to the main Terms of Service — two documents, two dates, and the distinction matters because it is the ToS, not the ADCT, that answers the Bright Data problem. Second, Reddit’s “9:1 rule” is not a Reddit rule; the eight platform rules contain no ratio, and the only official number is a “10% rule” described in a moderator-advice article as something “some communities” adopt. Third, the Gartner statistic that B2B buyers spend “17% of their time with suppliers” — which circulates widely and which we have cited ourselves — is not on Gartner’s page, and has been left out of this article accordingly.
How legal claims were checked. Court holdings were read from the filed documents on CourtListener and the courts’ own PDF servers, not from law-firm commentary; the one exception is the hiQ consent judgment, where the $500,000 figure and injunction terms come from law-firm write-ups because the docket entry was not retrievable, and the sentence carrying them says so. Platform terms are quoted from the live pages at the URLs linked, with the effective dates those pages print.
What this is not. This is a reading of published terms, public rulings and regulation text as they stood in August 2026, written by a software company. It is not legal advice, the case law here is thin and largely interlocutory, and the answer changes with your jurisdiction. If real money is riding on the question, pay a lawyer in the relevant country.
Conflict of interest. ClientRadar sells a product in the category this article assesses. The disclosure section above places our own product on the same ladder, by name, and states the clause we cannot claim to sit outside.
Platform terms and policies. 1. Meta Terms of Service (effective 1 January 2025) · 2. Meta Automated Data Collection Terms (effective 7 October 2024) · 3. Meta Community Standards: Spam · 4. Reddit Rules · 5. Reddit: How do I keep spam out of my community? · 6. LinkedIn User Agreement (effective 3 November 2025) · 7. X API pricing · 8. Meta Graph API documentation
Court decisions. 9. Van Buren v. United States, 593 U.S. (2021) · 10. hiQ Labs v. LinkedIn, 9th Cir., 18 April 2022 · 11. hiQ Labs v. LinkedIn, N.D. Cal., 4 November 2022 · 12. Meta v. Bright Data, summary judgment order, 23 January 2024 · 13. Meta v. Bright Data, stipulated dismissal and waiver of appeal · 14. Quinn Emanuel client alert on Bright Data · 15. Eric Goldman’s analysis · 16. Law-firm coverage of the hiQ consent judgment
Data protection. 17. GDPR Article 6 · 18. GDPR Article 14 · 19. ICO guidance on legitimate interests (last updated 23 March 2026) · 20. EDPB Guidelines 03/2026 on web scraping in the context of generative AI · 21. Irish DPC decision in the Facebook “Data Scraping” inquiry
Outreach and buyer economics. 22. Google email sender guidelines · 23. Outlook high-volume sender requirements · 24. Belkins B2B cold email response rates, 2026 study (vendor) · 25. Gartner B2B buying journey research · 26. 6sense 2025 Buyer Experience Report (vendor)
Terms, rulings and guidance are current as of August 2026 and will be revised here as the primary sources change. If you want the practical rather than the legal version of this, start with how to find clients in Facebook groups, finding clients on Reddit without getting banned, or the complete guide to social listening. The tool comparison prices ten products in this category at source, and the glossary pins down the adjacent terms.
Quick answers
- Is it legal to scrape Facebook groups?
- Automated collection of data from Facebook without Meta's prior written permission breaches Meta's Terms of Service, which since 1 January 2025 say so 'regardless of whether such automated access or collection is undertaken while logged-in to a Facebook account'. Breaching terms is a contract matter, not automatically a crime. Separately, if you store personal data about people in the EU or UK, GDPR applies whether or not Meta ever notices.
- Is there a legal way to get Facebook group data?
- Two, and both are narrow. You can ask Meta for express written permission through its formal authorisation process — the Automated Data Collection Terms state plainly that accepting the terms is not itself that permission. Or you can use Meta's own Graph API within its documented scopes. Neither route gives a freelancer bulk access to group posts or member lists, which is why the scraper market exists at all.
- What actually happens if you scrape Facebook?
- Meta's enforcement is contractual and unilateral: it can revoke permission, require deletion of collected data, restrict or remove accounts, groups and Pages, and terminate other agreements with you — 'at any time, including while we investigate you, with or without notice'. Civil action is possible but rare against individuals. The regulatory track is separate and runs through your own data-protection obligations, not Meta's.
- Is using a Chrome extension on LinkedIn against LinkedIn's terms?
- LinkedIn's User Agreement lists 'crawlers, browser plugins and add-ons or any other technology' among the means you may not use — but the prohibited purpose is 'to scrape or copy the Services, including profiles and other data from the Services'. Browser extensions appear inside a list of examples, not as a standalone ban. The clause turns on bulk extraction and copying, which is a distinction of degree that LinkedIn, not the extension's author, adjudicates.
- Is Reddit's 9:1 self-promotion rule real?
- No. Reddit's eight platform rules contain no ratio at all — no 9:1, no 90/10, no percentage. The nearest official text is a moderator-advice article which says 'some communities' abide by a '10% rule' and that it is 'ultimately up to you and your team to decide what works best for your community'. It is a per-community convention Reddit hands to moderators, not a platform rule you can breach.
- Do I need a GDPR legal basis to save a lead I found in a public post?
- Yes, if the post identifies a person and you store it. Most freelancers rely on legitimate interests under Article 6(1)(f), which the ICO breaks into a purpose test, a necessity test and a balancing test. The balancing test is the one that bites: interests are likely to override yours if the person 'wouldn't reasonably expect you to use their information'. Write the assessment down — it is a short document, not a legal project.
- Do I have to tell someone I collected their data from a public post?
- Usually yes. GDPR Article 14 covers personal data not obtained from the data subject, and its exemptions are an exhaustive list of four — the data subject already has the information, provision is impossible or disproportionate, disclosure is laid down in law, or professional secrecy applies. 'It was publicly available' is not among them. Article 14 actually runs the other way: you must disclose whether the data came from a publicly accessible source.
- What should I use instead of a Facebook group scraper?
- For finding clients, the honest answer is that you rarely needed the bulk data. What converts is noticing one buying-intent post early and replying usefully as yourself. Free keyword alerts, saved searches and a disciplined phrase list cover most of it. ClientRadar — our product, so weigh this accordingly — automates the watching but not the sending, and never exports member lists.
- Is any of this legal advice?
- No. This is a reading of published platform terms, public court records and regulation text as they stood in August 2026, written by a software company, not a law firm. Terms change without notice, the case law here is thin and mostly interlocutory, and the answer moves with your jurisdiction. If real money or a real risk is riding on it, pay a lawyer in the relevant country.
Andras B.