# Available Is Not Permission. Read the License Before You Build on That Dataset

An MIT badge covers the code, not the data. Two popular collection tools, one warning between them, and the three separate permission questions every owner should answer first.

Author: J.A. Watte
Published: August 22, 2026
Source: https://jwatte.com/blog/blog-available-is-not-permission-dataset-licensing/

---

Two of the most popular data collection tools on GitHub do roughly the same kind of work. One pulls business listings with reviews and contact details. The other pulls app store data. Between them they have about 8,500 stars and 1,600 forks, both are actively maintained, and both carry the MIT license.

One of them says this, once, near the bottom of a long page: "Please use this scraping tool responsibly and in accordance with applicable laws and regulations. Unauthorized scraping may violate terms of service."

The other says nothing at all.

Neither warning is really the point. The point is that the MIT badge sitting at the top of both projects is answering a question nobody asked. MIT tells you what you may do with **the code**. It says precisely nothing about whether you may collect the data, and nothing about what you may do with the data once you have it. Those are three different questions with three different answers, and the green license badge only covers the first one.

I went and read the actual license texts, the actual platform terms, and the actual court decisions rather than the summaries. Most of what circulates about this topic is wrong in a specific and expensive direction: people cite a famous case that says less than they think, and skip the boring contract question that is the one that actually bites.

## The three questions, kept separate

Before you build any process on a dataset or a collection tool, answer these independently. Answering the first one does not answer the others.

**One: may I run this code?** That is the software license. MIT, Apache, GPL, or nothing at all.

**Two: may I collect this data?** That is terms of service, access law, and the site's own machine-readable instructions. The software license has no bearing on it whatsoever.

**Three: may I use what I collected, for what I want to use it for?** That is the data license, plus a whole separate body of law if any of it describes a person.

Almost every mistake I have watched a small business make here is answering question one and assuming it settled two and three.

## Question one: the repository itself

Here is the part that surprises people most, and it is not close to ambiguous.

**A public GitHub repository with no license file gives you nothing.** GitHub's own documentation states it plainly: without a license, default copyright law applies, the author retains all rights, and no one may reproduce, distribute, or create derivative works from it. Being able to see it, clone it, and fork it is not permission to use it.

It gets narrower. GitHub's Terms of Service, section D.5, describes exactly what making a repository public grants everyone else: a license to "use, display, perform and reproduce (by forking) Your Content **through the Service**." Those last three words carry the whole clause. That grant covers viewing and forking on GitHub. It is not a grant to download the code and run it inside your business. As the terms then say, broader rights come only from the author adopting a license.

This matters more than it sounds, because the "no license" case is common among exactly the kind of useful, actively maintained project a small business wants to build on. In my last article I catalogued twelve free public data sources, and the widely used Reddit archive project in that set has no license file on either of its repositories despite well over a thousand stars. Popularity is not permission. Neither is a maintainer being friendly about it.

## Question two: what the law actually says about collecting

This is where the popular understanding is most badly off, so it is worth being precise about what was decided, by which court, and when.

### The hacking statute has narrowed. The contract has not.

Everyone cites **hiQ Labs v. LinkedIn** for the proposition that scraping public data is legal. Two things about that.

First, the famous Ninth Circuit decision (31 F.4th 1180, April 2022) was an appeal from a **preliminary injunction**, applying a likelihood-of-success standard. It was not a final merits ruling. And it exists in that form only because the Supreme Court vacated the earlier 2019 version and sent it back to be reconsidered in light of Van Buren.

Second, and this is the part that almost never gets repeated: **hiQ then lost the contract claim.** In November 2022 the trial court found LinkedIn's user agreement language unambiguous, rejecting hiQ's argument that other clauses created ambiguity. The case ended in a stipulated consent judgment the following month. So the honest summary of the most-cited scraping case in America is that the collector won on the anti-hacking statute and lost on the contract.

**Van Buren v. United States** (593 U.S. 374, decided June 2021, 6 to 3, binding nationwide) is the reason the statute narrowed. It adopted a gates-up-or-down reading of the Computer Fraud and Abuse Act. But read footnote 8, because it is the whole ballgame: the Court expressly declined to decide whether the inquiry turns only on technical limits, "or instead also looks to limits contained in contracts or policies." The Supreme Court deliberately left the contract question open. It is still open.

One more correction while we are here, because it circulates as folklore: it is often said that a civil claim under that statute requires $5,000 in losses. The statute permits a civil action where the conduct involves any one of five listed factors. The $5,000 loss threshold is one of those five alternatives, not a universal element.

### The most collector-friendly ruling in years still leaves the contract standing

On 4 August 2026 the Ninth Circuit decided **Amazon v. Perplexity** (No. 26-1444), and it went about as well for the technology side as such a case can go. The panel vacated an injunction that had barred Perplexity's AI assistant from reaching Amazon on users' behalf, holding Amazon unlikely to succeed on its federal and California computer-access claims. The reasoning was that when a user directs the assistant, "it was the user who 'accessed' Amazon's computers," because Perplexity's own servers never talk to Amazon's directly. Screenshots go from the user's own machine to Perplexity, and instructions come back to that same machine.

Now read the limits the panel wrote into it. The holding covers those two statutes only, and **expressly leaves open other claims, including breach of terms of service**. The analysis is fact-specific, and the court said a more autonomous agent, or one whose servers communicated directly with the site, could come out differently. And because it came up from a preliminary injunction, it is a likelihood assessment on the present record, not a final judgment.

So even the best recent outcome for the collecting side says, in the opinion itself, that the contract theory survives.

### And the theories are multiplying, not shrinking

Two recent developments matter for anyone assuming the fight is over.

In **Reddit v. Anthropic**, a federal judge in California sent the case back to state court in March 2026, finding that Reddit's state-law claims for breach of contract, unjust enrichment, tortious interference, and unfair competition assert rights that are not equivalent to copyright. They contain, in the court's framing, extra elements. Copyright preemption does not sweep them away. That is a district court ruling on jurisdiction, not a merits decision, but the direction is clear.

In **Reddit v. SerpApi**, a federal judge in New York on 31 July 2026 largely denied motions to dismiss, allowing Reddit's DMCA circumvention claim to proceed against both defendants and a trafficking claim against one. This is the anti-circumvention part of copyright law, aimed at getting around access controls, and it is a genuinely newer theory in this area. It is a motion-to-dismiss ruling, meaning the allegations are assumed true and nothing has been found. But claims that survive a motion to dismiss are claims that cost real money to defend.

## What the platforms actually say right now

Terms change, and they change specifically in response to losses. I read these on 22 August 2026.

**Meta** rewrote its terms effective 1 January 2025 to close the exact gap a court had found in the Bright Data case, extending the prohibition so it applies regardless of whether the collection happens while logged out. If you read that 2024 decision and concluded that logged-out collection of public Meta pages is fine, note that the decision interpreted terms that no longer exist in that form. It was also a single unreviewed district court ruling. Meta dropped its remaining claim and gave up its appeal rights, so no appeals court ever looked at it, and it binds nobody.

**Google's** umbrella terms are narrower than people assume. They prohibit collecting content by automated means **in violation of the site's machine-readable instructions**, which points at robots.txt rather than a flat ban. But the Maps-specific additional terms are much more direct: they bar mass downloading or creating bulk feeds of the content, and bar using Maps to create or augment any other mapping-related dataset. The paid platform terms go further still, with a section headed "No Scraping" that expressly forbids copying and saving business names, addresses, or user reviews.

That last clause is worth sitting with, because building a local business list with reviews is the single most common thing a small business wants from this category of tool.

**Amazon's** conditions of use grant a license for personal and non-commercial use only, and specifically exclude any collection and use of product listings and descriptions.

**LinkedIn's** user agreement bans collection tools, and separately bans using LinkedIn-derived information obtained through third parties such as data aggregators. That second clause is the less known one, and it means buying the data instead of collecting it yourself does not put you outside the agreement.

**Reddit's** data API terms make the free tier non-commercial by contract. Any commercial use requires a separate agreement. The 2023 change was not only about price.

**YouTube's** terms are the flattest of the set: no automated means, with exactly two exceptions, being public search engines obeying robots.txt, or prior written permission.

## Question three: what you may do with what you have

Assume you collected it cleanly, or better, downloaded it from an official bulk file. You still have a licensing question, and this is where the traps are subtle.

**The code license is not the data license.** These are routinely different, and the data side is usually the restrictive one. ImageNet is the standard example: the tooling around it is commonly permissive, while its own terms of access begin by restricting use to non-commercial research and education. Check the dataset's own terms page, not the repository badge.

**An aggregator does not launder rights.** Common Crawl's terms grant a limited license to their service, and say plainly that the crawled material may be subject to separate terms from the owners of that content. Someone else having collected it does not transfer permission to you.

**"NonCommercial" is about the purpose, not about you.** The CC BY-NC definition is "not primarily intended for or directed towards commercial advantage or monetary compensation." Being a nonprofit does not make your use non-commercial. And here is the uncomfortable part: no US appellate court has ever defined what NonCommercial means in a Creative Commons license. The one well-known case, Great Minds v. FedEx, was decided on agency law and expressly did not reach the definition. If your plan depends on a fine reading of that term, your plan depends on an open question.

**Share-alike is narrower than the fear and broader than the shrug.** The Open Database License, which governs OpenStreetMap, is the one people most often get wrong in both directions. Four clauses matter:

Making a map, a report, or an app screen from the data creates what the license calls a Produced Work, and section 4.5(b) says that expressly does not create a derivative database. So share-alike does **not** force you to open-source your product. That is the fear, and it is unfounded.

But if you **modify** the database and then publicly use anything made from it, section 4.6 requires you to offer recipients a machine-readable copy of the modified database, or a file of your alterations. That is the shrug, and it is real.

Section 4.5(c) says internal use within an organization is not use "to the public," so none of the share-alike machinery attaches to purely internal analysis. Run whatever you like on your own data for your own decisions.

And section 4.3 requires attribution on any publicly used Produced Work. In practice that is the obligation businesses actually breach, because it is the easy one to forget.

**CDLA Permissive 2.0 is the clean one.** Its only condition is including the agreement text if you redistribute the data, and it explicitly imposes nothing on the results of your analysis. When you have a choice, this and outright public domain dedications are the terms to prefer.

## The regime nobody budgets for: personal data

If any part of what you collected describes an identifiable person, you have entered a separate body of law that does not care at all how public the source was.

European regulators have been explicit. Draft EDPB guidance adopted in July 2026 states that when people make personal data available online, that does not mean they consented to it being collected, and that the absence of a robots.txt file is not consent either. That guidance is in public consultation and is not final, so treat it as direction of travel rather than settled rule.

Enforcement is not hypothetical. France's regulator fined a business-contact tool €240,000 in December 2024 over a browser extension that pulled professional contact details. The finding was narrower than "fined for scraping," and it turned on lawful basis and on the rights of the people in the database.

The example that should land hardest for a US small business involves no European law at all. In February 2025 California's privacy regulator settled with a company called Background Alert, and the outcome was that the company shut down. Its product was built entirely on **public records**. What drew enforcement was not the collecting. It was that the company drew inferences from those records to build profiles of people. Public source, lawful availability, and the product still ended.

The lesson generalizes cleanly. Business listing data is fine right up until it contains an owner's name, personal mobile number, or home address, at which point a different rulebook applies to the same file.

## What actually happens to a small business

I do not want to leave you with the impression that a plumbing company pulling a competitor list is about to be sued by Google. That is not the realistic risk, and pretending otherwise is its own kind of dishonesty.

The enforcement ladder runs roughly: rate limiting, then blocking, then account restriction, then a cease and desist letter, then litigation. Small businesses overwhelmingly encounter the middle rungs. LinkedIn treats automated tool use as its own named category of account restriction, which is exactly the outcome that hurts a small operator most: not a lawsuit, but losing the account the business runs on.

The Amazon and Perplexity fight shows the entire ladder in one case, technical block through injunction and appeal, and that is what it looks like when the target is large enough to be worth the trouble.

So the honest risk assessment for a small business is not usually a courtroom. It is losing an account, having a process silently break when a site changes, and building a growth channel on something you cannot defend if it works well enough to be noticed. The third one is the expensive one, because the cost arrives exactly when the thing is succeeding.

## What is clearly fine

Plenty. This should not be paralysing, and the good options are genuinely better than the risky ones.

Official APIs and published bulk files are fine, and the government sources are the best of them. Federal award data, court records, mortgage records, and FDA data are dedicated to the public domain outright, which is a stronger position than most paid vendors will give you in writing.

Datasets under CDLA Permissive or a public domain dedication are fine for commercial products.

OpenStreetMap-derived work is fine for internal analysis with no conditions, and fine publicly with attribution, as long as you handle the modified-database disclosure if you altered it.

Reading a competitor's public page yourself, or having an assistant read a handful of pages the way a person would, is a different activity from running a systematic collection process against a platform at volume, both practically and in how terms are written.

And your own data is entirely yours. Most businesses have not exhausted their own transaction history, support tickets, and quote records, which carry no licensing question at all.

For the specific free sources worth starting with, that is what my previous piece covers: [the twelve public data sources I verified](/blog/blog-free-public-data-repos-small-business/), with what each one replaces and what it costs you in caveats.

## A five-minute check before you build

Run this on any dataset or tool before it becomes load-bearing.

Find the license file. If there is not one, you have no permission, and that is the end of the analysis unless you ask the author.

Find the **data** license separately from the code license, and read the dataset's own terms page.

Search that license text for "non-commercial." If it appears, and you are a business, stop and get advice rather than reasoning your way past it.

Ask whether the data was published for this, or merely reachable. An official bulk download is a publication. A page you can load is not.

Ask whether any field describes a person. If yes, you have a privacy question on top of the licensing question, and it does not go away because the source was public.

Write down what you concluded and when, with the URL. Terms change, as Meta's January 2025 rewrite shows, and the version you agreed to is the version in force on the day you collected.

If you want the broader argument for running your own tools instead of renting them, [The $20 Dollar Agency](https://the20dollaragency.com/) is the long version. This article is the part about doing it without building on sand.

## Fact-check notes and sources

All sources below were read on 22 August 2026. Court decisions are cited with their level and posture, because a trial court order and an appellate ruling are very different things and most reporting blurs them.

**No license means no permission**: GitHub's guidance at [choosealicense.com/no-permission](https://choosealicense.com/no-permission/) and the [GitHub Terms of Service](https://docs.github.com/en/site-policy/github-terms/github-terms-of-service), section D.5, whose grant is limited to use "through the Service."

**Van Buren v. United States**, 593 U.S. 374 (2021), decided 3 June 2021, 6 to 3, binding nationwide. Footnote 8 expressly reserves whether limits in contracts or policies count. Opinion at [supremecourt.gov](https://www.supremecourt.gov/opinions/20pdf/19-783_k53l.pdf).

**hiQ Labs v. LinkedIn**: the Ninth Circuit decision is 31 F.4th 1180 (9th Cir., 18 April 2022), an appeal from a preliminary injunction, issued on remand after the Supreme Court vacated the 2019 opinion. Opinion at [CourtListener](https://www.courtlistener.com/opinion/6460342/hiq-labs-inc-v-linkedin-corporation/). The later trial court ruling on the contract claim is N.D. Cal. No. 17-cv-03301-EMC, order of 4 November 2022 (Chen, J.), which found the user agreement's anti-scraping language unambiguous. The case ended in a stipulated consent judgment in December 2022 and was never appealed. Trial court order via [FindLaw](https://caselaw.findlaw.com/court/us-dis-crt-n-d-cal/2182242.html).

**Amazon.com Services v. Perplexity AI**, No. 26-1444 (9th Cir., 4 August 2026). Case page at [CourtListener](https://www.courtlistener.com/opinion/10939432/amazoncom-services-llc-v-perplexity-ai-inc/); opinion PDF at [cdn.ca9.uscourts.gov](https://cdn.ca9.uscourts.gov/datastore/opinions/2026/08/04/26-1444.pdf). That PDF carries no extractable text layer, so the characterisation of the holding here, including the express limitation to the computer-access statutes and the preservation of contract and tort theories, follows [Cooley's analysis of the decision](https://www.cooley.com/news/insight/2026/2026-08-06-ninth-circuit-rules-on-ai-agent-access-to-third-party-websites-under-cfaa) rather than my own reading of the opinion text. Posture: vacated preliminary injunction, likelihood-of-success standard, remanded.

**Computer Fraud and Abuse Act civil actions**: 18 U.S.C. section 1030(g) and 1030(c)(4)(A)(i)(I) through (V) at [Cornell LII](https://www.law.cornell.edu/uscode/text/18/1030). The $5,000 loss threshold is one of five alternative factors.

**Meta Platforms v. Bright Data**, N.D. Cal. No. 3:23-cv-00077-EMC, order of 23 January 2024 (Chen, J.), a single district court ruling interpreting the terms as then written. The parties filed a joint stipulation dismissing the remaining count on 23 February 2024 and Meta waived appeal. Meta's current terms, effective 1 January 2025, are at [facebook.com/legal/terms](https://www.facebook.com/legal/terms).

**Reddit v. Anthropic**, N.D. Cal. No. 3:25-cv-05643-TLT, remand order March 2026, holding the state-law claims not equivalent to copyright. **Reddit v. SerpApi**, S.D.N.Y. No. 1:25-cv-08736, opinion of 31 July 2026 (Engelmayer, J.), denying dismissal of the DMCA circumvention claim. Both are trial court rulings; neither decides the merits.

**Platform terms, all read 22 August 2026**: [Google Terms of Service](https://policies.google.com/terms) (effective 30 July 2026); [Google Maps additional terms](https://www.google.com/help/terms_maps/) (effective 27 January 2026), prohibiting mass download and dataset augmentation; [Google Maps Platform terms](https://cloud.google.com/maps-platform/terms/), section 3.2.3(a); [Amazon Conditions of Use](https://www.amazon.com/gp/help/customer/display.html?nodeId=GLSBYFE9MGKKQXXM); [LinkedIn User Agreement](https://www.linkedin.com/legal/user-agreement) section 8.2.2 (effective 3 November 2025); [Reddit Data API Terms](https://www.redditinc.com/policies/data-api-terms); [YouTube Terms of Service](https://www.youtube.com/t/terms).

**Open Database License 1.0**, full text at [opendatacommons.org](https://opendatacommons.org/licenses/odbl/1-0/). Section 4.3 attribution, 4.5(b) Produced Work exclusion, 4.5(c) internal use, 4.6 disclosure of a modified database. OpenStreetMap attribution guidance at the [OSM Foundation](https://osmfoundation.org/wiki/Licence/Attribution_Guidelines).

**CC BY-NC 4.0** definition of NonCommercial at [creativecommons.org](https://creativecommons.org/licenses/by-nc/4.0/legalcode.en). **Great Minds v. FedEx Office**, 886 F.3d 91 (2d Cir., 21 March 2018), decided on agency law without defining NonCommercial. Opinion at [CourtListener](https://www.courtlistener.com/opinion/4479341/great-minds-v-fedex-office-print-servs-inc/).

**CDLA Permissive 2.0** full text at [cdla.dev](https://cdla.dev/permissive-2-0/). **ImageNet** terms of access at [image-net.org](https://www.image-net.org/download.php). **Common Crawl** terms at [commoncrawl.org](https://commoncrawl.org/terms-of-use).

**EDPB Guidelines 03/2026** on web scraping in the context of generative AI, version 1.0, adopted 7 July 2026, paragraph 45. This is **draft guidance in public consultation**, not final. [PDF at edpb.europa.eu](https://www.edpb.europa.eu/system/files/2026-07/edpb_guidelines_2020603_webscraping_v1_en_0.pdf).

**CNIL decision SAN-2024-020** of 5 December 2024, €240,000, adopted as lead authority in cooperation with counterpart European regulators. **California Privacy Protection Agency** settlement with Background Alert, Inc., board-approved 26 February 2025, announcement at [cppa.ca.gov](https://cppa.ca.gov/announcements/2025/20250227.html).

**The two tools described in the opening** are [gosom/google-maps-scraper](https://github.com/gosom/google-maps-scraper) (MIT, whose readme carries the single-sentence legal notice quoted) and [facundoolano/google-play-scraper](https://github.com/facundoolano/google-play-scraper) (MIT, no terms warning found in its readme). Star and fork counts are as of 22 August 2026. Neither is accused of wrongdoing here, and both are competently built. They are cited because their license badges are a good illustration of a question a badge cannot answer.

## Related reading

**[The Market Research You Were Quoted $4,000 For Is Sitting in a Public Repo](/blog/blog-free-public-data-repos-small-business/)**: the twelve free sources I verified, and what each one replaces.

**[DMCA Takedowns and Terms of Use for Developers](/blog/blog-dmca-takedown-and-terms-for-devs/)**: the practical mechanics when a notice actually arrives.

**[Verifying an Auditor's Findings Without Trusting the Tool](/blog/blog-ethical-scraping-self-auditing/)**: checking a vendor's report against the underlying data without overstepping.

**[Every Unsplash Photo On Your Site Legally Needs Attribution](/blog/blog-tool-image-licensing-credit-audit/)**: the same attribution problem in the place it bites most often.

**[Delete Yourself From 603 Data Brokers](/blog/delete-yourself-from-data-brokers/)**: what the personal-data side of this looks like when you are the record rather than the collector.

*This post is informational and is not legal advice. I am not a lawyer, court decisions are summarised here with their posture noted precisely because that posture matters, and none of this is a substitute for advice about your own situation. Terms of service and dataset licences change without notice, and several matters described above are actively being litigated. Mentions of third-party products, companies, and datasets are nominative fair use. No affiliation is implied.*


---

Canonical HTML: https://jwatte.com/blog/blog-available-is-not-permission-dataset-licensing/
RSS: https://jwatte.com/feed.xml
JSON Feed: https://jwatte.com/feed.json
Hero image: https://jwatte.com/images/blog-available-is-not-permission-dataset-licensing.webp
