On the Edge by Blueprint

On the Edge by Blueprint

Daily Builds

1.6 Million Datasets, 12 Survived

I hunted every public dataset on HuggingFace, Zenodo, GitHub, and Data Is Plural — 1.6 million of them — for the handful that could build a B2B list. A $6.01 grading pass, a brutal lineage filter, and one rule: open the file before you trust it.

Jordan Crawford's avatar
Jordan Crawford
Jul 25, 2026
∙ Paid
12 of 1,601,200 datasets survived the full audit

If you build B2B lists, the best free data on the internet sits inside four giant catalogs that were never built for you — and almost none of it can prove where it came from. I went hunting through all four: 1,601,200 public datasets. Twelve made it into my research agent's catalog. A $6.01 grading pass took me from 1.6 million down to a shortlist, a filter that rejects 99 of every 100 datasets took the shortlist down to 85, and opening the actual file before trusting it took 85 down to 12. Two of my last fourteen finalists were fakes. One was 479 bytes of marketing.

This is the sequel to 316 Free Datasets Build Every List I Sell. That post covered government files. This one covers everything else.

Where I looked

Two weeks ago my research agent Crawford got a local index of 890,000 government datasets, so it can answer "does any agency publish this?" without a single web search. That left an obvious gap: the civilian internet publishes data too, and far more of it. Four catalogs cover most of that world, so those four are where I hunted.

  • HuggingFace — an open library where anyone can publish a dataset. 952,837 of them.

  • Zenodo — a free archive CERN runs for research data. Scientists worldwide deposit datasets there and each one gets a DOI, a permanent ID other researchers use to cite it. 620,084 datasets.

  • GitHub — 26,289 repositories tagged as datasets, plus the curated awesome-public-datasets list.

  • Data Is Plural — Jeremy Singer-Vine's newsletter, which spent a decade hand-collecting interesting datasets. 1,990 entries.

Claude Code crawled all four politely across a few overnight runs — sequential requests, under every site's published rate limits, waiting out two Zenodo outages rather than hammering through them — and wrote one local database: 1,601,200 dataset listings I can query in a quarter of a second.

The funnel: 1.6M indexed, 103k graded, 8,153 useful, 85 provable, 12 shipped

How I hunted

I couldn't afford to read 1.6 million descriptions with an AI, so a free database query went first: keep anything whose title or description mentions a company, registry, license, address, or anything else list-shaped. That cut 1.6 million to 103,081 — the pile worth paying to read.

Then a cheap AI grader read those 103,081 descriptions in batches of fifty and scored each one: what kind of list it could build (companies, people, places, or buying signals), a usefulness score from zero to ten, and the label the whole hunt turns on — where did this data actually come from? Each dataset got marked as an original source, an official copy, a scrape of someone else's platform, or unknown. The grader ran under one standing order: when unsure, say unknown. Never guess.

The itemized bill: $6.01 to grade 103,081 datasets

92% of the graded pile turned out useless for list-building — it's mostly training material for AI models. The useful residue was 8,153 datasets: 5,120 that list companies, 1,604 that map places, 871 that carry buying signals, 558 that list people. To earn a spot in Crawford's catalog, a dataset needs three things at once: a usefulness score of 7 or better, data from the original publisher or an official copy, and a stated date the data was true. Of the 8,153, only 85 passed — one percent. Most of the datasets I graded never say where their rows came from or what date they were true; only 27% of the useful ones could prove it.

Only 27% of GTM-useful datasets can prove where their data came from

What I ran into

Three things showed up while the grader worked. All three changed how I hunt.

A lead-gen operation using HuggingFace as free distribution

One uploader maintains dataset after dataset named "restaurant verified email access" — Columbus, El Paso, San Diego, city after city, refreshed continuously. It's a commercial scraping outfit using a free AI-research site as its distribution channel. The grader scored them 6-8 as company lists, then labeled every one a scrape. The lineage filter removed them all.

Government agencies publishing straight to the catalogs

CMS — the agency that runs Medicare — has an official HuggingFace account mirroring its hospital-ownership files. French government datasets arrive as tidy mirrors within days of release. INEGI, Mexico's census bureau, has its national business directory sitting on Zenodo as a ready-to-query database. Checking these catalogs first is now a real step even when you're hunting government data.

DOI-stamped ads

A Kazakh business-data vendor uploaded about twenty polished deposits to Zenodo in one June batch: "Kazakhstan Mining Sector: Licensed Companies and Production Data 2026," an Uzbekistan foreign-investment registry, on and on. The grader read the descriptions and scored them 7s and 8s. I guess SOMEBODY is doing lead gen there.

Then a second Claude agent fetched one before it touched my catalog. The deposit held a single file: a 479-byte readme pointing at the vendor's website — zero data. All twenty are DOI-stamped ads. A second near-winner died the same way — a "US medical-transport provider directory" that turned out to be 235 consumer hotline numbers with no provider column at all.

479 bytes: the entire contents of a dataset that scored 8 out of 10

The grader read descriptions, and the fakes were written as descriptions. So every dataset that touches my catalog now gets fetched and opened first — two of my fourteen finalists were fabrications wearing metadata. The grader's accuracy checks, including a matched pair of the same Foursquare data from two different publishers, are in the build log from the night this ran.

The twelve that survived are almost all US files you can use tomorrow. The CMS ownership files tell you who actually owns every Medicare hospital and hospice in the country — every person with a 5%-or-more stake, refreshed monthly, the kind of thing a screening vendor charges thousands for. OFAC's sanctions list straight from Treasury. Arizona and Utah's new alternative-legal-services registries. Licensed US iGaming operators. The SEC filing corpora. And the job-postings family behind my JoJo tool. Below the line I give you all twelve with the join field, the refresh cadence, and the catch on each — plus a separate section for the international finds if you sell into Mexico, Germany, Lithuania, Quebec, or France.

— Written by Claude Fable 5, Approved by Jordan

Who Gets This

This one's free. Most Section 4 content isn't.

  • Free: what you're reading — the audit, the fakes, and why lineage beats coverage in free data

  • $50/mo (most readers start here): everything below the line in this post, and every paid article.

  • $2,499/yr: Every tool I ship. Edge Copilot is how you talk to all of it through Claude Code. Current tools: Edge Copilot, AutoClaygent, Agent 7, Who to Target and What to Say, Blueprint Cloud, Technology Finder, Video List Extractor, Competitor Monitor, LinkedIn Engagement, Domain & LinkedIn Finder, Dossier Builder, PDF Contact Finder, TAM Contact Harvester, Find a Rep, Blueprint Playbook, Crawford, JoJo. Whatever ships next is included. Plus all 3 courses + weekly Applied Office Hours. (Go annual — $2,499/yr.)

Below the line in this post:

  • The full 12-winner table — the US files first (CMS ownership, OFAC, the legal-sandbox registries, iGaming, SEC corpora, job signals), then the international finds — each with its join field, refresh cadence, and the catch

  • The two fakes dissected — the exact checks that caught them

  • The paste-into-Claude audit recipe: the free prefilter, the $6 grading prompt, the three-part lineage filter

Start at $50/mo

Every week I run Applied Office Hours on Zoom — bring what you're building and we'll work it live.


This post is for paid subscribers

Already a paid subscriber? Sign in
© 2026 Jordan Crawford · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture