Find your next best customer by knowing why your last best customer chose you
Every deal review I'd run used today's internet as proof about a deal that closed years ago. A grader caught it in a live engagement. Here's the fix — and the dated snapshot that pins 'SOC 2' to public view 799 days before a deal lost on trust.
Every win-loss review I'd ever run quietly used today's internet as proof about a deal that closed years ago. A customer's site says "SOC 2 compliant" today — a security certification enterprise buyers ask for — so the retro assumes it said that before the deal, and treats it as a signal you could have acted on. Until last month, my own backcasting — working backward from a closed deal to the signals that should have predicted it — couldn't tell the difference between a signal that was visible then and one that only showed up since. This is the discipline that fixed it: prove when a signal was visible, with a dated snapshot anyone can click, or admit you can't.
A good win-loss study exists to find the signal that came before the outcome. You take the deals you won and the ones you lost, and you go looking for the thing that separated them — the tell that was sitting in public before anyone signed anything. Find that tell, and you can go find the next batch of good-fit buyers before your competitor does.
Here's where it goes wrong, and it went wrong in a real engagement of mine. You pull a lost customer's website and you see a careers page, a compliance badge, a pricing tier, a new office. You write down: "They were scaling — we should have seen it." But you're looking at the site as it exists today. The version that was live when you lost the deal, two years ago, might have had none of that. You didn't check. I didn't check. The web you're reading is the present, and you're quietly filing it as the past.
An independent grader — brought in to pressure-test the work, with no stake in making it look good — caught exactly this. The finding was blunt: the backcasting was counting things it had no proof were visible at the time. Every "we could have known" claim was resting on a page that may not have existed yet. The whole study was built on evidence we'd never dated.
What a snapshot actually proves
The fix started with a question I typed into Claude Code — the AI coding tool I run all my research through: could its research agents go back in time and check what was actually visible while the deal was still open? The answer was sitting in a free library I'd walked past for years: the Internet Archive's Wayback Machine, which has been photographing the public web since the 1990s. If a page existed and the archive crawled it, there's a dated copy you can pull up.
So every claim about the past now carries one of three labels, and only two of them count.
Snapshot-proven. The archive has a dated copy from before the deal, and the signal is right there in it. This is the gold standard — a URL with a date on it that a skeptical buyer can click.
Record-dated. The signal comes from a record that carries its own date — a permit, a filing, a press release — and that date, adjusted to when the record actually surfaced in public (more on that adjustment below), lands before the deal.
Assumed. You only ever saw today's web. Maybe the signal was visible back then, maybe it wasn't. This label is allowed to exist. It's never allowed to count.
Only the first two are firm. An assumed signal sits in the file marked "we think," and it never makes it into the final list of tells you'd bet money on.
The one rule I'd tattoo on anyone doing this work: no snapshot does not mean the page wasn't there. A page the archive skipped — told to keep out, or simply never crawled — looks identical to a page that never existed: both come back empty. So "no snapshot" always means can't prove, never wasn't there. The discipline treats those as two different outcomes, and collapsing them is how you launder a guess into a fact.
A useful date is a tight window
A finding dated "sometime between 2022 and 2024" is close to worthless. You can't act on a two-year fog. The deliverable is the tight window.
The archive gives you that window for free, because its index can hand you just the distinct versions of a page — one row per real change, the identical re-crawls collapsed away. Line up the versions and you get a before and an after. The signal is missing in one snapshot and present in the next, so it appeared in the gap between them. That gap is the answer.
I ran the tool — call it the prover, since pinning visibility dates is its whole job — live on a page anyone can check. I asked it when the word "Claude" first showed up on the homepage of Anthropic, the company behind the Claude AI models, checked against a reference date — the "was it visible by this point?" cutoff — of June 1, 2023. It read about a dozen archived copies and pinned the window to a two-day gap: the name was absent on March 11, 2023, and present on the snapshot from March 13, 2023. Claude launched publicly on March 14. The site carried the name the day before the announcement, and the archive proves it to the day.
Take the compliance badge from the top of this post. I pointed the same tool at the security page of HubSpot, the marketing-software company, and asked when "SOC 2" first appeared there. The window came back tight: absent in early October 2015, present in the snapshot from October 25, 2015. Picture a deal you lost on trust and security concerns in early 2018. That compliance signal had been sitting in public for 799 days before you lost. The archive turns "they were probably compliant" into a number: how many days the tell sat in plain sight before the deal closed. That number — the days a signal sat in public before the deal closed — is what this discipline buys you.
A record goes public later than the date it carries
The date printed on a record comes before the date the world could actually see it, and the difference is easy to miss even once you've started dating your signals.
A construction permit gets filed in March. The public portal that lists it doesn't show it until June. If you date your signal to March, you've credited yourself with knowing something three months before anyone could have looked it up. So the "knowable-by" date is June, when the record showed up where the world could see it. Every record type gets that haircut, its date pushed later to the moment it became visible, and the safe move is to assume the slow version unless an archived copy of the portal itself proves the signal showed up earlier.
Map the ground before you guess
Dated proof needs somewhere to point — a source that existed, that the archive covered, that you can actually check. Which is why the biggest change was the order of operations.
The old way: brainstorm what signals might predict a good-fit buyer, then go hunting for data to support them. The new way flips it. Before writing a single guess, one research agent goes out per customer type and maps what that industry's public data actually looks like — the license boards, the state rosters, the permit portals, the regulator feeds. It confirms each source is real by pulling one live record, and it dates the archive's coverage of that source so you know how far back you can even see.
When the license record itself is locked behind a bot wall, the scout doesn't quit — it finds the state board's monthly roster instead and checks how far back the archive's copies of that go. There is no universal catalog of "here's every dataset for dentists." You earn the map one industry at a time, and it compounds: the map you build for one engagement is waiting for the next one in the same vertical.
Fewer findings, on purpose
This discipline produces fewer findings on purpose.
Some genuinely true signals die labeled "can't prove," because small companies are barely archived and the snapshot just isn't there. Losing those is the deliberate trade. And a dated snapshot only gets a finding into the room: each surviving signal still has to show, on a set of deals it was never tuned on, that it actually separates the wins from the losses. What survives is a shorter list where every item carries a dated URL a skeptical revenue leader can click and check for themselves. And the two ways a finding can die get reported separately, always. "Killed because the evidence was weak" sits in one column. "Killed because we couldn't prove when it was visible" sits in another. Keeping them apart is what stops thin archive coverage from quietly masquerading as "this signal wasn't there."
The LinkedIn hole
One source breaks the whole model, and you should know where the floor drops out. Login-walled pages are barely archived. I pointed the prover at a mid-size software company's LinkedIn page and got back a flat "no captures" — the archive has no usable copy of it at all, which matches an earlier spot-check that found zero readable snapshots of one such page over a seven-year span.
So the rule for those sources is deliberately lopsided. A snapshot that does exist and shows the signal proves the signal was there. A missing or login-walled snapshot proves nothing at all, because a locked page and a page that never existed look exactly the same to the archive. You can prove presence on LinkedIn. You can never prove absence.
The archive will block you
Pull pages too fast and the archive cuts your address off. While I was building this, the server dropped my connections cold, for stretches that lasted tens of minutes, because I was pulling pages too fast from one address.
That behavior is why the tool ships with a shared brake. Every copy of it running on the same machine spaces its requests three seconds apart, and the moment one of them gets refused, all of them back off together — sixty seconds, then two minutes, climbing to a fifteen-minute cap — so a swarm of research agents retreats as one organism instead of each one re-tripping the block. Anything already fetched gets saved to disk, so when the block clears, the run picks up where it stopped. If the archive refuses you during your own runs, that's the same wall. Wait it out.
What I got wrong in the first version
The first version of this shipped with a quiet bug I'm not proud of. A mismatch in how one field was named meant every backcast check silently came back "unknown" — the tool was doing the work and then dropping the answer on the floor, every time, without complaint. It looked like it was running. It wasn't reading. The version that shipped last week fixes it, along with dating each deal to its own correct reference point instead of stamping every deal with its close date.
When the Wayback Machine has a gap, its quieter sibling Common Crawl — photographing the web since 2008 — is the second place to look, and a dated record found there proves visibility the same way a snapshot does. And if you'd rather run all of this than rebuild it, annual subscribers already have it — the win-loss tool picked up this whole time-travel layer in an update on July 18.
The lesson under all of it is the same one behind never trusting what an agent tells you just because it sounds certain: a claim about the past is only as good as the dated proof under it. This is also one more case of the pattern that the best go-to-market data is free and sitting in plain sight — the archive of the entire web has been open the whole time.
— Written by Claude Fable 5, Approved by Jordan
Below is the geeky version. Copy it into Claude Code and rebuild the whole thing yourself.
Below the line in this post:
The index trick that lists every distinct version of a page for free — the whole thing rests on it
A binary search that pins a signal's first-seen date in about a dozen page reads
The shared brake that keeps a swarm of agents from getting your address blocked
The three visibility tags and two death columns that wire the discipline into any win-loss study
Or skip the rebuild: annual subscribers install the tool I actually built with one command — every tool I ship, all 3 courses, weekly Applied Office Hours. (Go annual — $2,499/yr.)
Every week I run Applied Office Hours on Zoom — bring what you're building and we'll work it live.







