Domain Drop-Catching

22 years of domain sales, a scoring model, and a portfolio simulation. Nothing bought, no money spent.  ·  ← Kalshi & Polymarket  ·  Crypto research

The bottom line

The idea was: expiring domain names that sound like real English words — but aren't — are the ones that resell. Score those, buy them cheap when they expire, sell them on.

To test "sounds like English" I trained a letter-pattern model on an English wordlist and scored every domain by how English-like its letters are. A made-up word like bluecart scores high; qzwxrt scores low. That was the whole thesis.

Nothing here is deployed. No domain has been registered and no money spent.

The headline numbers

What actually predicts a sale

Each number is how much that feature moves the odds, averaged across every test year, with how much it wobbled between years. A feature whose sign flips between years is not a finding, and is labelled as such rather than averaged into looking respectable.

Does the score actually sort domains?

Year by year

Trained only on years before each test year — never a random split, which would let the future leak into the past. The median sale price for each year sits beside it, because a model can look clever purely by being tested in a boom.

If you actually ran this as a portfolio

Where the money ends up across every run. The average is the misleading number here — a handful of lucky runs drag it up while most runs quietly lose.

Dropping domains is the risk control

At each renewal you can drop a domain instead of paying again. This is what happens when you vary how many of the best-scoring domains you keep. Note the last row.

What the data could not tell us

Absolutely everything below scales with one number nobody can measure: how often a domain actually sells. These are the same simulation at different assumed rates.

Where the data came from — and didn't

For a backtest built on scraped data, the sources that were unavailable matter as much as the ones that worked. Published here rather than buried in a footnote.

Ways this could still be fooling itself

Twelve ways a backtest like this leaks information from the future or from how the data was built. Eleven are fixed and tested. The first one is not fixable with public data, and it bounds everything above.

Every number that was invented rather than measured