Check our work before you talk to us
We have no reference customers yet. Rather than ask you to believe a capability claim, we published six datasets and the method behind them.
Every one is built from public-domain government records or from a Bitcoin archival node we operate ourselves. Every summary figure is a live spreadsheet formula pointing back at the rows it came from, so you can click any number and follow it. Nothing in them is a stored value pretending to be a calculation.
Free, CC BY 4.0, no signup, no email wall.
What's in them
OFAC digital currency designations
Every cryptocurrency address on the US Treasury SDN list, across nineteen chains, with the designated entity and sanctions programme behind each.
The reconciliation tab is the part worth your time. We parse Treasury's
SDN.XML directly and also ingest an independent open-source mirror as a
cross-check. They disagree, and the three addresses that differ are named
individually so you can adjudicate it yourself.
Fifteen sanctions regimes, reconciled
36,558 active designations across OFAC SDN and five other OFAC programmes, the EU consolidated list, UK OFSI, the UN Security Council list, four BIS lists and two State Department lists — on one spine, with a regime-by-regime divergence matrix.
90.7% (29,506 of 32,532 distinct names) appear on exactly one list. If you screen against a single jurisdiction, that figure is your blind spot.
Export-control designations and the OFAC gap
The BIS and State Department lists that stop shipments, which sanctions tooling aimed at finance mostly ignores. 92.7% (5,441 of 5,867 distinct names) appear on no OFAC list at all, each one named individually. Fifty years of listing history.
Structural CoinJoin detection from a full node
69,364 Bitcoin transactions classified as collaborative spends from transaction shape alone — no label list, no attribution vendor, no block explorer. The full classification rule is printed in the workbook: thresholds, branch order, confidence values. Reimplement it against your own node and you should get these rows back.
This workbook documents a defect in our own detector. Roughly three quarters of its output sits at denominations where batching and dust are far more likely than mixing. We publish the full corpus, the distribution that exposes the problem, and the subset we would actually stand behind.
Exposure is not guilt
A precedence-ranked model of how addresses relate to designated ones, with the model weight each tier contributes stated openly — including the address-poisoning victims we identify and deliberately score at zero.
Anyone can send an unsolicited sub-dollar payment from a designated address to any address they like. A model that scores the recipient for it can be weaponised by the sender.
Dark web infrastructure, measured
We hold a multi-million-document crawl corpus spanning Tor, I2P and Freenet. None of it is in this workbook, and that is the point.
What is in it is everything that can be said about the networks without publishing anything collected from them: a census of 14,398 Tor relays, months of I2P network-database observations from a floodfill router we operate, reachability measurements against 202,456 distinct hidden services, and a software census of what the hidden web actually runs on.
Onion addresses are hashed. Router hashes and site keys are dropped. Cross-site identifier reuse appears only as counts.
If you are weighing whether to trust us with sensitive collection, that file is the argument.
Why we publish the mistakes too
Two of the six workbooks lead with something we got wrong.
A control that has never caught anything is a control nobody has tested. Our pre-publication scan reads the rendered spreadsheets and refuses to release a file containing identifiers it should not — and the first thing it caught was us, about to publish addresses belonging to parties no government has designated and no court has named.
We would rather tell you that than have you find it.
Formats
Each dataset ships in a dark practitioner variant and a light print-safe
compliance variant, as both .ods and .xlsx, with a PDF of each compliance
version. .xlsx is included because chart fidelity through Excel's and Google
Sheets' ODS import is worse than LibreOffice's, and the charts are half the
point.
A SHA256SUMS manifest lets you verify you got what we published.