A thing I have come to believe after a year of working with scraped datasets: the hard part was never collection. It is provenance. Anyone can gather a hundred million rows. The question that actually determines whether the dataset is usable is whether you can say, for any given row, where it came from and when. Without that you cannot deduplicate honestly, you cannot handle takedown requests, and you cannot tell a downstream consumer what they are training on. Most data projects treat provenance as metadata to bolt on later. It is not. It is the schema.

BitFan
Public Service Atlas for Bittensor