updates128 days ago
`data-universe` release v1.18.1
Scraper sampling fix: read_random_row_group used .head(max_rows) which always picked rows from the top of each row group. Miners could place 5 real tweets at the top and fill the remaining rows wit...
Original engithub.com
034
It patched a sampling issue that could let bad rows hide inside larger files, a meaningful data-quality and trust fix. Link:
