arxiv:2412.15450
Bram Vanroy PRO
BramVanroy
AI & ML interests
Artificial intelligence, natural language processing, computational linguistics
Recent Activity
liked a dataset about 3 hours ago
lightonai/embeddings-pre-training posted an update about 4 hours ago
**I benchmarked HF buckets against https access for Common Crawl.**
Took me a while to get round to do this but I benchmarked access to Common Crawl via https vs hf buckets. Both experiments were run at night in Europe. I do not think other hardware problems were impacting the speeds since CPU processing time of the non-download pipeline components were highly similar (within 2% identical) and below only the WarcReader speeds of datatrove are used.
Experiment: selected 5 disjoint samples of 64 files each (randomly from the latest crawl; 20,499 docs/file). Those five batches were then processed by 32 single-core tasks with 4GB/core (five batches to calculate CIs). Paired experiment between using https and hf bucket.
- https: 40.0 [39.3-40.6] (seconds per WARC file)
- hf bucket: 172.0 [122.7-221.2]
That is a difference of about 4x in streaming speed. You'll see that https is also more stable (smaller CI).
I also ran raw throughput tests to the endpoints to measure rate limiting (64MiB transfer at 8/32/128/256 concurrent readers) and rate limiting seems not an issue for either: at any of those parallel reader numbers, their respective speeds stay about the same.
Note that, given CC scale, this is still a small test. Rate limiting may become more obvious when processing a full crawl. I do not know whether the https endpoint vs HF bucket will shut you out earlier with which limits. new activity 10 days ago
codefuse-ai/F2LLM-v2-160M:fix hidden size