New model reveals hidden structure in web crawls: persistent core vs dynamic shell
Common Crawl and GAW data show most URLs are ephemeral, only 20% persist.
Researchers introduce a two-component urn model to analyze longitudinal web crawls from Common Crawl (2020–2025) and the German Academic Web (GAW). They identify a persistent core fraction κ and a dynamic shell, reconciling discrepancies in traditional pairwise containment metrics. A residual on the shell's coverage parameter remains, signaling that the shell itself is not homogeneous.
- Two-component urn model separates a persistent core (fraction κ ≈ 20%) from a dynamic shell in longitudinal web crawls.
- Applied to Common Crawl (2020–2025, domain granularity) and German Academic Web (URL granularity), revealing non-uniform URL populations.
- Discovery curve U(s,T) and pairwise containment are reconciled as two projections of a single process, with disagreements signaling population heterogeneity.
Why It Matters
Better web crawl models mean more accurate archive coverage estimates, improving research reproducibility and search engine index quality.