HOW MANY QUERIES DO YOU NEED? BUDGET-GOVERNED LABEL SPARSITY IN SERP-DERIVED WEB SUPERVISION
Keywords:
search engine result pages, weak supervision, label sparsity, sampling budget, web measurement, class imbalance, reproducibilityAbstract
Work that predicts which web domains surface in search results tends to build its labels the same way: issue a set of queries, record what comes back, treat everything unseen as negative. How many queries gets reported as a detail, and is often settled by whatever a vendor’s free tier permits. It is not a detail. Holding a 1,000-domain frame and its features fixed while sweeping the budget from 5 to 500 queries, we find labelled positives growing as K0.47 against K0.87 for distinct domains seen. A popularity frame exhausts; the open web does not. Labels therefore cost about the square root of spend. Budget also decides what a study concludes, not merely how tightly. A fused model over Chrome UX Report, Common Crawl and live-crawl features loses to a zero-parameter popularity ranker at small budgets, draws level near 25 queries, and wins by 0.149 PR-AUC at 500. Same code, same features, opposite claim. Leave-one-source-out ablation puts the entire effect in Common Crawl together with CrUX; the live crawl, by a wide margin the costliest source we collected, adds nothing we can measure. Two labelling decisions usually compressed into a single clause, domain resolution and rank cutoff, shift the measured effect by nearly a factor of two. Re-issuing those same queries four months on returns result sets overlapping at mean Jaccard 0.49, so supervision depreciates too. We close with a formula converting a target prevalence into a query budget, and with two budgets below which claims about feature value outrun the data that produced them. Code, data and environment are released.


