Design and Analysis of a Novel Respondent-Driven Sampling Methodology for Estimation of Labor Violation Prevalence in Low-Wage Industries
Previous work utilizing RDS to sample low-wage workers has suffered from issues of seed bias, making inference difficult. To address this problem, we propose a new design that collects seeds in a probability sample, and study this design's resilience to network homophily, or the tendency for similar people to cluster within social networks. The structure of this design is novel in its focus on estimation within multiple sub-populations of interest (for example, low-wage industries), and in its formulation of complex constraints imposed on recruitment to limit bias. We study and model the population networks and recruitment sampling, propose a modified estimator, and, via simulation, analyze the validity of inference.
Results indicate that inference in this design is feasible, and that modifications to a popular RDS estimator to account for the sampling constraints improve the accuracy of estimation. While the accuracy of the estimator is promising, further improvements to this estimator and the network generation algorithm are likely necessary to properly assess the validity of inference. These improvements include incorporating the sub-population structure of the sampling more fully into the estimator and implementing non-uniform homophily effects estimation and correction within the estimator.

