aikyam school

Stability Selection via Bootstrapped Regularized Logistic Regression

RCTReview

High-dimensional text corpora contain thousands of sparse words, making standard regression models prone to overfitting when isolating key linguistic features.

Picture this

Imagine selecting the best players for a sports team by running hundreds of mini-scrimmages with randomly picked sub-groups of athletes; players who consistently perform well across almost every random match-up are the true star performers.

What the evidence says

High business content is reliably flagged by terms like product, project, business, market, idea, need, customer, service, plan, and cost, whereas low business content is flagged by hello, hi, token, leaderboard, status, and thanks.

Who was studied
N = 140,000+ messages written by 4,958 entrepreneurs across 49 African countries.
How
Bootstrapped fitting of 200 to 1,000 $l_1$-penalized LASSO logistic regressions on random sub-samples drawn with replacement, selecting words with non-zero coefficients in >50% of iterations.

What to do

Apply stability selection with $l_1$ LASSO penalization across repeated random bootstrap sub-samples to identify stable text regressors in high-dimensional communication data.

From the source

"We fit from 200 to 1,000 independent logistic-regression models on random subsets of the data... The regressors that have a score larger than 0.5 are declared to be informative."

Peer Networks and Entrepreneurship- a Pan-African RCT.pdf

Tags

  • stability selection
  • lasso regularization
  • text mining