LightGBM won tabular ML by attacking the hot loop from both sides: fewer rows with GOSS, fewer effective columns with EFB, all plugged into histogram-based split search and greedy leaf-wise growth.
LightGBM speeds up the repeated histogram-and-split loop rather than changing the basic boosting recipe. The picture below treats training cost as a moving product of row pressure, feature width, and candidate split scans.
Keep all high-gradient examples, subsample low-gradient ones, then reweight their contribution. Toggle the sampler to see how the selected population and estimated split gains change.
EFB treats sparse one-hot-ish columns like a conflict graph. If two features rarely fire together, the learner can pack them into a shared bundle and build fewer histograms.
The paper’s splashy speedups come from the full package, not a single isolated trick. Switch datasets to see how time, memory, and metric retention vary with sparsity and task shape.
Primary papers and comparison points from the episode. Links use the provided paper PDF, Google Scholar entries, and the related prior episode.