# Random Forests and Gradient Boosting — Machine Learning Basics

Source: https://www.skillbyai.com/en/machine-learning/e-ens

> Many trees beat one.

## Averaging and correcting

A **random forest** trains many decision trees on random samples of rows and features and averages their votes, which cancels much of the noise individual deep trees learn. **Gradient boosting** trains small trees one after another, each correcting the errors of the ensemble so far. For tabular data, these tree ensembles (including libraries such as XGBoost, LightGBM and CatBoost) are often the strongest practical choice, need little preprocessing, and handle non-linear relationships and feature interactions.

## Stronger models, chosen carefully

Ensembles of trees are strong defaults for tables; tune them fairly and inspect what they rely on.

![Three ideas: ensembles, search, importance.](assets/figures/machine-learning/section-6-map.svg) — Figure 6.1 — Ensembles, search and importance.

## One tree versus two ensembles, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. With 5-fold cross-validation on the breast cancer data, one decision tree scores 0.917, a 200-tree random forest 0.960 and gradient boosting 0.963.

```python
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.model_selection import cross_val_score
from sklearn.tree import DecisionTreeClassifier
X, y = load_breast_cancer(return_X_y=True)
for name, m in [("single decision tree", DecisionTreeClassifier(random_state=0)),
                ("random forest (200 trees)", RandomForestClassifier(n_estimators=200, random_state=0)),
                ("gradient boosting", GradientBoostingClassifier(random_state=0))]:
    print(f"{name:<26} 5-fold accuracy {cross_val_score(m, X, y, cv=5).mean():.3f}")
```

Output:

```
single decision tree       5-fold accuracy 0.917
random forest (200 trees)  5-fold accuracy 0.960
gradient boosting          5-fold accuracy 0.963
```

## Start with a forest

A random forest with default settings is a robust first serious model for tabular data and a good benchmark for anything fancier.

**Quiz:** How does a random forest reduce overfitting compared with one deep tree?

- [ ] It removes the test set
- [ ] It uses only one feature
- [x] It averages many trees trained on random subsets
- [ ] It trains a single shallow rule

*Answer:* It averages many trees trained on random subsets. Averaging cancels individual trees' noise.
