The sample is to the population as the bootstrap sample is to the sample. By resampling with replacement from our observed data, we approximate the sampling distribution of any statistic.
The standard deviation of the bootstrap distribution estimates the standard error of the statistic.
Draw a sample, then run the bootstrap. Compare the bootstrap SE to the true SE. Try different populations and statistics. The bootstrap works well even when the population is skewed or when the statistic has no simple formula for its SE.
The bootstrap gives us the sampling distribution of a statistic without relying on formulas. We can use that distribution to build confidence intervals.
Take the middle 95% (or other level) of the bootstrap distribution directly.
Use the bootstrap SE but assume the sampling distribution is normal. Works well when the bootstrap distribution is approximately bell-shaped.
A 95% confidence interval should contain the true parameter 95% of the time. The coverage experiment repeats the entire process many times and checks how often the CI actually captures the truth.
Coverage can be below the nominal level when the sample is small or when the population is heavily skewed. Try the right-skewed population with a small n to see this.
Compute a single CI to see it on the bootstrap distribution. Then run the coverage experiment to see how reliable that CI method is across many samples.
Split the data into K equally-sized folds. Train on K-1 folds, test on the held-out fold. Rotate through all K folds and average the test errors.
Small K (e.g., 3): each fold is large, so more bias (less training data per split) but lower variance. Large K (e.g., LOOCV): nearly unbiased but higher variance since folds overlap heavily.
K = 5 or K = 10 is the standard compromise.
Training error always decreases with model complexity. CV error reveals the true U-shape. The gap between the two is the key to understanding overfitting.
Generate data, then use the fold viewer slider to step through each fold's train/test split. Toggle the CV error curve to see training vs CV error across polynomial degrees. The minimum of the CV curve picks the right complexity.
Regularized regression adds a penalty to prevent overfitting. The penalty strength is controlled by a tuning parameter. Cross-validation finds the value that minimizes out-of-sample error.
Ridge (α=0): shrinks all coefficients toward zero but never removes them.
Lasso (α=1): shrinks some coefficients exactly to zero (variable selection).
Elastic Net (0<α<1): a blend of both.
Slide λ and watch coefficients shrink. Small λ = weak penalty (complex model). Large λ = strong penalty (simple model). The coefficient paths show the full trajectory.
λ1se is the largest λ whose CV error is within 1 standard error of the minimum. It produces a simpler model that performs nearly as well.
This app tunes λ for a fixed α. In practice, you can tune both jointly. Run CV over a grid of (α, λ) pairs and pick the combination with the lowest CV error. Packages like caret and tidymodels automate this.