60%
Model learns from this data
weights updated here
20%
Check during training
tune hyperparameters
20%
Final evaluation only
never touch during training
Test set = final exam. Peeking at it = cheating.
# Option 1: validation_split (simple)
model.fit(X_train, y_train,
epochs=50,
validation_split=0.2)
# Last 20% of X_train used as validation
# Option 2: pass validation data (better)
model.fit(X_train, y_train,
epochs=50,
validation_data=(X_val, y_val))
# Option 3: K-Fold (most robust)
from sklearn.model_selection import KFold
kfold = KFold(n_splits=5)
for train_idx, val_idx in kfold.split(X_train):
model.fit(X_train[train_idx],
y_train[train_idx],
validation_data=(
X_train[val_idx],
y_train[val_idx]))
- validation_split=0.2 = last 20% is validation
- Warning: data must not be shuffled before split (Keras uses last 20%)
- Better: split manually, then use validation_data
- K-Fold: most robust but slower (trains K models)
[Tip] For time series: never shuffle! Use past for train, future for val.
loss ----
val_loss ----
Both decrease, stay close
loss ------
val_loss / (goes up!)
Train good, validation bad
loss ---- (stays high)
val_loss ---- (stays high)
Both stay high, no learning
Gap between train and val = overfitting. Both high = underfitting.
import matplotlib.pyplot as plt
history = model.fit(X_train, y_train,
epochs=50,
validation_split=0.2)
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
# Loss
axes[0].plot(history.history['loss'],
label='Train')
axes[0].plot(history.history['val_loss'],
label='Val')
axes[0].set_title('Loss')
axes[0].legend()
# Accuracy
axes[1].plot(history.history['accuracy'],
label='Train')
axes[1].plot(history.history['val_accuracy'],
label='Val')
axes[1].set_title('Accuracy')
axes[1].legend()
plt.tight_layout()
plt.show()
- Plot both loss and accuracy side by side
- What to look for:
- Are both curves going down?
- Is the gap between train/val growing?
- Is val_loss starting to increase?
- Early stopping point: when val_loss stops decreasing
[Tip] Save best model only: keras.callbacks.ModelCheckpoint('best.keras', save_best_only=True)
Analogy: student who memorizes answers but can't solve new problems
Both train and val performance are poor. The model hasn't learned the pattern.
1 Dense(8)
164 parameters
Underfitting
2 Dense(64)
24,960 parameters
Good fit
5 Dense(256)
500k+ parameters
Overfitting
Start small, increase until val performance stops improving
from keras import regularizers
model = keras.Sequential([
layers.Dense(64, activation='relu',
input_shape=(784,),
kernel_regularizer=
regularizers.l2(0.01)),
layers.Dense(64, activation='relu',
kernel_regularizer=
regularizers.l2(0.01)),
layers.Dense(10, activation='softmax')
])
- Regularization adds penalty to loss function
- L2: \( +\lambda \sum w^2 \) = push weights toward zero
- L1: \( +\lambda \sum |w| \) = can make weights exactly zero (sparse)
- 0.01 = lambda (regularization strength)
- Larger lambda = stronger regularization = simpler model
model = keras.Sequential([
layers.Dense(128, activation='relu',
input_shape=(784,)),
layers.Dropout(0.5), # 50% off
layers.Dense(64, activation='relu'),
layers.Dropout(0.3), # 30% off
layers.Dense(10, activation='softmax')
])
- During training: randomly set X% of neurons to zero
- Forces network to not rely on any single neuron
- Dropout(0.5): 50% of neurons randomly off each batch
- Only active during training, not during inference
- Like training an ensemble of many smaller networks
[Tip] Typical: 0.2-0.5 for Dense layers. 0.1-0.3 for input.
early_stop = keras.callbacks.EarlyStopping(
monitor='val_loss', # what to watch
patience=5, # wait 5 epochs
restore_best_weights=True
)
model.fit(X_train, y_train,
epochs=100, # train "forever"
validation_split=0.2,
callbacks=[early_stop])
- monitor='val_loss' = watch validation loss
- patience=5 = stop if no improvement for 5 epochs
- restore_best_weights=True = revert to best epoch
- Train many epochs, but only keep the best version
[Tip] Always use EarlyStopping + ModelCheckpoint together for best results
checkpoint = keras.callbacks.ModelCheckpoint(
'best_model.keras',
monitor='val_loss',
save_best_only=True, # overwrite only better
save_weights_only=False
)
model.fit(X_train, y_train,
epochs=100,
validation_split=0.2,
callbacks=[checkpoint, early_stop])
# Load best model later
model = keras.models.load_model(
'best_model.keras')
- Saves model to file after each epoch
- save_best_only=True = only save if val improved
- Without this: final epoch saved (may be overfit!)
- .keras format = full model (architecture + weights + optimizer)
# Image augmentation with Keras
data_augmentation = keras.Sequential([
layers.RandomFlip("horizontal"),
layers.RandomRotation(0.1),
layers.RandomZoom(0.2),
layers.RandomContrast(0.2),
])
# Use in model
model = keras.Sequential([
data_augmentation, # <-- first layer
layers.Flatten(input_shape=(28,28)),
layers.Dense(128, activation='relu'),
layers.Dense(10, activation='softmax')
])
# Or use on-the-fly during training
datagen = keras.preprocessing.image\
.ImageDataGenerator(
rotation_range=10,
width_shift_range=0.1,
horizontal_flip=True)
- Creates new training samples from existing ones
- For images: flip, rotate, zoom, shift, contrast
- 1000 original images -> thousands of variations
- Only apply to training, not validation/test
- Greatly reduces overfitting when data is limited
[Tip] Augmentation = free data. Use it whenever you have less than ~10k images per class
from keras import regularizers
model = keras.Sequential([
layers.Dense(128, activation='relu',
input_shape=(784,),
kernel_regularizer=regularizers.l2(1e-4)),
layers.Dropout(0.3),
layers.Dense(64, activation='relu',
kernel_regularizer=regularizers.l2(1e-4)),
layers.Dropout(0.3),
layers.Dense(10, activation='softmax')
])
model.compile(optimizer='adam',
loss='categorical_crossentropy',
metrics=['accuracy'])
callbacks = [
keras.callbacks.EarlyStopping(
monitor='val_loss', patience=5,
restore_best_weights=True),
keras.callbacks.ModelCheckpoint(
'best.keras', save_best_only=True)
]
model.fit(X_train, y_train, epochs=200,
validation_split=0.2,
callbacks=callbacks)
- L2 regularization: small penalty (1e-4)
- Dropout: 30% after each hidden layer
- EarlyStopping: patience=5, restore best
- ModelCheckpoint: save best to file
- Train 200 epochs, but only keep the best version
[Tip] Start with this recipe, then tune if needed
uv init --no-package wk05-hwNext week: Feature Engineering and preprocessing