0 or 1
spam / not spam
cat / dog
1 of N
digit 0-9
fruit type (5 classes)
many labels
movie tags
image has: cat + outdoors
Today: binary + multi-class. Regression in Part 3.
model = keras.Sequential([
layers.Dense(64, activation='relu',
input_shape=(784,)),
layers.Dense(32, activation='relu'),
# 1 neuron + sigmoid = probability
layers.Dense(1, activation='sigmoid')
])
model.compile(
optimizer='adam',
loss='binary_crossentropy',
metrics=['accuracy']
)
- Output: 1 neuron, sigmoid
- Output range: (0, 1) = probability of class 1
- Loss: binary_crossentropy
\( L = -[y \log(\hat{y}) + (1-y)\log(1-\hat{y})] \)
- Decision: threshold at 0.5
- Metrics: accuracy, precision, recall
import numpy as np
from tensorflow import keras
from keras import layers
(X_train, y_train), (X_test, y_test) = \
keras.datasets.mnist.load_data()
# Flatten & normalize
X_train = X_train.reshape(-1, 784) / 255.0
X_test = X_test.reshape(-1, 784) / 255.0
# Binary: is this digit a 5?
y_train = (y_train == 5).astype(float)
y_test = (y_test == 5).astype(float)
model = keras.Sequential([
layers.Dense(128, activation='relu',
input_shape=(784,)),
layers.Dense(1, activation='sigmoid')
])
model.compile(optimizer='adam',
loss='binary_crossentropy',
metrics=['accuracy'])
model.fit(X_train, y_train, epochs=10,
batch_size=128, validation_split=0.2)
- MNIST: 60k train + 10k test, 28x28 grayscale images
- Normalize: pixel values 0-255 -> 0.0-1.0 (divide by 255)
- Flatten: 28x28 = 784 features
- Target: 1 if digit is 5, else 0
Epoch 10/10 - 375/375 - loss: 0.03 - accuracy: 0.99
[Tip] Binary MNIST is easy -- the real challenge is multi-class (Part 2)
TP / (TP + FP)
Of all predicted positive: how many are actually positive?
TP / (TP + FN)
Of all actual positive: how many did we catch?
- Accuracy can be misleading when classes are imbalanced
- 99% accuracy on spam = only 1% is spam? Model might say "not spam" every time
- Use precision/recall for imbalanced problems
model.compile(
optimizer='adam',
loss='binary_crossentropy',
metrics=['accuracy',
keras.metrics.Precision(),
keras.metrics.Recall()]
)
model = keras.Sequential([
layers.Dense(128, activation='relu',
input_shape=(784,)),
layers.Dense(64, activation='relu'),
# 10 neurons + softmax = probabilities
layers.Dense(10, activation='softmax')
])
model.compile(
optimizer='adam',
loss='categorical_crossentropy',
metrics=['accuracy']
)
- Output: N neurons (one per class)
- Softmax: turns logits into probabilities that sum to 1
\( \text{softmax}(z_i) = \frac{e^{z_i}}{\sum_j e^{z_j}} \)
- Loss: categorical_crossentropy
- Decision: argmax(output) = predicted class
[0, 0, 0, 0, 0, 1, 0, 0, 0, 0]
Position 5 = 1, all others = 0
[0, 0, 0, 1, 0, 0, 0, 0, 0, 0]
Position 3 = 1, all others = 0
- keras.utils.to_categorical(y, 10)
- Converts integer labels to one-hot vectors
y_train = keras.utils.to_categorical(y_train, 10) # Before: [5, 0, 3, 7, ...] # After: # [[0,0,0,0,0,1,0,0,0,0], # [1,0,0,0,0,0,0,0,0,0], # [0,0,0,1,0,0,0,0,0,0], # [0,0,0,0,0,0,0,1,0,0]]
sparse_categorical_crossentropy = skip one-hot, use integer labels directly
(X_train, y_train), (X_test, y_test) = \
keras.datasets.mnist.load_data()
X_train = X_train.reshape(-1,784)/255.0
X_test = X_test.reshape(-1,784)/255.0
y_train = keras.utils.to_categorical(y_train,10)
y_test = keras.utils.to_categorical(y_test,10)
model = keras.Sequential([
layers.Dense(128, activation='relu',
input_shape=(784,)),
layers.Dense(64, activation='relu'),
layers.Dense(10, activation='softmax')
])
model.compile(optimizer='adam',
loss='categorical_crossentropy',
metrics=['accuracy'])
model.fit(X_train, y_train, epochs=10,
batch_size=128, validation_split=0.2)
loss, acc = model.evaluate(X_test, y_test)
print(f"Test accuracy: {acc:.4f}")
# Test accuracy: ~0.9750
- Same data as binary, but all 10 classes
- Output: 10 probabilities (sum to 1.0)
- argmax = predicted digit
preds = model.predict(X_test[:5]) print(preds.argmax(axis=1)) # [7, 2, 1, 0, 4] print(y_test[:5].argmax(axis=1)) # [7, 2, 1, 0, 4] -- match!
[Tip] ~97.5% accuracy with just 2 Dense layers. Adding convolution would reach 99%+
[2.0, 1.0, 0.1, 3.5, 0.3]
[0.17, 0.06, 0.03, 0.70, 0.04]
Sum = 1.0
argmax = 3 (highest)
- Softmax = "winner take all": highest logit gets most probability
- Temperature controls sharpness:
\( \text{softmax}(z_i / T) \)
- T=1: normal. T=0.1: very sharp. T=10: very smooth
- In Keras: activation='softmax' = T=1
[Tip] Output of softmax + cross-entropy = stable numerically. Never compute log(softmax) manually.
Key difference: output layer has no activation = output any real number
# Single output (most common)
model = keras.Sequential([
layers.Dense(64, activation='relu',
input_shape=(8,)),
layers.Dense(32, activation='relu'),
layers.Dense(1) # NO activation!
])
model.compile(
optimizer='adam',
loss='mse', # Mean Squared Error
metrics=['mae'] # Mean Absolute Error
)
import numpy as np
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
data = fetch_california_housing()
X_train, X_test, y_train, y_test = \
train_test_split(data.data, data.target,
test_size=0.2)
# IMPORTANT: scale features
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
model = keras.Sequential([
layers.Dense(64, activation='relu',
input_shape=(8,)),
layers.Dense(32, activation='relu'),
layers.Dense(1)
])
model.compile(optimizer='adam',
loss='mse', metrics=['mae'])
model.fit(X_train, y_train, epochs=50,
batch_size=32, validation_split=0.2)
- 8 features: income, age, rooms, etc.
- Target: house value (in $100k units)
- Must scale features for neural networks (StandardScaler)
- Output: 1 number = predicted price
loss, mae = model.evaluate(X_test, y_test)
print(f"MAE: {mae:.3f}")
# MAE: ~0.32 (average error: $32k)
[Tip] MAE is easier to interpret: "off by $32k on average"
Output: Dense(1, sigmoid)
Loss: binary_crossentropy
email spam, disease, fraud
Output: Dense(N, softmax)
Loss: categorical_crossentropy
digit recognition, fruit type
Output: Dense(1)
Loss: mse or mae
house price, temperature
Only change: output layer and loss function
Hidden layers stay the same (Dense + ReLU)
Wrong loss = model learns the wrong thing. Always double-check!
model.summary()uv init --no-package wk04-hw for your projectNext week: Evaluation, Overfitting, and Underfitting