Learning Plan
/

Neural Networks:
Classification and Regression

binary - multi-class - multi-label - regression

Learning Outcomes

1. Build classifiers for binary and multi-class tasks
2. Understand output layer: softmax, sigmoid
3. Build regressors that predict continuous values
4. Use common Keras patterns: MNIST, California Housing
Week 4 -- 1145208 Deep Learning

Today's Roadmap

Part 1 - Classification
Binary classifiers
output layer, loss functions, MNIST binary
Part 2 - Multi-class
Multi-class classification
softmax, MNIST 10 classes, one-hot encoding
Part 3 - Regression
Regression networks
continuous output, California Housing, no activation
Part 4 - Wrap up
Summary + Homework
comparison table, review, exercises

Part 1: Classification Tasks

Binary

0 or 1

spam / not spam

cat / dog

Multi-class

1 of N

digit 0-9

fruit type (5 classes)

Multi-label

many labels

movie tags

image has: cat + outdoors

Today: binary + multi-class. Regression in Part 3.

Binary classifier: output layer

model = keras.Sequential([
    layers.Dense(64, activation='relu',
                 input_shape=(784,)),
    layers.Dense(32, activation='relu'),
    # 1 neuron + sigmoid = probability
    layers.Dense(1, activation='sigmoid')
])

model.compile(
    optimizer='adam',
    loss='binary_crossentropy',
    metrics=['accuracy']
)

- Output: 1 neuron, sigmoid

- Output range: (0, 1) = probability of class 1

- Loss: binary_crossentropy

\( L = -[y \log(\hat{y}) + (1-y)\log(1-\hat{y})] \)

- Decision: threshold at 0.5

- Metrics: accuracy, precision, recall

Example: MNIST binary (digit 5 vs not-5)

import numpy as np
from tensorflow import keras
from keras import layers

(X_train, y_train), (X_test, y_test) = \
    keras.datasets.mnist.load_data()

# Flatten & normalize
X_train = X_train.reshape(-1, 784) / 255.0
X_test  = X_test.reshape(-1, 784) / 255.0

# Binary: is this digit a 5?
y_train = (y_train == 5).astype(float)
y_test  = (y_test == 5).astype(float)

model = keras.Sequential([
    layers.Dense(128, activation='relu',
                 input_shape=(784,)),
    layers.Dense(1, activation='sigmoid')
])
model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])
model.fit(X_train, y_train, epochs=10,
          batch_size=128, validation_split=0.2)

- MNIST: 60k train + 10k test, 28x28 grayscale images

- Normalize: pixel values 0-255 -> 0.0-1.0 (divide by 255)

- Flatten: 28x28 = 784 features

- Target: 1 if digit is 5, else 0

Epoch 10/10 - 375/375 - loss: 0.03 - accuracy: 0.99

[Tip] Binary MNIST is easy -- the real challenge is multi-class (Part 2)

Beyond Accuracy: Precision and Recall

Precision

TP / (TP + FP)

Of all predicted positive: how many are actually positive?

Recall

TP / (TP + FN)

Of all actual positive: how many did we catch?

F1 Score

2 * (P * R) / (P + R)

Harmonic mean: balances precision and recall

- Accuracy can be misleading when classes are imbalanced

- 99% accuracy on spam = only 1% is spam? Model might say "not spam" every time

- Use precision/recall for imbalanced problems

model.compile(
    optimizer='adam',
    loss='binary_crossentropy',
    metrics=['accuracy',
             keras.metrics.Precision(),
             keras.metrics.Recall()]
)

Part 2: Multi-class Classification

model = keras.Sequential([
    layers.Dense(128, activation='relu',
                 input_shape=(784,)),
    layers.Dense(64, activation='relu'),
    # 10 neurons + softmax = probabilities
    layers.Dense(10, activation='softmax')
])

model.compile(
    optimizer='adam',
    loss='categorical_crossentropy',
    metrics=['accuracy']
)

- Output: N neurons (one per class)

- Softmax: turns logits into probabilities that sum to 1

\( \text{softmax}(z_i) = \frac{e^{z_i}}{\sum_j e^{z_j}} \)

- Loss: categorical_crossentropy

- Decision: argmax(output) = predicted class

One-hot Encoding

Label: 5

[0, 0, 0, 0, 0, 1, 0, 0, 0, 0]

Position 5 = 1, all others = 0

Label: 3

[0, 0, 0, 1, 0, 0, 0, 0, 0, 0]

Position 3 = 1, all others = 0

- keras.utils.to_categorical(y, 10)

- Converts integer labels to one-hot vectors

y_train = keras.utils.to_categorical(y_train, 10)
# Before: [5, 0, 3, 7, ...]
# After:
# [[0,0,0,0,0,1,0,0,0,0],
#  [1,0,0,0,0,0,0,0,0,0],
#  [0,0,0,1,0,0,0,0,0,0],
#  [0,0,0,0,0,0,0,1,0,0]]

sparse_categorical_crossentropy = skip one-hot, use integer labels directly

Full Example: MNIST 10 classes

(X_train, y_train), (X_test, y_test) = \
    keras.datasets.mnist.load_data()

X_train = X_train.reshape(-1,784)/255.0
X_test  = X_test.reshape(-1,784)/255.0
y_train = keras.utils.to_categorical(y_train,10)
y_test  = keras.utils.to_categorical(y_test,10)

model = keras.Sequential([
    layers.Dense(128, activation='relu',
                 input_shape=(784,)),
    layers.Dense(64, activation='relu'),
    layers.Dense(10, activation='softmax')
])
model.compile(optimizer='adam',
    loss='categorical_crossentropy',
    metrics=['accuracy'])
model.fit(X_train, y_train, epochs=10,
          batch_size=128, validation_split=0.2)
loss, acc = model.evaluate(X_test, y_test)
print(f"Test accuracy: {acc:.4f}")
# Test accuracy: ~0.9750

- Same data as binary, but all 10 classes

- Output: 10 probabilities (sum to 1.0)

- argmax = predicted digit

preds = model.predict(X_test[:5])
print(preds.argmax(axis=1))
# [7, 2, 1, 0, 4]
print(y_test[:5].argmax(axis=1))
# [7, 2, 1, 0, 4] -- match!

[Tip] ~97.5% accuracy with just 2 Dense layers. Adding convolution would reach 99%+

Softmax in detail

Raw logits (before softmax)

[2.0, 1.0, 0.1, 3.5, 0.3]

After softmax

[0.17, 0.06, 0.03, 0.70, 0.04]

Sum = 1.0

Predicted class

argmax = 3 (highest)

- Softmax = "winner take all": highest logit gets most probability

- Temperature controls sharpness:

\( \text{softmax}(z_i / T) \)

- T=1: normal. T=0.1: very sharp. T=10: very smooth

- In Keras: activation='softmax' = T=1

[Tip] Output of softmax + cross-entropy = stable numerically. Never compute log(softmax) manually.

Part 3: Regression -- predict continuous values

Classification

  • - Output: class label (discrete)
  • - Output layer: softmax / sigmoid
  • - Loss: cross-entropy
  • - Example: is this email spam?

Regression

  • - Output: number (continuous)
  • - Output layer: no activation (linear)
  • - Loss: MSE or MAE
  • - Example: predict house price

Key difference: output layer has no activation = output any real number

Regression output layer

# Single output (most common)
model = keras.Sequential([
    layers.Dense(64, activation='relu',
                 input_shape=(8,)),
    layers.Dense(32, activation='relu'),
    layers.Dense(1)  # NO activation!
])

model.compile(
    optimizer='adam',
    loss='mse',        # Mean Squared Error
    metrics=['mae']    # Mean Absolute Error
)

- Dense(1) = no activation = linear output

- Can output any real number: -50.3, 0, 1234567

- MSE: \( L = \frac{1}{N}\sum(y - \hat{y})^2 \)

- MAE: \( L = \frac{1}{N}\sum|y - \hat{y}| \)

- Never use sigmoid/softmax for regression -- it would cap output to (0, 1)

Example: California Housing price prediction

import numpy as np
from sklearn.datasets import fetch_california_housing
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

data = fetch_california_housing()
X_train, X_test, y_train, y_test = \
    train_test_split(data.data, data.target,
                     test_size=0.2)

# IMPORTANT: scale features
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test  = scaler.transform(X_test)

model = keras.Sequential([
    layers.Dense(64, activation='relu',
                 input_shape=(8,)),
    layers.Dense(32, activation='relu'),
    layers.Dense(1)
])
model.compile(optimizer='adam',
              loss='mse', metrics=['mae'])
model.fit(X_train, y_train, epochs=50,
          batch_size=32, validation_split=0.2)

- 8 features: income, age, rooms, etc.

- Target: house value (in $100k units)

- Must scale features for neural networks (StandardScaler)

- Output: 1 number = predicted price

loss, mae = model.evaluate(X_test, y_test)
print(f"MAE: {mae:.3f}")
# MAE: ~0.32 (average error: $32k)

[Tip] MAE is easier to interpret: "off by $32k on average"

Pattern Summary

Binary Classification

Output: Dense(1, sigmoid)

Loss: binary_crossentropy

email spam, disease, fraud

Multi-class

Output: Dense(N, softmax)

Loss: categorical_crossentropy

digit recognition, fruit type

Regression

Output: Dense(1)

Loss: mse or mae

house price, temperature

Only change: output layer and loss function

Hidden layers stay the same (Dense + ReLU)

Loss Function Cheat Sheet

Classification

binary_crossentropy -- binary, output sigmoid
categorical_crossentropy -- multi-class, output softmax
sparse_categorical_crossentropy -- multi-class, labels as integers

Regression

mse -- Mean Squared Error, penalizes large errors more
mae -- Mean Absolute Error, more robust to outliers
huber -- combines MSE + MAE, robust + smooth

Wrong loss = model learns the wrong thing. Always double-check!

Common Pitfalls

Do

  • - Normalize inputs: scale to 0-1 or standardize
  • - Start simple: 1-2 hidden layers, 32-128 units
  • - Monitor val_loss: stop if it goes up
  • - Check shape: model.summary()

Don't

  • - Don't use softmax for regression
  • - Don't use sigmoid for multi-class
  • - Don't train without validation split
  • - Don't forget to scale test data (use train scaler)

Summary + Homework

What we covered

Homework (due next week)

  • - Build a classifier for Fashion MNIST (10 classes of clothing)
  • - Architecture: 784 -> 256(relu) -> 128(relu) -> 10(softmax)
  • - Try: different hidden layer sizes (32 vs 256)
  • - Plot training curves, compare train vs val accuracy
  • - Use uv init --no-package wk04-hw for your project

Next week: Evaluation, Overfitting, and Underfitting