Learning Plan
/

Introduction to Keras and TensorFlow

setup - layers - compile - train - evaluate

Learning Outcomes

1. Set up a TensorFlow/Keras project with uv
2. Build a neural network using the Sequential API
3. Explain layers, activation, optimizer, loss
4. Train, evaluate, and make predictions with Keras
Week 3 -- 1145208 Deep Learning

Today's Roadmap

Part 1 - Setup
Install and configure
uv init, add TensorFlow, verify import
Part 2 - Build
Create a model
Sequential, Dense, activation functions
Part 3 - Train
Compile and fit
loss, optimizer, metrics, model.fit()
Part 4 - Wrap up
Evaluate and predict
model.evaluate(), model.predict(), summary + homework

Part 1: Why TensorFlow + Keras?

TensorFlow

  • - Open-source library by Google
  • - Deploy anywhere: server, mobile, edge
  • - Production-ready ecosystem
  • - Keras is the official high-level API

Keras

  • - High-level API: simple, readable, fast
  • - 3 lines of code = working model
  • - Great for learning and prototyping
  • - Built into TensorFlow as tf.keras

Keras = friendly interface / TensorFlow = powerful engine underneath

Project setup with uv

# 1. Create new project
$ uv init --no-package wk03
$ cd wk03

# 2. Add TensorFlow
$ uv add tensorflow

# 3. Run your code
$ uv run python main.py

- uv init --no-package wk03 = create project folder with config

- uv add tensorflow = install TF + dependencies into venv

- uv run = run in the project's venv (no "wrong env" bugs)

- No need to manually create venv -- uv handles it

[Tip] Project structure: wk03/ contains main.py, pyproject.toml, uv.lock

Verify: import TensorFlow + Keras

# main.py
import tensorflow as tf
from tensorflow import keras

print("TensorFlow:", tf.__version__)
print("Keras:", keras.__version__)

# Quick test: tensor operations
x = tf.constant([[1.0, 2.0],
                 [3.0, 4.0]])
print(x @ x.T)   # matrix multiply

- import tensorflow as tf = standard import

- from tensorflow import keras = Keras API

- TF tensor works like NumPy array + GPU support

- If you see version numbers = ready to go

$ uv run main.py
TensorFlow: 2.x.x
Keras: 3.x.x
tf.Tensor(
[[ 7. 10.]
 [15. 22.]], shape=(2,2))

Part 2: Sequential API -- stack layers

from tensorflow import keras
from keras import layers

model = keras.Sequential([
    # Hidden layer 1
    layers.Dense(64, activation='relu',
                 input_shape=(10,)),
    # Hidden layer 2
    layers.Dense(32, activation='relu'),
    # Output layer
    layers.Dense(1, activation='sigmoid')
])

model.summary()

- Sequential() = stack layers one after another

- Dense(units) = fully connected layer (every neuron connects to every neuron in next layer)

- activation = non-linearity function

- input_shape=(10,) = tell the model the shape of one input sample

[Tip] First layer must have input_shape. Later layers infer automatically.

Layers and Neurons

Input(10)  Dense(64)   Dense(32)   Dense(1)
 [x1] --> [h1] -----> [h3] ----> [y]
 [x2] --> [h2] -----> [h4]
 ...       ...         ...
[x10] --> [h64]

input_shape=(10,)

Each sample has 10 features

e.g. 10 measurements of a patient

Dense(64, relu)

64 neurons, each learns a pattern

Params: 10*64 + 64 = 704

Dense(1, sigmoid)

1 neuron, output probability

Binary classification (0 or 1)

More neurons = more capacity to learn, but also more risk of overfitting

Activation Functions

ReLU

f(x) = max(0, x)

Most popular for hidden layers

Fast, simple, effective

Sigmoid

\( \sigma(z) = \frac{1}{1+e^{-z}} \)

Output: (0, 1) = probability

Used in output layer for binary classification

Softmax

\( \sigma(z_i) = \frac{e^{z_i}}{\sum e^{z_j}} \)

Output sums to 1 (probability distribution)

Used for multi-class classification

Hidden layers: use ReLU / Output: match the task

Counting Parameters

model = keras.Sequential([
    layers.Dense(64, activation='relu',
                 input_shape=(10,)),
    layers.Dense(32, activation='relu'),
    layers.Dense(1, activation='sigmoid')
])

model.summary()
# Layer (type)         Output     Param
# dense (Dense)        (None,64)  704   <-- 10*64+64
# dense_1 (Dense)      (None,32)  2080  <-- 64*32+32
# dense_2 (Dense)      (None,1)   33    <-- 32*1+1
# Total params: 2,817

- Formula: Dense(in, out) = in*out + out (weights + bias)

- Layer 1: 10*64 + 64 = 704

- Layer 2: 64*32 + 32 = 2080

- Layer 3: 32*1 + 1 = 33

- Total: 2,817 trainable parameters

[Tip] model.summary() = instant parameter count -- use it to sanity-check model size

Part 3: Compile -- set up learning

model.compile(
    optimizer='adam',
    loss='binary_crossentropy',
    metrics=['accuracy']
)

- optimizer = algorithm that adjusts weights

- Adam = adaptive learning rate, good default

- SGD = classic, slower but predictable

- loss = how wrong the model is

- binary_crossentropy = binary classification

- categorical_crossentropy = multi-class

- mse = regression

- metrics = what to monitor (not used for training)

model.fit() -- train the model

import numpy as np

# Sample data: 1000 samples, 10 features
X_train = np.random.randn(1000, 10)
y_train = (X_train[:, 0] > 0).astype(float)

# Train!
history = model.fit(
    X_train, y_train,
    epochs=20,
    batch_size=32,
    validation_split=0.2
)

- epochs = how many times to see all data

- batch_size = samples per gradient update

- validation_split = hold out 20% for validation

- history = dict with loss/metrics per epoch

Epoch 1/20 - 25/25 - 0s - loss: 0.72 - accuracy: 0.51
Epoch 2/20 - 25/25 - 0s - loss: 0.48 - accuracy: 0.82
...
Epoch 20/20 - 25/25 - 0s - loss: 0.21 - accuracy: 0.93

[Tip] Watch val_loss -- if it goes up while loss goes down = overfitting

Reading the history

# history contains training curves
import matplotlib.pyplot as plt

plt.plot(history.history['loss'],
         label='train loss')
plt.plot(history.history['val_loss'],
         label='val loss')
plt.xlabel('Epoch')
plt.ylabel('Loss')
plt.legend()
plt.show()

# Also:
# history.history['accuracy']
# history.history['val_accuracy']

- loss decreasing = model is learning

- val_loss decreasing = generalizes to unseen data

- loss down, val_loss up = overfitting

- both flat = model is too simple (underfitting)

[Tip] Early stopping: stop when val_loss stops improving. Use keras.callbacks.EarlyStopping()

Part 4: model.evaluate() -- test performance

# Test data (unseen during training)
X_test = np.random.randn(200, 10)
y_test = (X_test[:, 0] > 0).astype(float)

# Evaluate
loss, accuracy = model.evaluate(X_test, y_test)
print(f"Test loss: {loss:.4f}")
print(f"Test accuracy: {accuracy:.4f}")

# Output:
# 7/7 [======] - 0s - loss: 0.19 - accuracy: 0.94
# Test loss: 0.1923
# Test accuracy: 0.9400

- evaluate() = run model on test set, return loss + metrics

- Uses the same loss/metrics from compile()

- Returns tuple: (loss, metric1, metric2, ...)

- Test accuracy should be close to validation accuracy

[Tip] Never peek at test data during training -- it's your final exam, not practice

model.predict() -- use the model

# New data
X_new = np.array([[0.5, -0.3, 0.8, 0.1,
                    -0.2, 0.6, -0.1, 0.4,
                    0.9, -0.5]])

# Predict
pred = model.predict(X_new)
print(pred)       # [[0.87]]
print(pred[0][0]) # 0.87 = probability

# Convert to class
if pred[0][0] > 0.5:
    print("Class 1")
else:
    print("Class 0")

- predict() = forward pass on new data

- Returns probability (if sigmoid output)

- Threshold: pred > 0.5 = class 1

- For multi-class: np.argmax(pred, axis=1)

[Tip] In production, save model: model.save('my_model.keras')

Complete Example -- end to end

import numpy as np
from tensorflow import keras
from keras import layers

# 1. Data
X = np.random.randn(1000, 10)
y = (X[:, 0] * X[:, 1] > 0).astype(float)

# 2. Model
model = keras.Sequential([
    layers.Dense(64, activation='relu', input_shape=(10,)),
    layers.Dense(32, activation='relu'),
    layers.Dense(1, activation='sigmoid')
])

# 3. Compile
model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])

# 4. Train
model.fit(X[:800], y[:800], epochs=20, validation_split=0.2)

# 5. Evaluate
loss, acc = model.evaluate(X[800:], y[800:])
print(f"Test accuracy: {acc:.4f}")

# 6. Predict
preds = model.predict(X[800:5])
print("Predictions:", (preds > 0.5).astype(int).ravel())

Keras vs PyTorch (preview)

Keras (this week)

model = Sequential([
  Dense(64, relu, input_shape=(10,)),
  Dense(1, sigmoid)])
model.compile(optimizer='adam',
              loss='binary_crossentropy')
model.fit(X, y, epochs=10)

+ Simple, fast prototyping

+ Good defaults, less boilerplate

PyTorch (week 9+)

model = nn.Sequential(
  nn.Linear(10, 64), nn.ReLU(),
  nn.Linear(64, 1), nn.Sigmoid())
opt = optim.Adam(model.parameters())
for epoch in range(10):
    # ... manual training loop

+ Full control, easier to debug

+ Preferred in research

Learn both: Keras first (easy start) -> PyTorch later (deeper understanding)

Summary + Homework

What we covered

  • - Project setup: uv init --no-package -> uv add tensorflow
  • - Sequential API: stack Dense layers
  • - Activation: ReLU (hidden), Sigmoid/Softmax (output)
  • - compile(optimizer, loss, metrics)
  • - fit() -> evaluate() -> predict()

Homework (due next week)

  • - Build a classifier for MNIST digits (use keras.datasets.mnist)
  • - Architecture: 784 -> 128(relu) -> 10(softmax)
  • - Try different: optimizers (sgd vs adam), batch_size (16 vs 128)
  • - Plot training curves, report test accuracy
  • - Use uv init --no-package wk03-hw for your project

Next week: Neural networks -- Classification and Regression