model.evaluate(), model.predict(), summary + homeworktf.kerasKeras = friendly interface / TensorFlow = powerful engine underneath
# 1. Create new project $ uv init --no-package wk03 $ cd wk03 # 2. Add TensorFlow $ uv add tensorflow # 3. Run your code $ uv run python main.py
- uv init --no-package wk03 = create project folder with config
- uv add tensorflow = install TF + dependencies into venv
- uv run = run in the project's venv (no "wrong env" bugs)
- No need to manually create venv -- uv handles it
[Tip] Project structure: wk03/ contains main.py, pyproject.toml, uv.lock
# main.py
import tensorflow as tf
from tensorflow import keras
print("TensorFlow:", tf.__version__)
print("Keras:", keras.__version__)
# Quick test: tensor operations
x = tf.constant([[1.0, 2.0],
[3.0, 4.0]])
print(x @ x.T) # matrix multiply
- import tensorflow as tf = standard import
- from tensorflow import keras = Keras API
- TF tensor works like NumPy array + GPU support
- If you see version numbers = ready to go
$ uv run main.py TensorFlow: 2.x.x Keras: 3.x.x tf.Tensor( [[ 7. 10.] [15. 22.]], shape=(2,2))
from tensorflow import keras
from keras import layers
model = keras.Sequential([
# Hidden layer 1
layers.Dense(64, activation='relu',
input_shape=(10,)),
# Hidden layer 2
layers.Dense(32, activation='relu'),
# Output layer
layers.Dense(1, activation='sigmoid')
])
model.summary()
- Sequential() = stack layers one after another
- Dense(units) = fully connected layer (every neuron connects to every neuron in next layer)
- activation = non-linearity function
- input_shape=(10,) = tell the model the shape of one input sample
[Tip] First layer must have input_shape. Later layers infer automatically.
Input(10) Dense(64) Dense(32) Dense(1) [x1] --> [h1] -----> [h3] ----> [y] [x2] --> [h2] -----> [h4] ... ... ... [x10] --> [h64]
Each sample has 10 features
e.g. 10 measurements of a patient
64 neurons, each learns a pattern
Params: 10*64 + 64 = 704
1 neuron, output probability
Binary classification (0 or 1)
More neurons = more capacity to learn, but also more risk of overfitting
\( \sigma(z) = \frac{1}{1+e^{-z}} \)
Output: (0, 1) = probability
Used in output layer for binary classification
\( \sigma(z_i) = \frac{e^{z_i}}{\sum e^{z_j}} \)
Output sums to 1 (probability distribution)
Used for multi-class classification
Hidden layers: use ReLU / Output: match the task
model = keras.Sequential([
layers.Dense(64, activation='relu',
input_shape=(10,)),
layers.Dense(32, activation='relu'),
layers.Dense(1, activation='sigmoid')
])
model.summary()
# Layer (type) Output Param
# dense (Dense) (None,64) 704 <-- 10*64+64
# dense_1 (Dense) (None,32) 2080 <-- 64*32+32
# dense_2 (Dense) (None,1) 33 <-- 32*1+1
# Total params: 2,817
- Formula: Dense(in, out) = in*out + out (weights + bias)
- Layer 1: 10*64 + 64 = 704
- Layer 2: 64*32 + 32 = 2080
- Layer 3: 32*1 + 1 = 33
- Total: 2,817 trainable parameters
[Tip] model.summary() = instant parameter count -- use it to sanity-check model size
model.compile(
optimizer='adam',
loss='binary_crossentropy',
metrics=['accuracy']
)
- optimizer = algorithm that adjusts weights
- Adam = adaptive learning rate, good default
- SGD = classic, slower but predictable
- loss = how wrong the model is
- binary_crossentropy = binary classification
- categorical_crossentropy = multi-class
- mse = regression
- metrics = what to monitor (not used for training)
model.fit() -- train the modelimport numpy as np
# Sample data: 1000 samples, 10 features
X_train = np.random.randn(1000, 10)
y_train = (X_train[:, 0] > 0).astype(float)
# Train!
history = model.fit(
X_train, y_train,
epochs=20,
batch_size=32,
validation_split=0.2
)
- epochs = how many times to see all data
- batch_size = samples per gradient update
- validation_split = hold out 20% for validation
- history = dict with loss/metrics per epoch
Epoch 1/20 - 25/25 - 0s - loss: 0.72 - accuracy: 0.51 Epoch 2/20 - 25/25 - 0s - loss: 0.48 - accuracy: 0.82 ... Epoch 20/20 - 25/25 - 0s - loss: 0.21 - accuracy: 0.93
[Tip] Watch val_loss -- if it goes up while loss goes down = overfitting
# history contains training curves
import matplotlib.pyplot as plt
plt.plot(history.history['loss'],
label='train loss')
plt.plot(history.history['val_loss'],
label='val loss')
plt.xlabel('Epoch')
plt.ylabel('Loss')
plt.legend()
plt.show()
# Also:
# history.history['accuracy']
# history.history['val_accuracy']
- loss decreasing = model is learning
- val_loss decreasing = generalizes to unseen data
- loss down, val_loss up = overfitting
- both flat = model is too simple (underfitting)
[Tip] Early stopping: stop when val_loss stops improving. Use keras.callbacks.EarlyStopping()
model.evaluate() -- test performance# Test data (unseen during training)
X_test = np.random.randn(200, 10)
y_test = (X_test[:, 0] > 0).astype(float)
# Evaluate
loss, accuracy = model.evaluate(X_test, y_test)
print(f"Test loss: {loss:.4f}")
print(f"Test accuracy: {accuracy:.4f}")
# Output:
# 7/7 [======] - 0s - loss: 0.19 - accuracy: 0.94
# Test loss: 0.1923
# Test accuracy: 0.9400
- evaluate() = run model on test set, return loss + metrics
- Uses the same loss/metrics from compile()
- Returns tuple: (loss, metric1, metric2, ...)
- Test accuracy should be close to validation accuracy
[Tip] Never peek at test data during training -- it's your final exam, not practice
model.predict() -- use the model# New data
X_new = np.array([[0.5, -0.3, 0.8, 0.1,
-0.2, 0.6, -0.1, 0.4,
0.9, -0.5]])
# Predict
pred = model.predict(X_new)
print(pred) # [[0.87]]
print(pred[0][0]) # 0.87 = probability
# Convert to class
if pred[0][0] > 0.5:
print("Class 1")
else:
print("Class 0")
- predict() = forward pass on new data
- Returns probability (if sigmoid output)
- Threshold: pred > 0.5 = class 1
- For multi-class: np.argmax(pred, axis=1)
[Tip] In production, save model: model.save('my_model.keras')
import numpy as np
from tensorflow import keras
from keras import layers
# 1. Data
X = np.random.randn(1000, 10)
y = (X[:, 0] * X[:, 1] > 0).astype(float)
# 2. Model
model = keras.Sequential([
layers.Dense(64, activation='relu', input_shape=(10,)),
layers.Dense(32, activation='relu'),
layers.Dense(1, activation='sigmoid')
])
# 3. Compile
model.compile(optimizer='adam',
loss='binary_crossentropy',
metrics=['accuracy'])
# 4. Train
model.fit(X[:800], y[:800], epochs=20, validation_split=0.2)
# 5. Evaluate
loss, acc = model.evaluate(X[800:], y[800:])
print(f"Test accuracy: {acc:.4f}")
# 6. Predict
preds = model.predict(X[800:5])
print("Predictions:", (preds > 0.5).astype(int).ravel())
model = Sequential([
Dense(64, relu, input_shape=(10,)),
Dense(1, sigmoid)])
model.compile(optimizer='adam',
loss='binary_crossentropy')
model.fit(X, y, epochs=10)
+ Simple, fast prototyping
+ Good defaults, less boilerplate
model = nn.Sequential(
nn.Linear(10, 64), nn.ReLU(),
nn.Linear(64, 1), nn.Sigmoid())
opt = optim.Adam(model.parameters())
for epoch in range(10):
# ... manual training loop
+ Full control, easier to debug
+ Preferred in research
Learn both: Keras first (easy start) -> PyTorch later (deeper understanding)
uv init --no-package -> uv add tensorflowcompile(optimizer, loss, metrics)fit() -> evaluate() -> predict()keras.datasets.mnist)uv init --no-package wk03-hw for your projectNext week: Neural networks -- Classification and Regression