แผนการเรียนรู้
/

Mathematical Building Blocks

tensors - operations - gradients - differentiation

ผลลัพธ์การเรียนรู้

1. สร้างและจัดการ tensor ได้
2. ทำการคำนวณบน tensor (arithmetic, dot product, reshaping)
3. เข้าใจ broadcasting ใน NumPy/PyTorch
4. คำนวณ gradient และเข้าใจ autodiff
สัปดาห์ที่ 2 -- 1145208 การเรียนรู้เชิงลึก

เส้นทางของวันนี้

Part 1 - Data Structure
Tensor คืออะไร?
scalar, vector, matrix, tensor -- สร้าง จัดรูป แปลงมิติ
Part 2 - Operations
คำนวณบน Tensor
element-wise, dot product, matrix multiply, broadcasting
Part 3 - Calculus
Gradient และ Autodiff
partial derivative, chain rule, computational graph
Part 4 - tying it together
จากสมการสู่โค้ด
step-by-step: build a tiny model with tensors only

Part 1: Tensor คืออะไร?

Tensor = กล่องเก็บตัวเลขหลายมิติ (เหมือน NumPy array)

Scalar (0D)

x = 3.14

ตัวเลขเดียว

ndim = 0

Vector (1D)

x = [1, 2, 3]

แถวเดียว

ndim = 1, shape=(3,)

Matrix (2D)

x = [[1,2],[3,4]]

ตาราง

ndim = 2, shape=(2,2)

3D+ Tensor

x = [[[1,2],[3,4]]]

กล่องซ้อนกล่อง

ndim = 3, shape=(1,2,2)

ใน Deep Learning: รูปภาพ = 3D tensor (height, width, channels) / batch ข้อมูล = 4D tensor

สร้าง Tensor ด้วย NumPy

import numpy as np

# สร้าง tensor หลายแบบ
x = np.array([1, 2, 3])          # vector
y = np.zeros((3, 4))             # matrix 3x4, เติม 0
z = np.ones((2, 3, 4))           # 3D tensor, เติม 1
r = np.random.randn(5, 5)        # normal distribution

# สอบถาม shape
print(x.shape)     # (3,)
print(y.ndim)      # 2
print(z.dtype)     # float64

- np.array() = สร้างจาก list

- np.zeros() / np.ones() = สร้าง tensor เปล่า

- np.random.randn() = สุ่มจาก normal distribution

- shape = ขนาดในแต่ละมิติ (tuple)

- ndim = จำนวนมิติ

- dtype = ชนิดข้อมูล (float32, float64, int64)

[Tip] ใน DL มักใช้ float32 เพราะประหยัด memory + เร็วกว่าบน GPU

Tensor ใน PyTorch

import torch

# สร้าง tensor (เหมือน NumPy แต่เพิ่ม requires_grad)
x = torch.tensor([1.0, 2.0, 3.0])
y = torch.zeros(3, 4)
z = torch.ones(2, 3, 4)
r = torch.randn(5, 5)

# 询问属性
print(x.shape)       # torch.Size([3])
print(x.dtype)       # torch.float32

# 的关键: requires_grad
w = torch.tensor(2.0, requires_grad=True)
# บอก PyTorch: "ช่วยจำ operation ทุกอย่างที่ทำกับ w ด้วย"

- PyTorch tensor = NumPy array + autograd

- requires_grad=True = เปิดคำนวณ gradient อัตโนมัติ

- shape, ndim, dtype ใช้เหมือน NumPy

- เปลี่ยนจาก NumPy: tensor.numpy()

- เปลี่ยนเป็น NumPy: torch.from_numpy(arr)

[Tip] PyTorch tensor + NumPy array มี API ใกล้เคียงกันมาก -- ถ้ารู้ NumPy ก็ใช้ PyTorch ได้เลย

Reshaping Tensor: เปลี่ยนรูปแบบไม่เปลี่ยนค่า

x = torch.arange(12)
# tensor([ 0,  1,  2,  3,  4,  5,
#           6,  7,  8,  9, 10, 11])

# reshape
x.reshape(3, 4)    # -> 3 rows x 4 cols
x.reshape(4, 3)    # -> 4 rows x 3 cols
x.reshape(2, 2, 3) # -> 3D tensor
x.reshape(-1, 4)   # -1 = ให้ NumPy คำนวณเอง

# flatten (2D -> 1D)
x.flatten()         # -> tensor([0..11])

# transpose
m = torch.arange(6).reshape(2, 3)
m.T                 # -> shape (3, 2)

- reshape() = เปลี่ยน shape โดยไม่เปลี่ยนค่า

- -1 = ให้ framework คำนวณมิตินั้นให้ (ต้องหารกันลงตัว)

- flatten() = ยุบเป็น 1 มิติ

- .T = transpose (สลับมิติ)

[Tip] ใน DL มัก reshape ก่อนป้อนเข้าโมเดล เช่น รูป (28,28) -> flatten เป็น vector ขนาด 784

Part 2: Element-wise Operations

a = torch.tensor([1.0, 2.0, 3.0])
b = torch.tensor([4.0, 5.0, 6.0])

# arithmetic
a + b       # -> [5., 7., 9.]
a * b       # -> [4., 10., 18.]
a - b       # -> [-3., -3., -3.]
a / b       # -> [0.25, 0.4, 0.5]

# power
a ** 2      # -> [1., 4., 9.]

# comparison
a > 2       # -> [False, False, True]

# math functions
torch.exp(a)
torch.log(a)
torch.abs(a - b)

- ทุก operation ทำ ทีละ element (ขนานกัน)

- ผลลัพธ์ = tensor ใหม่ (ไม่เปลี่ยนตัวเดิม)

- คล้าย NumPy: a + b ทำงานเหมือนกัน

- ใน DL: activation function = element-wise operation

- Math operations มีให้เลือกหลายแบบ

Dot Product -- หัวใจของ neural network

dot(a, b) = a[0]*b[0] + a[1]*b[1] + ... = \( \sum_i a_i \cdot b_i \)

a = torch.tensor([1.0, 2.0, 3.0])
b = torch.tensor([4.0, 5.0, 6.0])

# dot product = scalar เดียว
torch.dot(a, b)
# -> 1*4 + 2*5 + 3*6 = 32

# manual version
(a * b).sum()
# -> tensor(32.)

# matrix-vector: ใช้ @ operator
W = torch.randn(3, 4)   # matrix 3x4
x = torch.randn(4)      # vector ขนาด 4
y = W @ x               # -> vector ขนาด 3
# y[i] = dot(W[i], x)

- Dot product = วัด "ความคล้าย" ระหว่าง 2 vectors

- ใน neural network: z = w @ x + b = weighted sum

- @ operator = matrix multiply (ใน Python 3.5+)

- ถ้า a และ b ชี้ทิศเดียวกัน = dot product มาก

[Tip] Cosine similarity = dot product / (norm(a) * norm(b)) = dot product ที่ normalize แล้ว

Matrix Multiplication: เทรนทีเดียวหลายจุด

# ข้อมูล N จุด, แต่ละจุดมี 3 features
X = torch.randn(5, 3)   # (N=5, features=3)

# weight matrix: 3 inputs -> 2 outputs
W = torch.randn(3, 2)   # (in=3, out=2)

# matrix multiply: ครั้งเดียวได้ 5 จุด
Z = X @ W               # (5,3) @ (3,2) = (5,2)

# เทียบกับ loop
for i in range(5):
    z[i] = X[i] @ W     # <-- ช้ามาก!

- Matrix multiply = dot product หลายคู่พร้อมกัน

- (N, K) @ (K, M) = (N, M) -- K ต้องตรงกัน

- ทำ 1 ครั้งได้ทุกจุด = เร็วกว่า loop 100x

- นี่คือเหตุผลที่ DL ต้องการ GPU -- GPU คำนวณ matrix ได้ parallel

[Tip] nn.Linear ใน PyTorch = ทำ x @ W.T + b ให้เราอัตโนมัติ

Broadcasting: เติมมิติอัตโนมัติ

# บวก matrix กับ vector
m = torch.ones(3, 4)    # (3, 4)
v = torch.tensor([1., 2., 3., 4.])  # (4,)

# Broadcasting: v ถูก copy เป็น 3 แถว
result = m + v
# -> [[2., 3., 4., 5.],
#      [2., 3., 4., 5.],
#      [2., 3., 4., 5.]]

# broadcast ได้เมื่อ:
# - มิติตรงกัน, หรือ
# - มิติใดมิติหนึ่งเป็น 1

a = torch.ones(3, 1)    # (3, 1)
b = torch.ones(1, 4)    # (1, 4)
c = a + b               # (3, 4) -- ทั้งสองมิติ broadcast

- Broadcasting = "คัดลอก" tensor ให้ตรงมิติอัตโนมัติ

- ไม่ต้องเขียน loop -- เร็วกว่ามาก

- กฎ: มิติจากหลังไปหน้า ต้องตรงกัน หรือเป็น 1

- ตัวอย่าง: เพิ่ม bias ทุกจุดพร้อมกัน

[Tip] ถ้า error บอก "operands could not be broadcast together" = shape ไม่ตรงตามกฎ -- ต้อง reshape ก่อน

Aggregation: รวม tensor เป็นค่าเดียว

x = torch.tensor([[1., 2.],
                   [3., 4.]])

x.sum()           # -> 10. (รวมทุกค่า)
x.mean()          # -> 2.5
x.max()           # -> 4.
x.min()           # -> 1.

# along axis
x.sum(axis=0)     # -> [4., 6.]  (รวมตาม column)
x.sum(axis=1)     # -> [3., 7.]  (รวมตาม row)

# argmax = index ของค่ามากสุด
x.argmax()        # -> 3 (index ของ 4.0)

- axis=0 = รวมตามแนวตั้ง (ลด row)

- axis=1 = รวมตามแนวนอน (ลด column)

- ใน DL: loss function มัก sum/mean ทั้ง batch

- argmax() = หา class ที่ทำนายได้มากสุด (classification)

Part 3: Gradient คืออะไร?

\(f'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}\)

Gradient = vector ของ partial derivatives = ทิศทางที่ function ชันมากที่สุด

# ตัวอย่าง: f(x) = x^2
# f'(x) = 2x

def f(x):
    return x ** 2

# numerical derivative
def numerical_grad(f, x, h=1e-4):
    return (f(x + h) - f(x - h)) / (2 * h)

print(numerical_grad(f, 3.0))  # ~6.0
print(numerical_grad(f, -2.0)) # ~-4.0

- Gradient = ทิศทางที่ function ไต่ขึ้นเร็วที่สุด

- ใน DL: gradient ของ loss respect w = "ขยับ w แล้ว loss เปลี่ยนเท่าไหร่"

- gradient > 0 = ขยับ w ขึ้น -> loss ขึ้น -> ต้องขยับลง

- gradient < 0 = ขยับ w ลง -> loss ขึ้น -> ต้องขยับขึ้น

[Tip] Gradient descent = ไป ทวน gradient (-negative direction) = loss ลดลง

Automatic Differentiation: ให้เครื่องคำนวณให้

import torch

x = torch.tensor(3.0, requires_grad=True)

# forward pass: สร้าง computational graph
y = x ** 2          # y = x^2
z = 2 * y + 1       # z = 2x^2 + 1

# backward pass: คำนวณ gradient อัตโนมัติ
z.backward()

# dz/dx = 4x = 4*3 = 12
print(x.grad)       # tensor(12.)

#  เปรียบเทียบ: numerical vs autodiff
#  numerical:  (f(x+h) - f(x-h)) / 2h
#  autodiff:    แม่นยำ 100%, เร็ว, ไม่ต้องเลือก h

- requires_grad=True = ให้ PyTorch ติดตาม operation

- .backward() = เดินย้อน graph คำนวณ gradient ทุกจุด

- .grad = เก็บค่า gradient ไว้ให้

- ข้อดี: แม่นยำ + เร็ว + ไม่ต้องเขียนสูตรเอง

- ข้อเสีย: ต้อง forward pass ก่อนถึงจะ backward ได้

[Tip] Autodiff = หัวใจของ deep learning framework ทุกตัว (PyTorch, TensorFlow, JAX)

Chain Rule: คูณ gradient ย้อนกลับ

\( \frac{df}{dx} = \frac{df}{dg} \cdot \frac{dg}{dx} \)

ถ้า \( f = f(g(x)) \) แล้วก็ \( \frac{df}{dx} \) = คูณ gradient ทีละชั้น

# ตัวอย่าง: y = (2x + 1)^2
#  LET u = 2x + 1,  y = u^2

# _dy/du = 2u,  du/dx = 2
#  _dy/dx = 2u * 2 = 4u = 4(2x+1)

import torch
x = torch.tensor(1.0, requires_grad=True)
u = 2 * x + 1     # u = 3
y = u ** 2         # y = 9
y.backward()
print(x.grad)      # tensor(12.) = 4*3

- Chain rule = หัวใจของ backpropagation

- Neural network = function ซ้อน function หลายชั้น

- backward() ใช้ chain rule คูณ gradient ย้อนกลับทีละชั้น

- ชั้น越多 = chain rule ยิ่งยาว -- นี่คือ "deep" ใน deep learning

[Tip] ไม่ต้องเขียน chain rule เอง -- .backward() ทำให้หมด!

Part 4: สร้างโมเดลเล็กๆ ด้วย Tensor ล้วน

# tiny model: y = sigmoid(w @ x + b)
import torch

# ข้อมูล: 4 จุด, 2 features
X = torch.randn(4, 2)
y_true = torch.tensor([[1.],[0.],[1.],[0.]])

# parameters (สุ่ม)
w = torch.randn(2, 1, requires_grad=True)
b = torch.zeros(1, requires_grad=True)

# training loop (10 รอบ)
lr = 0.1
for i in range(10):
    # forward
    z = X @ w + b
    y_pred = torch.sigmoid(z)
    loss = ((y_pred - y_true)**2).mean()
    # backward + update
    loss.backward()
    w.data -= lr * w.grad
    b.data -= lr * b.grad
    w.grad.zero_()
    b.grad.zero_()
    print(f"round {i}: loss={loss.item():.4f}")

- ทุกอย่างเป็น tensor -- ไม่ต้องมี framework

- forward: z = X @ w + b แล้ว sigmoid

- backward: loss.backward() คำนวณ gradient ทุกตัว

- update: w.data -= lr * w.grad ปรับ parameter

- ล้าง gradient: w.grad.zero_()

[Tip] โค้ดนี้ = ต้นแบบของ training loop ที่จะใช้ตลอดคอร์ส -- เข้าใจตรงนี้ = เข้าใจ DL

เปรียบเทียบ: NumPy vs PyTorch

NumPy

  • - array operations สำหรับ data science
  • - ไม่มี GPU acceleration
  • - ไม่มี autograd
  • - เร็วพอสำหรับ data preprocessing
  • - API ที่ PyTorch เลียนแบบ

PyTorch

  • - GPU acceleration (CUDA, MPS)
  • - Autograd = คำนวณ gradient อัตโนมัติ
  • - Dynamic computational graph
  • - API คล้าย NumPy มาก
  • - เป็นที่นิยมใน research + production

รู้ NumPy = พร้อมใช้ PyTorch เลย / รู้ Tensor ทั้ง 2 = พร้อมเข้าโลก DL

สรุป + การบ้าน

สิ่งที่ทำวันนี้

  • - Tensor = กล่องเก็บตัวเลขหลายมิติ (scalar -> vector -> matrix -> 3D+)
  • - Operations: element-wise, dot product, matrix multiply, broadcasting
  • - Gradient = ทิศทางที่ function ชันมากที่สุด
  • - Autodiff + Chain rule = คำนวณ gradient ย้อนกลับอัตโนมัติ

การบ้าน (ส่งสัปดาห์หน้า)

  • - สร้าง tensor 5 แบบ แล้ว print shape, ndim, dtype
  • - คำนวณ dot product ระหว่าง 2 vectors ด้วยมือ + ด้วย torch
  • - เขียน numerical gradient สำหรับ \( f(x) = x^3 - 2x + 1 \) ที่ x=2
  • - รัน tiny model (สไลด์ 14) จน loss ลด -- ปรับ lr แล้วดูผลต่าง

สัปดาห์หน้า: Introduction to Keras and TensorFlow