Tensor = กล่องเก็บตัวเลขหลายมิติ (เหมือน NumPy array)
x = 3.14
ตัวเลขเดียว
ndim = 0
x = [1, 2, 3]
แถวเดียว
ndim = 1, shape=(3,)
x = [[1,2],[3,4]]
ตาราง
ndim = 2, shape=(2,2)
x = [[[1,2],[3,4]]]
กล่องซ้อนกล่อง
ndim = 3, shape=(1,2,2)
ใน Deep Learning: รูปภาพ = 3D tensor (height, width, channels) / batch ข้อมูล = 4D tensor
import numpy as np # สร้าง tensor หลายแบบ x = np.array([1, 2, 3]) # vector y = np.zeros((3, 4)) # matrix 3x4, เติม 0 z = np.ones((2, 3, 4)) # 3D tensor, เติม 1 r = np.random.randn(5, 5) # normal distribution # สอบถาม shape print(x.shape) # (3,) print(y.ndim) # 2 print(z.dtype) # float64
- np.array() = สร้างจาก list
- np.zeros() / np.ones() = สร้าง tensor เปล่า
- np.random.randn() = สุ่มจาก normal distribution
- shape = ขนาดในแต่ละมิติ (tuple)
- ndim = จำนวนมิติ
- dtype = ชนิดข้อมูล (float32, float64, int64)
[Tip] ใน DL มักใช้ float32 เพราะประหยัด memory + เร็วกว่าบน GPU
import torch # สร้าง tensor (เหมือน NumPy แต่เพิ่ม requires_grad) x = torch.tensor([1.0, 2.0, 3.0]) y = torch.zeros(3, 4) z = torch.ones(2, 3, 4) r = torch.randn(5, 5) # 询问属性 print(x.shape) # torch.Size([3]) print(x.dtype) # torch.float32 # 的关键: requires_grad w = torch.tensor(2.0, requires_grad=True) # บอก PyTorch: "ช่วยจำ operation ทุกอย่างที่ทำกับ w ด้วย"
- PyTorch tensor = NumPy array + autograd
- requires_grad=True = เปิดคำนวณ gradient อัตโนมัติ
- shape, ndim, dtype ใช้เหมือน NumPy
- เปลี่ยนจาก NumPy: tensor.numpy()
- เปลี่ยนเป็น NumPy: torch.from_numpy(arr)
[Tip] PyTorch tensor + NumPy array มี API ใกล้เคียงกันมาก -- ถ้ารู้ NumPy ก็ใช้ PyTorch ได้เลย
x = torch.arange(12) # tensor([ 0, 1, 2, 3, 4, 5, # 6, 7, 8, 9, 10, 11]) # reshape x.reshape(3, 4) # -> 3 rows x 4 cols x.reshape(4, 3) # -> 4 rows x 3 cols x.reshape(2, 2, 3) # -> 3D tensor x.reshape(-1, 4) # -1 = ให้ NumPy คำนวณเอง # flatten (2D -> 1D) x.flatten() # -> tensor([0..11]) # transpose m = torch.arange(6).reshape(2, 3) m.T # -> shape (3, 2)
- reshape() = เปลี่ยน shape โดยไม่เปลี่ยนค่า
- -1 = ให้ framework คำนวณมิตินั้นให้ (ต้องหารกันลงตัว)
- flatten() = ยุบเป็น 1 มิติ
- .T = transpose (สลับมิติ)
[Tip] ใน DL มัก reshape ก่อนป้อนเข้าโมเดล เช่น รูป (28,28) -> flatten เป็น vector ขนาด 784
a = torch.tensor([1.0, 2.0, 3.0]) b = torch.tensor([4.0, 5.0, 6.0]) # arithmetic a + b # -> [5., 7., 9.] a * b # -> [4., 10., 18.] a - b # -> [-3., -3., -3.] a / b # -> [0.25, 0.4, 0.5] # power a ** 2 # -> [1., 4., 9.] # comparison a > 2 # -> [False, False, True] # math functions torch.exp(a) torch.log(a) torch.abs(a - b)
- ทุก operation ทำ ทีละ element (ขนานกัน)
- ผลลัพธ์ = tensor ใหม่ (ไม่เปลี่ยนตัวเดิม)
- คล้าย NumPy: a + b ทำงานเหมือนกัน
- ใน DL: activation function = element-wise operation
- Math operations มีให้เลือกหลายแบบ
dot(a, b) = a[0]*b[0] + a[1]*b[1] + ... = \( \sum_i a_i \cdot b_i \)
a = torch.tensor([1.0, 2.0, 3.0]) b = torch.tensor([4.0, 5.0, 6.0]) # dot product = scalar เดียว torch.dot(a, b) # -> 1*4 + 2*5 + 3*6 = 32 # manual version (a * b).sum() # -> tensor(32.) # matrix-vector: ใช้ @ operator W = torch.randn(3, 4) # matrix 3x4 x = torch.randn(4) # vector ขนาด 4 y = W @ x # -> vector ขนาด 3 # y[i] = dot(W[i], x)
- Dot product = วัด "ความคล้าย" ระหว่าง 2 vectors
- ใน neural network: z = w @ x + b = weighted sum
- @ operator = matrix multiply (ใน Python 3.5+)
- ถ้า a และ b ชี้ทิศเดียวกัน = dot product มาก
[Tip] Cosine similarity = dot product / (norm(a) * norm(b)) = dot product ที่ normalize แล้ว
# ข้อมูล N จุด, แต่ละจุดมี 3 features
X = torch.randn(5, 3) # (N=5, features=3)
# weight matrix: 3 inputs -> 2 outputs
W = torch.randn(3, 2) # (in=3, out=2)
# matrix multiply: ครั้งเดียวได้ 5 จุด
Z = X @ W # (5,3) @ (3,2) = (5,2)
# เทียบกับ loop
for i in range(5):
z[i] = X[i] @ W # <-- ช้ามาก!
- Matrix multiply = dot product หลายคู่พร้อมกัน
- (N, K) @ (K, M) = (N, M) -- K ต้องตรงกัน
- ทำ 1 ครั้งได้ทุกจุด = เร็วกว่า loop 100x
- นี่คือเหตุผลที่ DL ต้องการ GPU -- GPU คำนวณ matrix ได้ parallel
[Tip] nn.Linear ใน PyTorch = ทำ x @ W.T + b ให้เราอัตโนมัติ
# บวก matrix กับ vector m = torch.ones(3, 4) # (3, 4) v = torch.tensor([1., 2., 3., 4.]) # (4,) # Broadcasting: v ถูก copy เป็น 3 แถว result = m + v # -> [[2., 3., 4., 5.], # [2., 3., 4., 5.], # [2., 3., 4., 5.]] # broadcast ได้เมื่อ: # - มิติตรงกัน, หรือ # - มิติใดมิติหนึ่งเป็น 1 a = torch.ones(3, 1) # (3, 1) b = torch.ones(1, 4) # (1, 4) c = a + b # (3, 4) -- ทั้งสองมิติ broadcast
- Broadcasting = "คัดลอก" tensor ให้ตรงมิติอัตโนมัติ
- ไม่ต้องเขียน loop -- เร็วกว่ามาก
- กฎ: มิติจากหลังไปหน้า ต้องตรงกัน หรือเป็น 1
- ตัวอย่าง: เพิ่ม bias ทุกจุดพร้อมกัน
[Tip] ถ้า error บอก "operands could not be broadcast together" = shape ไม่ตรงตามกฎ -- ต้อง reshape ก่อน
x = torch.tensor([[1., 2.],
[3., 4.]])
x.sum() # -> 10. (รวมทุกค่า)
x.mean() # -> 2.5
x.max() # -> 4.
x.min() # -> 1.
# along axis
x.sum(axis=0) # -> [4., 6.] (รวมตาม column)
x.sum(axis=1) # -> [3., 7.] (รวมตาม row)
# argmax = index ของค่ามากสุด
x.argmax() # -> 3 (index ของ 4.0)
- axis=0 = รวมตามแนวตั้ง (ลด row)
- axis=1 = รวมตามแนวนอน (ลด column)
- ใน DL: loss function มัก sum/mean ทั้ง batch
- argmax() = หา class ที่ทำนายได้มากสุด (classification)
\(f'(x) = \lim_{h \to 0} \frac{f(x+h) - f(x)}{h}\)
Gradient = vector ของ partial derivatives = ทิศทางที่ function ชันมากที่สุด
# ตัวอย่าง: f(x) = x^2
# f'(x) = 2x
def f(x):
return x ** 2
# numerical derivative
def numerical_grad(f, x, h=1e-4):
return (f(x + h) - f(x - h)) / (2 * h)
print(numerical_grad(f, 3.0)) # ~6.0
print(numerical_grad(f, -2.0)) # ~-4.0
- Gradient = ทิศทางที่ function ไต่ขึ้นเร็วที่สุด
- ใน DL: gradient ของ loss respect w = "ขยับ w แล้ว loss เปลี่ยนเท่าไหร่"
- gradient > 0 = ขยับ w ขึ้น -> loss ขึ้น -> ต้องขยับลง
- gradient < 0 = ขยับ w ลง -> loss ขึ้น -> ต้องขยับขึ้น
[Tip] Gradient descent = ไป ทวน gradient (-negative direction) = loss ลดลง
import torch x = torch.tensor(3.0, requires_grad=True) # forward pass: สร้าง computational graph y = x ** 2 # y = x^2 z = 2 * y + 1 # z = 2x^2 + 1 # backward pass: คำนวณ gradient อัตโนมัติ z.backward() # dz/dx = 4x = 4*3 = 12 print(x.grad) # tensor(12.) # เปรียบเทียบ: numerical vs autodiff # numerical: (f(x+h) - f(x-h)) / 2h # autodiff: แม่นยำ 100%, เร็ว, ไม่ต้องเลือก h
- requires_grad=True = ให้ PyTorch ติดตาม operation
- .backward() = เดินย้อน graph คำนวณ gradient ทุกจุด
- .grad = เก็บค่า gradient ไว้ให้
- ข้อดี: แม่นยำ + เร็ว + ไม่ต้องเขียนสูตรเอง
- ข้อเสีย: ต้อง forward pass ก่อนถึงจะ backward ได้
[Tip] Autodiff = หัวใจของ deep learning framework ทุกตัว (PyTorch, TensorFlow, JAX)
\( \frac{df}{dx} = \frac{df}{dg} \cdot \frac{dg}{dx} \)
ถ้า \( f = f(g(x)) \) แล้วก็ \( \frac{df}{dx} \) = คูณ gradient ทีละชั้น
# ตัวอย่าง: y = (2x + 1)^2 # LET u = 2x + 1, y = u^2 # _dy/du = 2u, du/dx = 2 # _dy/dx = 2u * 2 = 4u = 4(2x+1) import torch x = torch.tensor(1.0, requires_grad=True) u = 2 * x + 1 # u = 3 y = u ** 2 # y = 9 y.backward() print(x.grad) # tensor(12.) = 4*3
- Chain rule = หัวใจของ backpropagation
- Neural network = function ซ้อน function หลายชั้น
- backward() ใช้ chain rule คูณ gradient ย้อนกลับทีละชั้น
- ชั้น越多 = chain rule ยิ่งยาว -- นี่คือ "deep" ใน deep learning
[Tip] ไม่ต้องเขียน chain rule เอง -- .backward() ทำให้หมด!
# tiny model: y = sigmoid(w @ x + b)
import torch
# ข้อมูล: 4 จุด, 2 features
X = torch.randn(4, 2)
y_true = torch.tensor([[1.],[0.],[1.],[0.]])
# parameters (สุ่ม)
w = torch.randn(2, 1, requires_grad=True)
b = torch.zeros(1, requires_grad=True)
# training loop (10 รอบ)
lr = 0.1
for i in range(10):
# forward
z = X @ w + b
y_pred = torch.sigmoid(z)
loss = ((y_pred - y_true)**2).mean()
# backward + update
loss.backward()
w.data -= lr * w.grad
b.data -= lr * b.grad
w.grad.zero_()
b.grad.zero_()
print(f"round {i}: loss={loss.item():.4f}")
- ทุกอย่างเป็น tensor -- ไม่ต้องมี framework
- forward: z = X @ w + b แล้ว sigmoid
- backward: loss.backward() คำนวณ gradient ทุกตัว
- update: w.data -= lr * w.grad ปรับ parameter
- ล้าง gradient: w.grad.zero_()
[Tip] โค้ดนี้ = ต้นแบบของ training loop ที่จะใช้ตลอดคอร์ส -- เข้าใจตรงนี้ = เข้าใจ DL
รู้ NumPy = พร้อมใช้ PyTorch เลย / รู้ Tensor ทั้ง 2 = พร้อมเข้าโลก DL
สัปดาห์หน้า: Introduction to Keras and TensorFlow