สัปดาห์ที่ 12
Word Embeddings และ Sequence Models
1. Text Representation: One-hot, TF-IDF
2. Word Embeddings: Word2Vec, GloVe
3. PyTorch nn.Embedding
4. Text Classification with LSTM
5. Sentiment Analysis Example
- Variable length: ประโยคยาวสั้นไม่เท่ากัน
- Discrete: คำเป็น categorical ไม่ใช่ตัวเลข
- Sequential: ลำดับคำสำคัญ (\"not good\" != \"good\")
- Ambiguous: คำเดียวกันมีหลายความหมาย
[ปัญหา] Machine ไม่เข้าใจคำ = ต้องแปลงคำเป็นตัวเลขก่อน
# Text: variable length, discrete
texts = [
"I love this movie", # 4 words
"This is terrible", # 3 words
"The film was absolutely great" # 5 words
]
# Problem: NN ต้องการ fixed-size input
# Solution: Embeddings + Padding
# ลำดับสำคัญ!
"not good" -> negative
"good" -> positive
"not bad" -> positive (double negative)
# Vocabulary: {"I": 0, "love": 1, "this": 2, "movie": 3}
# One-hot vector (size = vocab_size)
"I" -> [1, 0, 0, 0]
"love" -> [0, 1, 0, 0]
"this" -> [0, 0, 1, 0]
"movie"-> [0, 0, 0, 1]
# Problem: sparse, high-dimensional
# 50K words -> 50K-dimensional vectors!
# All vectors are equidistant
# (no semantic similarity)
import torch
vocab = {"I": 0, "love": 1, "this": 2, "movie": 3}
one_hot = torch.zeros(4, 4)
for word, idx in vocab.items():
one_hot[idx, idx] = 1.0
print(one_hot)
# tensor([[1., 0., 0., 0.], <- "I"
# [0., 1., 0., 0.], <- "love"
# [0., 0., 1., 0.], <- "this"
# [0., 0., 0., 1.]]) <- "movie"
- One-hot: vector ที่มี 1 ที่ตำแหน่งของคำ
- ขนาดเท่ากับ vocabulary
- ปัญหา: sparse, ไม่มี semantic meaning
- ทุกคำอยู่ห่างกันเท่ากัน
[สำคัญ] One-hot ไม่จับความหมายของคำ = ต้องใช้ embeddings
- TF (Term Frequency): คำนี้บ่อยแค่ไหนใน document
- IDF (Inverse Document Frequency): คำนี้หายากแค่ไหน
- TF-IDF = TF x IDF: คำสำคัญได้ค่าสูง
- ดีกว่า one-hot เพราะจับ importance
[เคล็ดลับ] "the", "is" ได้ IDF ต่ำ (อยู่ในทุกเอกสาร), คำเฉพาะได้ IDF สูง
from sklearn.feature_extraction.text import TfidfVectorizer
docs = [
"I love this movie",
"This movie is great",
"I hate this movie"
]
vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(docs)
print(vectorizer.get_feature_names_out())
# ['great' 'hate' 'is' 'love' 'movie' 'this']
print(X.toarray())
# [[0. 0. 0. 0.82 0.58 0.0 ] <- "love"
# [0.62 0. 0.62 0. 0.44 0.44] <- "great"
# [0. 0.82 0. 0. 0.58 0.0 ]] <- "hate"
# "movie" appears in all docs -> low TF-IDF
# "love"/"hate" -> high TF-IDF (important)
- Embedding: dense vector ที่จับความหมายของคำ
- คำที่ความหมายใกล้กัน = vector ใกล้กัน
- มิติคงที่ (เช่น 100-300 dimensions)
- เรียนรู้จากข้อมูลจำนวนมาก
[ตัวอย่าง] king - man + woman = queen (semantic arithmetic)
# Embedding: dense, low-dimensional
# "love" -> [0.2, -0.5, 0.8, ..., 0.1] (300-d)
# "like" -> [0.3, -0.4, 0.7, ..., 0.2] (300-d)
# "hate" -> [-0.7, 0.6, -0.3, ..., -0.5] (300-d)
# Similarity via cosine
import numpy as np
love = np.array([0.2, -0.5, 0.8, 0.1])
like = np.array([0.3, -0.4, 0.7, 0.2])
hate = np.array([-0.7, 0.6, -0.3, -0.5])
def cosine_sim(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
print(f"love-like: {cosine_sim(love, like):.3f}")
print(f"love-hate: {cosine_sim(love, hate):.3f}")
# love-like: 0.987 (similar)
# love-hate: -0.952 (opposite)
# Word2Vec: predict context from word
# "The cat sat on the mat"
# -> predict "sat" from ["cat", "on"]
# -> or predict "cat" from context ["The", "sat"]
from gensim.models import Word2Vec
sentences = [
["I", "love", "deep", "learning"],
["deep", "learning", "is", "great"],
["I", "love", "natural", "language", "processing"],
]
model = Word2Vec(sentences, vector_size=100,
window=5, min_count=1)
# Get embedding
vec = model.wv["love"]
print(vec.shape) # (100,)
# Find similar words
print(model.wv.most_similar("love"))
# [('deep', 0.85), ('great', 0.72), ...]
# Semantic arithmetic
result = model.wv.most_similar(
positive=['king', 'woman'],
negative=['man'])
# [('queen', 0.92), ...]
- CBOW: predict word from context
- Skip-gram: predict context from word
- vector_size: ขนาด embedding (100-300)
- window: ขนาด context window
[สำคัญ] Word2Vec เรียนรู้จาก co-occurrence = คำที่อยู่ด้วยกันมีความหมายใกล้กัน
- GloVe: ใช้ global co-occurrence matrix
- ต่างจาก Word2Vec ที่ใช้ local context
- Pre-trained vectors สำเร็จรูป
- ขนาด: 6B tokens, 400K words, 50/100/200/300d
[เปรียบเทียบ] Word2Vec เร็วกว่า / GloVe จับ global statistics ได้ดีกว่า
# Load pre-trained GloVe
import numpy as np
def load_glove(path):
embeddings = {}
with open(path, 'r', encoding='utf-8') as f:
for line in f:
values = line.split()
word = values[0]
vector = np.array(values[1:], dtype='float32')
embeddings[word] = vector
return embeddings
glove = load_glove('glove.6B.100d.txt')
# Similar words
def most_similar(word, top_k=5):
vec = glove[word]
sims = {}
for w, v in glove.items():
if w != word:
sims[w] = np.dot(vec, v) / (
np.linalg.norm(vec) * np.linalg.norm(v))
return sorted(sims.items(), key=lambda x: -x[1])[:top_k]
print(most_similar("computer"))
# [('software', 0.85), ('pc', 0.82), ...]
import torch
import torch.nn as nn
# Create embedding layer
vocab_size = 10000
embed_dim = 128
embedding = nn.Embedding(vocab_size, embed_dim)
print(embedding.weight.shape)
# torch.Size([10000, 128])
# Convert words to indices
word_to_idx = {"I": 0, "love": 1, "deep": 2, "learning": 3}
sentence = ["I", "love", "deep", "learning"]
indices = torch.tensor([word_to_idx[w] for w in sentence])
print(indices) # tensor([0, 1, 2, 3])
# Get embeddings
embeds = embedding(indices)
print(embeds.shape) # torch.Size([4, 128])
# Use pre-trained GloVe
def from_pretrained(embeddings_matrix):
vocab_size, dim = embeddings_matrix.shape
layer = nn.Embedding(vocab_size, dim)
layer.weight = nn.Parameter(
torch.tensor(embeddings_matrix, dtype=torch.float32),
requires_grad=False) # Freeze weights
return layer
- nn.Embedding(vocab_size, embed_dim): lookup table
- Input: indices (LongTensor)
- Output: dense vectors
- requires_grad=False: freeze pre-trained vectors
[สำคัญ] ต้อง padding sequences ให้ยาวเท่ากันก่อน
- Input: sequence of word indices
- Embedding: convert to dense vectors
- RNN/LSTM/GRU: process sequence
- FC + Softmax: classify
[สถาปัตยกรรม] Embedding -> LSTM -> FC -> Sigmoid (binary) / Softmax (multi-class)
class TextClassifier(nn.Module):
def __init__(self, vocab_size, embed_dim,
hidden_dim, output_dim):
super().__init__()
self.embedding = nn.Embedding(vocab_size, embed_dim)
self.lstm = nn.LSTM(embed_dim, hidden_dim,
batch_first=True)
self.fc = nn.Linear(hidden_dim, output_dim)
def forward(self, x):
# x: (batch, seq_len) - word indices
embeds = self.embedding(x)
# embeds: (batch, seq_len, embed_dim)
out, (hn, cn) = self.lstm(embeds)
# out: (batch, seq_len, hidden_dim)
out = self.fc(hn[-1])
# out: (batch, output_dim)
return torch.sigmoid(out)
model = TextClassifier(10000, 128, 64, 1)
# Binary sentiment classification
from torch.nn.utils.rnn import pad_sequence
from torch.utils.data import Dataset, DataLoader
class TextDataset(Dataset):
def __init__(self, texts, labels, vocab):
self.texts = texts
self.labels = labels
self.vocab = vocab
def __len__(self):
return len(self.texts)
def __getitem__(self, idx):
tokens = self.texts[idx].split()
indices = [self.vocab.get(t, 0) for t in tokens]
return torch.tensor(indices), self.labels[idx]
def collate_fn(batch):
texts, labels = zip(*batch)
# Pad sequences to same length
texts_padded = pad_sequence(texts, batch_first=True,
padding_value=0)
labels = torch.tensor(labels, dtype=torch.float32)
return texts_padded, labels
dataset = TextDataset(texts, labels, vocab)
loader = DataLoader(dataset, batch_size=32,
collate_fn=collate_fn)
- Padding: เติม 0 ให้ sequences ยาวเท่ากัน
- pad_sequence: pad batch ให้ยาวเท่า max length
- collate_fn: กำหนดวิธีรวม samples เป็น batch
- Attention mask: บอก model ว่าตำแหน่งไหนเป็น padding
[สำคัญ] Padding position ไม่ควร effect prediction = ต้อง mask
class SentimentLSTM(nn.Module):
def __init__(self, vocab_size, embed_dim,
hidden_dim, output_dim):
super().__init__()
self.embedding = nn.Embedding(vocab_size, embed_dim)
self.lstm = nn.LSTM(embed_dim, hidden_dim,
num_layers=2, batch_first=True,
dropout=0.5)
self.fc = nn.Linear(hidden_dim, output_dim)
self.dropout = nn.Dropout(0.5)
def forward(self, x):
embeds = self.dropout(self.embedding(x))
out, (hn, cn) = self.lstm(embeds)
out = self.dropout(out[:, -1, :])
return self.fc(out)
# Training
model = SentimentLSTM(25000, 128, 64, 1)
criterion = nn.BCEWithLogitsLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
for epoch in range(10):
model.train()
for batch_text, batch_label in loader:
pred = model(batch_text).squeeze()
loss = criterion(pred, batch_label)
optimizer.zero_grad()
loss.backward()
optimizer.step()
- ใช้ Dropout ป้องกัน overfitting
- BCEWithLogitsLoss: sigmoid + BCE ในที่เดียว
- ใช้ out[:, -1, :]: hidden state ตัวสุดท้าย
- 2 layers LSTM + dropout=0.5
[เคล็ดลับ] ใช้ pre-trained embedding จะเร็วกว่า train จากศูนย์
- IMDB: 50K movie reviews (25K train, 25K test)
- Binary sentiment: positive / negative
- ข้อความยาวไม่เท่ากัน ( averaging 231 words)
- Standard benchmark สำหรับ text classification
[ข้อมูล] ใช้ vocab 25K words, pad to 256 tokens
from torchtext.datasets import IMDB
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator
tokenizer = get_tokenizer('basic_english')
def yield_tokens(data_iter):
for text, label in data_iter:
yield tokenizer(text)
# Load data
train_iter, test_iter = IMDB(split=('train', 'test'))
# Build vocab
vocab = build_vocab_from_iterator(
yield_tokens(train_iter),
specials=["<unk>", "<pad>"],
max_size=25000)
vocab.set_default_index(vocab["<unk>"])
# Text pipeline
text_pipeline = lambda x: vocab(tokenizer(x))
# Example
text = "This movie is great!"
indices = text_pipeline(text)
print(indices) # [123, 456, 789, 1011]
def train_epoch(model, loader, optimizer, criterion):
model.train()
total_loss = 0
for text, label in loader:
optimizer.zero_grad()
pred = model(text).squeeze()
loss = criterion(pred, label.float())
loss.backward()
torch.nn.utils.clip_grad_norm_(
model.parameters(), max_norm=1.0)
optimizer.step()
total_loss += loss.item()
return total_loss / len(loader)
def evaluate(model, loader, criterion):
model.eval()
total_loss = 0
correct = 0
with torch.no_grad():
for text, label in loader:
pred = model(text).squeeze()
loss = criterion(pred, label.float())
total_loss += loss.item()
predicted = (pred > 0.5).long()
correct += (predicted == label).sum().item()
accuracy = correct / len(loader.dataset)
return total_loss / len(loader), accuracy
# Run
for epoch in range(5):
loss = train_epoch(model, train_loader, optimizer, criterion)
val_loss, acc = evaluate(model, test_loader, criterion)
print(f"Epoch {epoch}: loss={loss:.4f} acc={acc:.4f}")
- clip_grad_norm_: ป้องกัน exploding gradients
- วัดทั้ง loss และ accuracy
- ใช้ model.eval() ตอน evaluate
- 5-10 epochs พอสำหรับ IMDB
[สำคัญ] Gradient clipping สำคัญมากสำหรับ RNN/LSTM
import matplotlib.pyplot as plt
from sklearn.metrics import classification_report
# Predict
model.eval()
all_preds = []
all_labels = []
with torch.no_grad():
for text, label in test_loader:
pred = model(text).squeeze()
predicted = (pred > 0.5).long()
all_preds.extend(predicted.tolist())
all_labels.extend(label.tolist())
# Classification report
print(classification_report(all_labels, all_preds,
target_names=['Negative', 'Positive']))
# Plot loss curves
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
plt.plot(train_losses, label='Train')
plt.plot(val_losses, label='Val')
plt.legend()
plt.title('Loss')
plt.subplot(1, 2, 2)
plt.plot(train_accs, label='Train')
plt.plot(val_accs, label='Val')
plt.legend()
plt.title('Accuracy')
plt.show()
# Predict custom text
def predict_sentiment(text):
model.eval()
indices = text_pipeline(text)
tensor = torch.tensor([indices])
with torch.no_grad():
pred = torch.sigmoid(model(tensor)).item()
return "Positive" if pred > 0.5 else "Negative"
print(predict_sentiment("This movie is amazing!"))
# Positive
- ใช้ classification_report ดู precision, recall, f1
- Plot loss/accuracy curves
- สร้าง predict_sentiment function
- ต้อง preprocessing เหมือนตอน train
[เมตริก] Accuracy 85-90% สำหรับ IMDB = ดีแล้ว
import numpy as np
def load_glove(path, vocab, embed_dim):
# Create embedding matrix
matrix = np.zeros((len(vocab), embed_dim))
with open(path, 'r', encoding='utf-8') as f:
for line in f:
word, vector = line.split(maxsplit=1)
if word in vocab:
idx = vocab[word]
matrix[idx] = np.fromstring(vector, sep=' ')
return matrix
# Load GloVe
glove_matrix = load_glove('glove.6B.100d.txt', vocab, 100)
# Set embedding weights
model.embedding.weight = nn.Parameter(
torch.tensor(glove_matrix, dtype=torch.float32),
requires_grad=True) # Fine-tune
# Or freeze
model.embedding.weight = nn.Parameter(
torch.tensor(glove_matrix, dtype=torch.float32),
requires_grad=False) # ไม่ train ใหม่
# Fine-tuning strategies:
# 1. Train from scratch: ต้องมีข้อมูลเยอะ
# 2. Freeze embeddings: ข้อมูลน้อย
# 3. Fine-tune: ปล่อยให้ train ต่อ (best)
- ใช้ GloVe/Word2Vec สำเร็จรูป
- Freeze: ไม่ train embeddings (ข้อมูลน้อย)
- Fine-tune: ปล่อยให้ train ต่อ (ข้อมูลเยอะ)
- Train from scratch: เริ่มใหม่ทั้งหมด
[เคล็ดลับ] Fine-tuning มักได้ผลดีที่สุดถ้าข้อมูลเพียงพอ
uv add django torch torchtext
uv run django-admin startproject wk12 .
uv run manage.py startapp dashboard
# dashboard/views.py
import json, torch
from django.http import StreamingHttpResponse
def train(request, model_type):
def event_stream():
# Build model based on selection
if model_type == 'lstm':
model = SentimentLSTM(25000, 128, 64, 1)
elif model_type == 'gru':
model = SentimentGRU(25000, 128, 64, 1)
criterion = nn.BCEWithLogitsLoss()
optimizer = torch.optim.Adam(model.parameters())
for epoch in range(10):
# Train on IMDB subset
train_loss = 0
for batch_text, batch_label in train_loader:
pred = model(batch_text).squeeze()
loss = criterion(pred, batch_label.float())
optimizer.zero_grad()
loss.backward()
optimizer.step()
train_loss += loss.item()
# Evaluate
val_acc = evaluate(model, test_loader)
yield f"data: {json.dumps({'epoch': epoch, 'loss': train_loss/len(train_loader), 'acc': val_acc})}\n\n"
return StreamingHttpResponse(event_stream(), content_type='text/event-stream')
- Django: web framework สำหรับ backend
- SSE: real-time training updates
- เลือก model: LSTM, GRU
- แสดงผล loss และ accuracy แบบ live
[โครงสร้าง] manage.py -> settings.py -> urls.py -> views.py -> templates/
from django.http import StreamingHttpResponse
import json, time
def train_stream(request, model_type):
def event_stream():
# Simplified training on small subset
model = build_model(model_type)
criterion = nn.BCEWithLogitsLoss()
optimizer = torch.optim.Adam(model.parameters())
for epoch in range(20):
model.train()
epoch_loss = 0
correct = 0
total = 0
for texts, labels in mini_loader:
optimizer.zero_grad()
pred = model(texts).squeeze()
loss = criterion(pred, labels.float())
loss.backward()
torch.nn.utils.clip_grad_norm_(
model.parameters(), max_norm=1.0)
optimizer.step()
epoch_loss += loss.item()
predicted = (pred > 0.5).long()
correct += (predicted == labels).sum().item()
total += labels.size(0)
accuracy = correct / total
avg_loss = epoch_loss / len(mini_loader)
# Send update to frontend
yield f"data: {json.dumps({\n"
f" 'epoch': {epoch},\n"
f" 'loss': {avg_loss:.4f},\n"
f" 'accuracy': {accuracy:.4f}\n"
f"})}\n\n"
time.sleep(0.1)
resp = StreamingHttpResponse(event_stream(),
content_type='text/event-stream')
resp['Cache-Control'] = 'no-cache'
return resp
- StreamingHttpResponse: ส่ง data ทีละ epoch
- วัดทั้ง loss และ accuracy
- clip_grad_norm_: ป้องกัน exploding gradients
- time.sleep(0.1): ให้ frontend อ่านทัน
[สำคัญ] SSE = one-way (server -> client), real-time updates
<!-- dashboard/templates/index.html -->
<div x-data="{ model: 'lstm', training: false }">
<h1>Sentiment Analysis Dashboard</h1>
<select x-model="model">
<option value="lstm">LSTM</option>
<option value="gru">GRU</option>
</select>
<button @click="startTrain()">
Start Training
</button>
<div>
<p>Epoch: <span x-text="epoch"></span></p>
<p>Loss: <span x-text="loss"></span></p>
<p>Accuracy: <span x-text="acc"></span></p>
</div>
<canvas id="chart" width="420" height="200"></canvas>
</div>
<script>
function startTrain() {
const es = new EventSource('/train/' + model);
es.onmessage = function(e) {
const m = JSON.parse(e.data);
epoch = m.epoch;
loss = m.loss.toFixed(4);
acc = m.accuracy.toFixed(4);
drawChart(m.epoch, m.loss, m.accuracy);
};
}
</script>
- ใช้ <select> เลือก LSTM/GRU
- แสดง epoch, loss, accuracy แบบ live
- ใช้ Canvas วาด loss/accuracy curves
- ใช้ EventSource รับ SSE updates
[เคล็ดลับ] ใช้ Alpine.js สำหรับ reactive UI
สร้าง LSTM sentiment classifier: