แผนการเรียนรู้
/

Deep Learning for Text

สัปดาห์ที่ 12

Word Embeddings และ Sequence Models

1. Text Representation: One-hot, TF-IDF

2. Word Embeddings: Word2Vec, GloVe

3. PyTorch nn.Embedding

4. Text Classification with LSTM

5. Sentiment Analysis Example

Learning Roadmap

1
Text Basics
One-hot, TF-IDF, Bag of Words
2
Embeddings
Word2Vec, GloVe, nn.Embedding
3
Sequence Models
LSTM/GRU for text classification
4
Sentiment Analysis
IMDB dataset project

ทำไม Text ถึงยาก

- Variable length: ประโยคยาวสั้นไม่เท่ากัน

- Discrete: คำเป็น categorical ไม่ใช่ตัวเลข

- Sequential: ลำดับคำสำคัญ (\"not good\" != \"good\")

- Ambiguous: คำเดียวกันมีหลายความหมาย

[ปัญหา] Machine ไม่เข้าใจคำ = ต้องแปลงคำเป็นตัวเลขก่อน

# Text: variable length, discrete
texts = [
    "I love this movie",           # 4 words
    "This is terrible",            # 3 words
    "The film was absolutely great" # 5 words
]

# Problem: NN ต้องการ fixed-size input
# Solution: Embeddings + Padding

# ลำดับสำคัญ!
"not good" -> negative
"good"     -> positive
"not bad"  -> positive (double negative)

One-Hot Encoding

# Vocabulary: {"I": 0, "love": 1, "this": 2, "movie": 3}

# One-hot vector (size = vocab_size)
"I"    -> [1, 0, 0, 0]
"love" -> [0, 1, 0, 0]
"this" -> [0, 0, 1, 0]
"movie"-> [0, 0, 0, 1]

# Problem: sparse, high-dimensional
# 50K words -> 50K-dimensional vectors!

# All vectors are equidistant
# (no semantic similarity)
import torch
vocab = {"I": 0, "love": 1, "this": 2, "movie": 3}
one_hot = torch.zeros(4, 4)
for word, idx in vocab.items():
    one_hot[idx, idx] = 1.0
print(one_hot)
# tensor([[1., 0., 0., 0.],  <- "I"
#         [0., 1., 0., 0.],  <- "love"
#         [0., 0., 1., 0.],  <- "this"
#         [0., 0., 0., 1.]]) <- "movie"

- One-hot: vector ที่มี 1 ที่ตำแหน่งของคำ

- ขนาดเท่ากับ vocabulary

- ปัญหา: sparse, ไม่มี semantic meaning

- ทุกคำอยู่ห่างกันเท่ากัน

[สำคัญ] One-hot ไม่จับความหมายของคำ = ต้องใช้ embeddings

TF-IDF

- TF (Term Frequency): คำนี้บ่อยแค่ไหนใน document

- IDF (Inverse Document Frequency): คำนี้หายากแค่ไหน

- TF-IDF = TF x IDF: คำสำคัญได้ค่าสูง

- ดีกว่า one-hot เพราะจับ importance

[เคล็ดลับ] "the", "is" ได้ IDF ต่ำ (อยู่ในทุกเอกสาร), คำเฉพาะได้ IDF สูง

from sklearn.feature_extraction.text import TfidfVectorizer

docs = [
    "I love this movie",
    "This movie is great",
    "I hate this movie"
]

vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(docs)

print(vectorizer.get_feature_names_out())
# ['great' 'hate' 'is' 'love' 'movie' 'this']

print(X.toarray())
# [[0.   0.   0.   0.82 0.58 0.0 ]  <- "love"
#  [0.62 0.   0.62 0.   0.44 0.44]  <- "great"
#  [0.   0.82 0.   0.   0.58 0.0 ]] <- "hate"

# "movie" appears in all docs -> low TF-IDF
# "love"/"hate" -> high TF-IDF (important)

Word Embeddings

- Embedding: dense vector ที่จับความหมายของคำ

- คำที่ความหมายใกล้กัน = vector ใกล้กัน

- มิติคงที่ (เช่น 100-300 dimensions)

- เรียนรู้จากข้อมูลจำนวนมาก

[ตัวอย่าง] king - man + woman = queen (semantic arithmetic)

# Embedding: dense, low-dimensional
# "love" -> [0.2, -0.5, 0.8, ..., 0.1]  (300-d)
# "like" -> [0.3, -0.4, 0.7, ..., 0.2]  (300-d)
# "hate" -> [-0.7, 0.6, -0.3, ..., -0.5] (300-d)

# Similarity via cosine
import numpy as np

love = np.array([0.2, -0.5, 0.8, 0.1])
like = np.array([0.3, -0.4, 0.7, 0.2])
hate = np.array([-0.7, 0.6, -0.3, -0.5])

def cosine_sim(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

print(f"love-like: {cosine_sim(love, like):.3f}")
print(f"love-hate: {cosine_sim(love, hate):.3f}")
# love-like: 0.987 (similar)
# love-hate: -0.952 (opposite)

Word2Vec

# Word2Vec: predict context from word
# "The cat sat on the mat"
#   -> predict "sat" from ["cat", "on"]
#   -> or predict "cat" from context ["The", "sat"]

from gensim.models import Word2Vec

sentences = [
    ["I", "love", "deep", "learning"],
    ["deep", "learning", "is", "great"],
    ["I", "love", "natural", "language", "processing"],
]

model = Word2Vec(sentences, vector_size=100,
                 window=5, min_count=1)

# Get embedding
vec = model.wv["love"]
print(vec.shape)  # (100,)

# Find similar words
print(model.wv.most_similar("love"))
# [('deep', 0.85), ('great', 0.72), ...]

# Semantic arithmetic
result = model.wv.most_similar(
    positive=['king', 'woman'],
    negative=['man'])
# [('queen', 0.92), ...]

- CBOW: predict word from context

- Skip-gram: predict context from word

- vector_size: ขนาด embedding (100-300)

- window: ขนาด context window

[สำคัญ] Word2Vec เรียนรู้จาก co-occurrence = คำที่อยู่ด้วยกันมีความหมายใกล้กัน

GloVe (Global Vectors)

- GloVe: ใช้ global co-occurrence matrix

- ต่างจาก Word2Vec ที่ใช้ local context

- Pre-trained vectors สำเร็จรูป

- ขนาด: 6B tokens, 400K words, 50/100/200/300d

[เปรียบเทียบ] Word2Vec เร็วกว่า / GloVe จับ global statistics ได้ดีกว่า

# Load pre-trained GloVe
import numpy as np

def load_glove(path):
    embeddings = {}
    with open(path, 'r', encoding='utf-8') as f:
        for line in f:
            values = line.split()
            word = values[0]
            vector = np.array(values[1:], dtype='float32')
            embeddings[word] = vector
    return embeddings

glove = load_glove('glove.6B.100d.txt')

# Similar words
def most_similar(word, top_k=5):
    vec = glove[word]
    sims = {}
    for w, v in glove.items():
        if w != word:
            sims[w] = np.dot(vec, v) / (
                np.linalg.norm(vec) * np.linalg.norm(v))
    return sorted(sims.items(), key=lambda x: -x[1])[:top_k]

print(most_similar("computer"))
# [('software', 0.85), ('pc', 0.82), ...]

nn.Embedding in PyTorch

import torch
import torch.nn as nn

# Create embedding layer
vocab_size = 10000
embed_dim = 128

embedding = nn.Embedding(vocab_size, embed_dim)
print(embedding.weight.shape)
# torch.Size([10000, 128])

# Convert words to indices
word_to_idx = {"I": 0, "love": 1, "deep": 2, "learning": 3}
sentence = ["I", "love", "deep", "learning"]
indices = torch.tensor([word_to_idx[w] for w in sentence])
print(indices)  # tensor([0, 1, 2, 3])

# Get embeddings
embeds = embedding(indices)
print(embeds.shape)  # torch.Size([4, 128])

# Use pre-trained GloVe
def from_pretrained(embeddings_matrix):
    vocab_size, dim = embeddings_matrix.shape
    layer = nn.Embedding(vocab_size, dim)
    layer.weight = nn.Parameter(
        torch.tensor(embeddings_matrix, dtype=torch.float32),
        requires_grad=False)  # Freeze weights
    return layer

- nn.Embedding(vocab_size, embed_dim): lookup table

- Input: indices (LongTensor)

- Output: dense vectors

- requires_grad=False: freeze pre-trained vectors

[สำคัญ] ต้อง padding sequences ให้ยาวเท่ากันก่อน

Text Classification Architecture

- Input: sequence of word indices

- Embedding: convert to dense vectors

- RNN/LSTM/GRU: process sequence

- FC + Softmax: classify

[สถาปัตยกรรม] Embedding -> LSTM -> FC -> Sigmoid (binary) / Softmax (multi-class)

class TextClassifier(nn.Module):
    def __init__(self, vocab_size, embed_dim,
                 hidden_dim, output_dim):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, embed_dim)
        self.lstm = nn.LSTM(embed_dim, hidden_dim,
                           batch_first=True)
        self.fc = nn.Linear(hidden_dim, output_dim)

    def forward(self, x):
        # x: (batch, seq_len) - word indices
        embeds = self.embedding(x)
        # embeds: (batch, seq_len, embed_dim)
        out, (hn, cn) = self.lstm(embeds)
        # out: (batch, seq_len, hidden_dim)
        out = self.fc(hn[-1])
        # out: (batch, output_dim)
        return torch.sigmoid(out)

model = TextClassifier(10000, 128, 64, 1)
# Binary sentiment classification

Padding และ Collate Function

from torch.nn.utils.rnn import pad_sequence
from torch.utils.data import Dataset, DataLoader

class TextDataset(Dataset):
    def __init__(self, texts, labels, vocab):
        self.texts = texts
        self.labels = labels
        self.vocab = vocab

    def __len__(self):
        return len(self.texts)

    def __getitem__(self, idx):
        tokens = self.texts[idx].split()
        indices = [self.vocab.get(t, 0) for t in tokens]
        return torch.tensor(indices), self.labels[idx]

def collate_fn(batch):
    texts, labels = zip(*batch)
    # Pad sequences to same length
    texts_padded = pad_sequence(texts, batch_first=True,
                                 padding_value=0)
    labels = torch.tensor(labels, dtype=torch.float32)
    return texts_padded, labels

dataset = TextDataset(texts, labels, vocab)
loader = DataLoader(dataset, batch_size=32,
                    collate_fn=collate_fn)

- Padding: เติม 0 ให้ sequences ยาวเท่ากัน

- pad_sequence: pad batch ให้ยาวเท่า max length

- collate_fn: กำหนดวิธีรวม samples เป็น batch

- Attention mask: บอก model ว่าตำแหน่งไหนเป็น padding

[สำคัญ] Padding position ไม่ควร effect prediction = ต้อง mask

LSTM for Sentiment Analysis

class SentimentLSTM(nn.Module):
    def __init__(self, vocab_size, embed_dim,
                 hidden_dim, output_dim):
        super().__init__()
        self.embedding = nn.Embedding(vocab_size, embed_dim)
        self.lstm = nn.LSTM(embed_dim, hidden_dim,
                           num_layers=2, batch_first=True,
                           dropout=0.5)
        self.fc = nn.Linear(hidden_dim, output_dim)
        self.dropout = nn.Dropout(0.5)

    def forward(self, x):
        embeds = self.dropout(self.embedding(x))
        out, (hn, cn) = self.lstm(embeds)
        out = self.dropout(out[:, -1, :])
        return self.fc(out)

# Training
model = SentimentLSTM(25000, 128, 64, 1)
criterion = nn.BCEWithLogitsLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)

for epoch in range(10):
    model.train()
    for batch_text, batch_label in loader:
        pred = model(batch_text).squeeze()
        loss = criterion(pred, batch_label)
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

- ใช้ Dropout ป้องกัน overfitting

- BCEWithLogitsLoss: sigmoid + BCE ในที่เดียว

- ใช้ out[:, -1, :]: hidden state ตัวสุดท้าย

- 2 layers LSTM + dropout=0.5

[เคล็ดลับ] ใช้ pre-trained embedding จะเร็วกว่า train จากศูนย์

IMDB Dataset

- IMDB: 50K movie reviews (25K train, 25K test)

- Binary sentiment: positive / negative

- ข้อความยาวไม่เท่ากัน ( averaging 231 words)

- Standard benchmark สำหรับ text classification

[ข้อมูล] ใช้ vocab 25K words, pad to 256 tokens

from torchtext.datasets import IMDB
from torchtext.data.utils import get_tokenizer
from torchtext.vocab import build_vocab_from_iterator

tokenizer = get_tokenizer('basic_english')

def yield_tokens(data_iter):
    for text, label in data_iter:
        yield tokenizer(text)

# Load data
train_iter, test_iter = IMDB(split=('train', 'test'))

# Build vocab
vocab = build_vocab_from_iterator(
    yield_tokens(train_iter),
    specials=["<unk>", "<pad>"],
    max_size=25000)
vocab.set_default_index(vocab["<unk>"])

# Text pipeline
text_pipeline = lambda x: vocab(tokenizer(x))

# Example
text = "This movie is great!"
indices = text_pipeline(text)
print(indices)  # [123, 456, 789, 1011]

Training Loop สำหรับ Text

def train_epoch(model, loader, optimizer, criterion):
    model.train()
    total_loss = 0
    for text, label in loader:
        optimizer.zero_grad()
        pred = model(text).squeeze()
        loss = criterion(pred, label.float())
        loss.backward()
        torch.nn.utils.clip_grad_norm_(
            model.parameters(), max_norm=1.0)
        optimizer.step()
        total_loss += loss.item()
    return total_loss / len(loader)

def evaluate(model, loader, criterion):
    model.eval()
    total_loss = 0
    correct = 0
    with torch.no_grad():
        for text, label in loader:
            pred = model(text).squeeze()
            loss = criterion(pred, label.float())
            total_loss += loss.item()
            predicted = (pred > 0.5).long()
            correct += (predicted == label).sum().item()
    accuracy = correct / len(loader.dataset)
    return total_loss / len(loader), accuracy

# Run
for epoch in range(5):
    loss = train_epoch(model, train_loader, optimizer, criterion)
    val_loss, acc = evaluate(model, test_loader, criterion)
    print(f"Epoch {epoch}: loss={loss:.4f} acc={acc:.4f}")

- clip_grad_norm_: ป้องกัน exploding gradients

- วัดทั้ง loss และ accuracy

- ใช้ model.eval() ตอน evaluate

- 5-10 epochs พอสำหรับ IMDB

[สำคัญ] Gradient clipping สำคัญมากสำหรับ RNN/LSTM

Evaluation และ Visualization

import matplotlib.pyplot as plt
from sklearn.metrics import classification_report

# Predict
model.eval()
all_preds = []
all_labels = []
with torch.no_grad():
    for text, label in test_loader:
        pred = model(text).squeeze()
        predicted = (pred > 0.5).long()
        all_preds.extend(predicted.tolist())
        all_labels.extend(label.tolist())

# Classification report
print(classification_report(all_labels, all_preds,
                           target_names=['Negative', 'Positive']))

# Plot loss curves
plt.figure(figsize=(10, 4))
plt.subplot(1, 2, 1)
plt.plot(train_losses, label='Train')
plt.plot(val_losses, label='Val')
plt.legend()
plt.title('Loss')

plt.subplot(1, 2, 2)
plt.plot(train_accs, label='Train')
plt.plot(val_accs, label='Val')
plt.legend()
plt.title('Accuracy')
plt.show()

# Predict custom text
def predict_sentiment(text):
    model.eval()
    indices = text_pipeline(text)
    tensor = torch.tensor([indices])
    with torch.no_grad():
        pred = torch.sigmoid(model(tensor)).item()
    return "Positive" if pred > 0.5 else "Negative"

print(predict_sentiment("This movie is amazing!"))
# Positive

- ใช้ classification_report ดู precision, recall, f1

- Plot loss/accuracy curves

- สร้าง predict_sentiment function

- ต้อง preprocessing เหมือนตอน train

[เมตริก] Accuracy 85-90% สำหรับ IMDB = ดีแล้ว

Pre-trained Embeddings

import numpy as np

def load_glove(path, vocab, embed_dim):
    # Create embedding matrix
    matrix = np.zeros((len(vocab), embed_dim))
    with open(path, 'r', encoding='utf-8') as f:
        for line in f:
            word, vector = line.split(maxsplit=1)
            if word in vocab:
                idx = vocab[word]
                matrix[idx] = np.fromstring(vector, sep=' ')
    return matrix

# Load GloVe
glove_matrix = load_glove('glove.6B.100d.txt', vocab, 100)

# Set embedding weights
model.embedding.weight = nn.Parameter(
    torch.tensor(glove_matrix, dtype=torch.float32),
    requires_grad=True)  # Fine-tune

# Or freeze
model.embedding.weight = nn.Parameter(
    torch.tensor(glove_matrix, dtype=torch.float32),
    requires_grad=False)  # ไม่ train ใหม่

# Fine-tuning strategies:
# 1. Train from scratch: ต้องมีข้อมูลเยอะ
# 2. Freeze embeddings: ข้อมูลน้อย
# 3. Fine-tune: ปล่อยให้ train ต่อ (best)

- ใช้ GloVe/Word2Vec สำเร็จรูป

- Freeze: ไม่ train embeddings (ข้อมูลน้อย)

- Fine-tune: ปล่อยให้ train ต่อ (ข้อมูลเยอะ)

- Train from scratch: เริ่มใหม่ทั้งหมด

[เคล็ดลับ] Fine-tuning มักได้ผลดีที่สุดถ้าข้อมูลเพียงพอ

Django Dashboard + SSE

uv add django torch torchtext

uv run django-admin startproject wk12 .
uv run manage.py startapp dashboard

# dashboard/views.py
import json, torch
from django.http import StreamingHttpResponse

def train(request, model_type):
    def event_stream():
        # Build model based on selection
        if model_type == 'lstm':
            model = SentimentLSTM(25000, 128, 64, 1)
        elif model_type == 'gru':
            model = SentimentGRU(25000, 128, 64, 1)

        criterion = nn.BCEWithLogitsLoss()
        optimizer = torch.optim.Adam(model.parameters())

        for epoch in range(10):
            # Train on IMDB subset
            train_loss = 0
            for batch_text, batch_label in train_loader:
                pred = model(batch_text).squeeze()
                loss = criterion(pred, batch_label.float())
                optimizer.zero_grad()
                loss.backward()
                optimizer.step()
                train_loss += loss.item()

            # Evaluate
            val_acc = evaluate(model, test_loader)

            yield f"data: {json.dumps({'epoch': epoch, 'loss': train_loss/len(train_loader), 'acc': val_acc})}\n\n"

    return StreamingHttpResponse(event_stream(), content_type='text/event-stream')

- Django: web framework สำหรับ backend

- SSE: real-time training updates

- เลือก model: LSTM, GRU

- แสดงผล loss และ accuracy แบบ live

[โครงสร้าง] manage.py -> settings.py -> urls.py -> views.py -> templates/

SSE Training Stream

from django.http import StreamingHttpResponse
import json, time

def train_stream(request, model_type):
    def event_stream():
        # Simplified training on small subset
        model = build_model(model_type)
        criterion = nn.BCEWithLogitsLoss()
        optimizer = torch.optim.Adam(model.parameters())

        for epoch in range(20):
            model.train()
            epoch_loss = 0
            correct = 0
            total = 0

            for texts, labels in mini_loader:
                optimizer.zero_grad()
                pred = model(texts).squeeze()
                loss = criterion(pred, labels.float())
                loss.backward()
                torch.nn.utils.clip_grad_norm_(
                    model.parameters(), max_norm=1.0)
                optimizer.step()

                epoch_loss += loss.item()
                predicted = (pred > 0.5).long()
                correct += (predicted == labels).sum().item()
                total += labels.size(0)

            accuracy = correct / total
            avg_loss = epoch_loss / len(mini_loader)

            # Send update to frontend
            yield f"data: {json.dumps({\n"
                  f"  'epoch': {epoch},\n"
                  f"  'loss': {avg_loss:.4f},\n"
                  f"  'accuracy': {accuracy:.4f}\n"
                  f"})}\n\n"
            time.sleep(0.1)

    resp = StreamingHttpResponse(event_stream(),
        content_type='text/event-stream')
    resp['Cache-Control'] = 'no-cache'
    return resp

- StreamingHttpResponse: ส่ง data ทีละ epoch

- วัดทั้ง loss และ accuracy

- clip_grad_norm_: ป้องกัน exploding gradients

- time.sleep(0.1): ให้ frontend อ่านทัน

[สำคัญ] SSE = one-way (server -> client), real-time updates

Frontend: Text Classification Dashboard

<!-- dashboard/templates/index.html -->
<div x-data="{ model: 'lstm', training: false }">
  <h1>Sentiment Analysis Dashboard</h1>

  <select x-model="model">
    <option value="lstm">LSTM</option>
    <option value="gru">GRU</option>
  </select>

  <button @click="startTrain()">
    Start Training
  </button>

  <div>
    <p>Epoch: <span x-text="epoch"></span></p>
    <p>Loss: <span x-text="loss"></span></p>
    <p>Accuracy: <span x-text="acc"></span></p>
  </div>

  <canvas id="chart" width="420" height="200"></canvas>
</div>

<script>
function startTrain() {
  const es = new EventSource('/train/' + model);
  es.onmessage = function(e) {
    const m = JSON.parse(e.data);
    epoch = m.epoch;
    loss = m.loss.toFixed(4);
    acc = m.accuracy.toFixed(4);
    drawChart(m.epoch, m.loss, m.accuracy);
  };
}
</script>

- ใช้ <select> เลือก LSTM/GRU

- แสดง epoch, loss, accuracy แบบ live

- ใช้ Canvas วาด loss/accuracy curves

- ใช้ EventSource รับ SSE updates

[เคล็ดลับ] ใช้ Alpine.js สำหรับ reactive UI

Part 5: เปรียบเทียบ Approaches

One-Hot / TF-IDF

  • - Simple, fast
  • - No semantic meaning
  • - Sparse vectors
  • - Good baseline

Word Embeddings

  • - Dense vectors
  • - Semantic similarity
  • - Pre-trained available
  • - Better than one-hot

LSTM/GRU + Embedding

  • - Capture sequence
  • - Best for classification
  • - Handle variable length
  • - State-of-the-art baseline

Summary + Homework

Key Takeaways

  • - Text ต้องแปลงเป็นตัวเลขก่อนใช้กับ NN
  • - One-hot/TF-IDF ง่ายแต่ไม่มี semantic
  • - Word embeddings จับความหมายของคำ
  • - Word2Vec, GloVe: pre-trained vectors
  • - LSTM/GRU จับ sequence ได้ดี
  • - Padding + collate_fn สำหรับ variable length

Homework

สร้าง LSTM sentiment classifier:

  • 1. ใช้ IMDB dataset (torchtext)
  • 2. ใช้ pre-trained GloVe embeddings
  • 3. LSTM(128, 2 layers) + Dropout(0.5)
  • 4. Train 5 epochs, Adam lr=1e-3
  • 5. แสดง accuracy > 85%
  • 6. สร้าง predict_sentiment() function
  • 7. บันทึก model ด้วย state_dict()