पाठ 15 / 25
Naive Bayes Classification
Bayes' rule with a simplifying assumption.
Assume features are independent given the class
Naive Bayes classifies by computing, for each class, the prior times the probability of each observed feature given that class, assuming features are independent given the class (rarely true, but often good enough). For text, features are words; smoothing (adding one to counts) prevents a single unseen word from zeroing a class. Working in log probabilities avoids numerical underflow. It is fast, needs little data and was the classic spam filter; its weakness is that it cannot capture word order or interactions.
A six-message spam filter, run
I ran this with plain Python 3 (standard library only), with fixed random seeds where randomness is used. Trained on six tiny messages, the filter labels free money as spam and report tomorrow as ham. claim your prize at noon comes out ham because at and noon appear in ham messages and your is unseen, showing how sensitive tiny models are to their training data.
import math
from collections import Counter
train = [("win free money now", "spam"), ("free prize claim now", "spam"), ("cheap money offer", "spam"),
("meeting at noon tomorrow", "ham"), ("project report attached", "ham"), ("lunch tomorrow at noon", "ham")]
classes = {"spam", "ham"}
words = {c: Counter(w for t, l in train if l == c for w in t.split()) for c in classes}
vocab = {w for t, _ in train for w in t.split()}
prior = {c: sum(l == c for _, l in train) / len(train) for c in classes}
def score(text, c): # log P(c) + sum log P(word | c), with add-one smoothing
total = sum(words[c].values())
return math.log(prior[c]) + sum(math.log((words[c][w] + 1) / (total + len(vocab))) for w in text.split())
for text in ["free money", "report tomorrow", "claim your prize at noon"]:
s = {c: score(text, c) for c in classes}
print(f"{text!r:<28} -> {max(s, key=s.get)} (log-score spam {s['spam']:.2f}, ham {s['ham']:.2f})")
Output:
'free money' -> spam (log-score spam -5.09, ham -7.28) 'report tomorrow' -> ham (log-score spam -7.28, ham -5.49) 'claim your prize at noon' -> ham (log-score spam -15.79, ham -14.98)
Use log probabilities
Multiplying many small probabilities underflows to zero; add logarithms instead.
त्वरित जाँच: What is the "naive" assumption in naive Bayes?
- Features are independent of each other given the class
- All classes are equally likely
- There is only one feature
- Words never repeat
Answer
Features are independent of each other given the class — A strong but useful simplification.