Unit 07: SVM, Cross-Validation, and Ensembles
The last classical-ML push before deep learning. This unit adds the professional-practice layer: how to evaluate a model honestly, how to tune hyperparameters without overfitting to your test set, and how to combine many weak models into a stronger one. It also introduces the first learned text representations — word embeddings — as a bridge toward the sequence models coming later.
Concepts You’ll Learn About
- Arithmetic coding — optimal prefix-free codes; connecting Shannon entropy to practical compression
- Support Vector Machines — maximum-margin classifiers; the kernel trick; soft-margin SVMs; C as a regularization parameter
- Cross-validation — k-fold CV; why a held-out test set is not enough for hyperparameter search
- Grid search — systematic hyperparameter sweep; combining with CV to avoid data leakage
- Word embeddings — learned dense representations; TF-IDF vs. embeddings; first exposure to representations that encode meaning
- Ensemble methods — bagging, random forests, gradient boosting; why combining models reduces variance
Topics
-
Arithmetic codes — lecture on optimal coding; connects directly to the Shannon entropy from Unit 06.
-
SVM theory and lab — the margin, support vectors, and the kernel trick; then a lab applying SVM to a classification problem. View Download Run View Download Run
-
Cross-validation and grid search — applied to Twitter airline sentiment (twitter_training.csv) with SVM + TF-IDF; includes the
mnist.pk.gzdataset as a secondary target. View Download Run -
Word embeddings and fake news — moving from bag-of-words to dense representations; applying them to a fake-news classification task. word2vec embeddings View Download Run View Download Run
-
Ensemble methods — bagging, random forests, gradient boosting, and stacking; applied to a dataset of your choice. View Download Run Assignment: apply ensemble methods to a previously-analyzed dataset and submit.
-
AET Challenge Day — end-of-semester competition event.
What’s next
Unit 08 is a short focused unit on anomaly detection — the second-quarter capstone, closing with an open-ended quarter project.