This is the anticipated schedule for the course.
The schedule is subject to change.
Each lecture lists its assigned readings and, where applicable, optional readings.
Each student will present one paper of their choosing.
Students who sign up for 04/06 will receive extra credit.
More papers will be announced if needed.
| # |
Lecture |
Date |
Slides |
Presenters |
| 1 |
Introduction and Dark Silicon
Assigned readings:
|
03/30/2026 |
Intro and Dark Silicon |
Hadi |
| 2 |
Scaling Laws
Assigned readings:
- Kaplan et al., "Scaling Laws for Neural Language Models", 2020
- Hoffmann et al., "Training Compute-Optimal Large Language Models" (Chinchilla), NeurIPS 2022
- Goodfellow, Bengio & Courville, "Deep Learning" Chapter 8: Optimization for Training Deep Models, 2016
- He et al., "Deep Residual Learning for Image Recognition" (ResNet), CVPR 2016
- Kingma & Ba, "Adam: A Method for Stochastic Optimization", ICLR 2015
|
04/01/2026 |
Scaling Laws |
Hanyang |
| 3 |
Learning Paradigms + Alignment
Assigned readings:
- Esmaeilzadeh et al., "Dark Silicon and the End of Multicore Scaling", ISCA 2011 (25-Year Retrospective) (no critique required)
- Goodfellow, Bengio & Courville, "Deep Learning" Chapters 5 + 15: Supervised/Unsupervised Learning, 2016
- Radford et al., "Improving Language Understanding by Generative Pre-Training" (GPT-1), 2018
- Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (InstructGPT/RLHF), NeurIPS 2022
- DeepSeek-AI, "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning", 2025
Optional readings:
|
04/06/2026 |
GPT-1, InstructGPT/RLHF, DeepSeek-R1 |
Ethan Jenkins, Yuvanand Saravanan, Hantian Lin |
| — |
Learning Paradigms + Alignment (continued) |
04/08/2026 |
|
|
| — |
Gradient Descent / Backpropagation
Assigned readings:
|
04/13/2026 |
|
Hadi |
| 4 |
Attention, Transformers + Diffusion
Assigned readings:
Optional readings:
|
04/15/2026 |
Attention, DiT |
Hadi, Hanyang |
| — |
Attention, Transformers + Diffusion (continued) |
04/20/2026 |
Attention |
Hadi, Hanyang |
| 5 |
Model Capabilities: MoE, Multimodal, Reasoning
Assigned readings:
|
04/22/2026 |
LLaVA, Chain-of-Thought |
Tejus Singh, Sankalpa Hota |
| 6 |
Model Capabilities (continued) + Agentic Systems
Assigned readings:
|
04/27/2026 |
Agents, Coconut, MemGPT |
Hanyang, Wentao Chen, Chi-Han Chiu |
| 7 |
Distributed Training
Assigned readings:
|
04/29/2026 |
Megatron-LM, Llama 3 Training |
Aditya Kamat, Nikhita Neelakanta |
| 8 |
Efficient Inference Architectures
Assigned readings:
Optional readings:
|
05/04/2026 |
GQA |
Sanskriti Goyal |
| 9 |
Model Compression: Quantization, LoRA, Sparsity
Assigned readings:
|
05/06/2026 |
GPTQ, LoRA, Sparse DNNs |
Shivam Yadav, Hanyang, Mohamed Ibrahim |
| — |
Model Compression: Quantization, LoRA, Sparsity (continued)
Optional readings:
|
05/11/2026 |
|
|
| 10 |
GPU Architecture + FlashAttention / Kernel Optimizations
Assigned readings:
Optional readings:
|
05/13/2026 |
Hopper GPU, FlashAttention |
Amaan M., Nisha |
| 11 |
Domain-Specific Accelerators: TPUs, Groq, PIM
Assigned readings:
|
05/18/2026 |
TPU |
John |
| — |
Domain-Specific Accelerators: TPUs, Groq, PIM (continued) |
05/20/2026 |
Groq |
Shoumik, TBD |
| — |
NO CLASS — Memorial Day |
05/25/2026 |
|
|
| 12 |
Compiler Optimizations
Assigned readings:
Optional readings:
|
05/27/2026 |
TVM, MLIR |
Jiahui Cheng, Abhinandan Sharma |
| 13 |
KV Cache Management + Prefill-Decode Scheduling
Assigned readings:
|
06/01/2026 |
Mooncake |
Shan Ananth |
| 14 |
Network Infrastructure + DPUs
Assigned readings:
Optional readings:
|
06/03/2026 |
ShadowServe |
TBD |