Three phases, ten weeks.
Each phase pairs derivation-heavy lectures with a project on a real dataset. Your peer pod grades assignments first, then they're reviewed in office hours.
Phase 01Weeks 1–3
Deterministic Learning & Linear Foundations
- Geometry of high-dimensional spaces, hyperplanes
- SVD and PCA, derived from scratch
- Ridge, Lasso, and dual representations, including the kernel trick
Phase 02Weeks 4–6
Probabilistic Machine Learning
- Maximum likelihood vs. maximum a posteriori
- Gaussian processes and exponential families
- Expectation–maximization and latent variable models
Phase 03Weeks 7–10
Generative Modeling & Research Practice
- Variational inference and the ELBO
- A short paper-reproduction sprint in your pod
- Capstone project on a real dataset, presented to the cohort
How it runs.
Live classes
Two a week, worked through at the board. Derivations first, code second.
Peer pods
Pods of 5–6 on Discord handle doubt discussion between classes and review each other's assignments first.
Real coursework
10 assignments, each mapped to a phase topic, and each one completed on time earns back ₹200 of your fee. AI tools can help you understand faster; they don't write your code.
Cohort 2: the current AI landscape.
Once the foundations are solid, Cohort 2 goes deep on how today's frontier systems work, starting from computer vision and building up to the architectures behind current research.
10 weeks2 live classes / week
20 seatssame pod structure
Cohort 1required prerequisite
TBAstarts after Cohort 1 wraps
Unit 01Weeks 1–2
Foundations of Computer Vision
- Convolutions, receptive fields, and pooling from first principles
- Classic architectures: ResNets, feature pyramids, batch norm
- Classification, detection, and segmentation as problem setups
Unit 02Weeks 3–4
Transformer Architectures
- Self-attention and multi-head attention, derived from scratch
- Positional encodings, layer norm placement, residual streams
- Encoder-decoder vs. decoder-only designs, and scaling laws
Unit 03Weeks 5–6
Vision-Language Models
- Contrastive image-text pretraining (CLIP-style objectives)
- Cross-modal attention and joint embedding spaces
- Instruction-tuned VLMs and visual question answering
Unit 04Weeks 7–8
Omni-Modal VLMs
- Unified tokenization across text, image, audio, and video
- Multimodal fusion strategies and shared latent spaces
- Evaluating omni-modal systems across benchmarks
Unit 05Weeks 9–10
JEPA and World Models
- Joint Embedding Predictive Architectures: predicting in latent space, not pixel space
- Self-supervised video prediction and representation learning
- World models for planning, and a capstone project on a real system
Ready for week one?
Every applicant starts with a 30-minute placement check.