Training and Deploying Large-Scale Models

MVA & M2 Math-Mod 2026–2027 — 2nd semester

Instructor: E. Oyallon
Format: 8 sessions
Distributed training Scaling strategies Modern toolchains

Course objective

This course introduces the foundations and practices of training modern Large Language Models (LLMs) at scale. You will learn how deep learning models are trained across multiple GPUs, nodes, and clusters—and why distributed training is essential for today’s largest AI systems.

We will cover:

  • Core techniques for distributed training
  • Modern frameworks and scaling strategies
  • Practical implementations with real-world toolchains
  • Theoretical underpinnings of large-scale learning
  • Inference and applications

As LLMs grow in complexity and impact, understanding how they are built and deployed has become essential for researchers and engineers. This series bridges engineering and theory.

Organization

  • The course is offered jointly in the MVA and M2 Math-Mod programs.
  • 8 sessions during the second semester.
  • Lectures combine conceptual material with practical labs.
  • MVA, M2 Math-Mod students and external auditors should register using this form.
  • All attendees must register in advance to participate in the course.

The detailed programme, lab material, and practical information will be announced progressively.

Lectures & Labs

Topics and lab material will be announced later.

# Topic Date Labs
1TBA14/01/27TBA
2TBA21/01/27TBA
3TBA28/01/27TBA
4TBA25/02/27TBA
5TBA04/03/27TBA
6TBA11/03/27TBA
7TBA18/03/27TBA
8TBA01/04/27TBA

Grading

The exact grading scheme and deadlines will be confirmed at the beginning of the course. The current plan is to keep a combination of homework and a small project.

HW1 — TBA HW2 — TBA Project — TBA

The project will be related to the lecture topics and may involve studying, reproducing, or extending ideas from a recent research paper. Detailed instructions and deadlines: TBA.