Multimodal, Foundation, and Large Language Models
Foundation Models
Scaling pretraining
Foundation models are large pretrained models adapted to many downstream tasks. Their training often relies on massive self-supervised corpora, scaling laws, and post-training methods such as instruction tuning and alignment.
Common components
Capabilities and risks
| Capability | Benefit | Risk |
|---|---|---|
| Instruction following | Better usability | Over-reliance on prompt format |
| Retrieval augmentation | Improved factual grounding | Retriever errors propagate |
| Multimodal fusion | Richer perception | Alignment across modalities is hard |
| Tool use | Extends system capabilities | Error cascades and security issues |
Evaluation is hard
Benchmark performance may overestimate real-world reliability because of contamination, shortcut learning, and narrow test distributions.
What is one major advantage of retrieval-augmented generation?
Retrieval supplies external evidence that can improve factuality and domain coverage.
Correct answer: It can ground model outputs in external documents
Why is post-training important for foundation models?
Pretraining alone usually optimizes generic next-token prediction rather than useful downstream behavior.
Correct answer: Because it aligns a pretrained model with user instructions, task requirements, and safety constraints.
Adapting foundation models
Prompting
- No parameter updates
- Fast and cheap
- Sensitive to prompt quality
Fine-tuning
- Updates model parameters
- Can specialize behavior
- Needs curated data
Research agenda for frontier model analysis
-
1
Step 1: Measure scaling behavior across data and compute.
-
2
Step 2: Probe in-context learning and task transfer.
-
3
Step 3: Evaluate hallucination, calibration, and faithfulness.
-
4
Step 4: Study multimodal grounding and tool use.
-
5
Step 5: Analyze emergent abilities and their causal origins.