Natural Language Processing and Large Language Models
Language as Data
Core NLP Tasks
Natural language processing includes tasks such as classification, translation, summarization, question answering, and information extraction. These tasks require handling ambiguity, context, and meaning.
Tokens and Embeddings
Text is typically split into tokens, which are then mapped to vectors called embeddings. Embeddings let models work with semantic relationships in continuous space.
Transformers
Transformers use attention mechanisms to weigh relationships between tokens. Self-attention allows the model to compare each token with others in the same sequence and build contextual representations.
What is the main purpose of tokenization?
Tokenization breaks text into smaller pieces that models can process.
Correct answer: To convert text into model-readable units
What does attention help a language model do?
Attention weights interactions between tokens based on relevance.
Correct answer: Focus on the most relevant parts of the input when producing a representation or output.
Context window
A model can only condition on a limited amount of text at once, which affects memory, coherence, and long-document handling.
Classic NLP vs Transformer-Based NLP
Classic NLP
- Often uses hand-engineered features
- Can work well on smaller tasks
- Usually less powerful on complex language understanding
Transformer-based NLP
- Learns rich representations from data
- Excels at many large-scale tasks
- Requires more compute and careful training
Embeddings are best described as:
Embeddings map discrete items into continuous vectors that capture meaning and similarity.
Correct answer: Vector representations of tokens or concepts