Machine Learning (ML)
Machine learning is the practice of writing programs that learn a function from data, via a loss function and an optimizer, instead of following hand-written rules.
Machine learning (ML) is the practice of building systems that learn a function from examples, rather than following hand-written rules. The core recipe: a model (a function with adjustable parameters), a loss functionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training. (a number measuring how wrong its predictions are), and an optimizer (usually gradient descentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient.) that adjusts the parameters to reduce that loss. Deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. is the subfield of ML built on neural networksNeural NetworkA neural network is layers of simple weighted-sum-plus-nonlinearity units (neurons) chained together, trained by gradient descent and backpropagation..
How it works
The standard workflow is the same whether the model is a linear regression or a transformerTransformerThe transformer is the neural network architecture built around self-attention, introduced in 2017, underlying essentially all modern LLMs.:
- Split the data into train, validation, and test sets — see train/validation/test splitTrain/Validation/Test SplitSplitting data into training, validation, and test sets lets you measure a model's ability to generalize to new data instead of just memorizing what it trained on..
- Choose a model family and initialize its parameters.
- Loop: predict on a batch, score with the loss, compute gradients, update the parameters.
- Check validation error to pick hyperparameters and decide when to stop.
- Report on the test set once.
The main learning settings differ in what signal is available. Supervised learning has input–label pairs; unsupervised learning has inputs only and looks for structure; self-supervised learning manufactures labels from the input itself (predict the next token); reinforcement learning receives a reward after acting.
When it breaks
- Data leakage. Any signal that reaches training but will not exist at prediction time — a target-derived feature, duplicated rows across splits, normalizing before splitting — produces validation scores that evaporate in production.
- Non-random splits. Time series must be split chronologically; grouped data (multiple rows per user) must be split by group, or the model memorizes the group.
- Optimizing the wrong objective. The loss you can differentiate is rarely the outcome the business cares about.
- Stale models. Input distributions drift, and a model that is never retrained degrades quietly with no error signal.
See also: Loss FunctionLoss FunctionA loss function is a single number measuring how wrong a model's predictions are, which gradient descent minimizes during training., Gradient DescentGradient DescentGradient descent is the optimization algorithm that trains models by repeatedly stepping parameters in the opposite direction of the loss function's gradient., OverfittingOverfittingOverfitting is when a model fits its training data (including noise) so closely that it fails to generalize to new, unseen data.
Learn more: ML Fundamentals · Wikipedia: Machine learning
Mentioned in
Lessons where this comes up in context.
- AI SecurityAdversarial misuse of a deployed AI system — prompt injection, jailbreaks, data exfiltration via tool use, adversarial examples, and the red-teaming practice that hunts for all of them
- Computer VisionConvolutions, pooling, CNNs, transfer learning
- Evaluation & BenchmarksHow models are actually scored — benchmark suites, leaderboards, eval methodology, and why a high benchmark score doesn't guarantee good real-world performance
- Generative ModelsGANs, VAEs, and diffusion models — how AI generates new images, audio, and video, as opposed to classifying or understanding existing content
- History & LandscapeSymbolic AI to expert systems to statistical ML to deep learning to the LLM era
- LLMsTokenization, embeddings, pretraining vs fine-tuning, RLHF basics
- ML FundamentalsSupervised/unsupervised learning, loss functions, gradient descent
- Neural Networks & BackpropFrom Karpathy's micrograd approach — building a tiny neural net and stepping through forward/backward passes
- Optimization & Training DynamicsWhy plain gradient descent isn't what actually trains modern models — Adam's mechanics, why warmup exists, how learning-rate schedules are shaped, and what batch size actually trades off at scale
- Probability & Statistics FoundationsDistributions, Bayes' theorem, and maximum likelihood estimation — the math that loss functions and uncertainty in ML are actually built on
- Reinforcement LearningMDPs, reward, policy and value functions, Q-learning — the third major ML paradigm, and the actual mechanism behind RLHF
- Tooling & The Dev StackLanguages, frameworks, and where they fit — what you'd actually touch to build and ship a model
AI (Artificial Intelligence)
AI is the field of building systems that perform tasks normally requiring human intelligence — reasoning, perception, language, and decision-making.
Deep Learning
Deep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones.