Hugging Face
Hugging Face is an ecosystem — transformers, datasets, and the Hub — providing pretrained models and standardized tooling on top of frameworks like PyTorch.
Hugging Face is the ecosystem layer most practitioners build on top
of a framework like PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation.: the transformers
library provides thousands of pretrained model architectures with a
consistent load/fine-tune/run API; the datasets library and the
Hugging Face Hub give standardized access to public datasets and
pretrained weights. accelerate and PEFT (including LoRA) handle
multi-GPU distribution and efficient fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.
respectively.
How it works
The Hub is a Git-backed repository host: model weights, configs, and
tokenizer files live under a repo id like org/model-name. Calling
AutoModelForCausalLM.from_pretrained("org/model-name") resolves that
id, downloads the files into a local cache, reads config.json to pick
the right architecture class, and loads the weights into it. A matching
AutoTokenizer handles tokenizationTokenizationTokenization converts raw text into a sequence of integers a model can process, typically via subword schemes like byte-pair encoding (BPE). with the
exact vocabulary the model was trained on.
The abstraction is deliberately thin — what you get back is an ordinary
nn.Module you can train, inspect, or wrap yourself. Trainer and
accelerate sit above that for training loops and
distributed trainingDistributed Training (FSDP, DeepSpeed, Megatron-LM)Distributed training frameworks split a model's parameters, gradients, and optimizer state across many GPUs so models too large for one GPU can still be trained., and PEFT
injects adapter layers so only a small fraction of parameters receive
gradient updates.
When it breaks
- Tokenizer/model drift. Loading a tokenizer from a different repo than the weights produces no error, just garbage output, because token ids silently map to the wrong embeddingsEmbeddingAn embedding is a learned vector of numbers representing a token, word, or passage of text such that similar meanings end up close together in vector space..
- Chat templates matter. Instruction-tuned models expect a specific
prompt format. Skipping
apply_chat_templateyields plausible-looking but noticeably worse generations. - Repos are mutable.
mainon a Hub repo can change under you; reproducible runs need a pinnedrevisioncommit hash. - Cache and quota surprises. The default cache grows without bound and gated or private repos fail at download time with an auth error rather than at import time.
See also: PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation., Fine-tuningFine-tuningFine-tuning continues training a pretrained model on a smaller, curated dataset to teach it a specific behavior, such as following instructions.
Learn more: Tooling & The Dev Stack · Hugging Face Hub
Mentioned in
Lessons where this comes up in context.
JAX
JAX is a Google-associated deep learning framework built around function transformations (grad, jit, vmap), popular in research and large-scale training.
Distributed Training (FSDP, DeepSpeed, Megatron-LM)
Distributed training frameworks split a model's parameters, gradients, and optimizer state across many GPUs so models too large for one GPU can still be trained.