TensorFlow
TensorFlow is a deep learning framework, dominant in the mid-2010s, still widely used in production and edge deployment (via TensorFlow Lite).
TensorFlow (Google) was the dominant deep learningDeep LearningDeep learning is machine learning using multi-layer neural networks, which learn their own features from raw data instead of relying on hand-engineered ones. framework in the mid-2010s. It's still widely used in existing production systems and in mobile/edge deployment (via TensorFlow Lite), but has lost most of the research and new-project share to PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation. over the last several years. Keras is TensorFlow's high-level model API.
How it works
TensorFlow's original design was define-then-run: you built a static
dataflow graph of operations, then executed it inside a Session. TF 2
flipped the default to eager execution, with @tf.function opting a
Python function back into a traced graph that the runtime can optimize,
fuse, and place across devices. That graph is also the artifact that
makes deployment portable — a SavedModel bundles the graph and weights
so the same model can run under TensorFlow Serving, in a browser via
TensorFlow.js, or converted to a TFLite flatbuffer for phones and
microcontrollers, usually with quantizationQuantizationQuantization reduces a model's numerical precision (e.g. 16-bit to 4-bit) to shrink memory footprint and speed up inference, at some cost to accuracy.
applied.
Keras layers sit on top: model.fit wraps the training loop,
tf.data builds input pipelines that prefetch and shard, and gradients
come from tf.GradientTape.
When it breaks
- Tracing surprises. Code inside
@tf.functionruns once at trace time. Python counters, prints, and list appends do not behave as they read, and changing input shapes retraces the function. - Version churn. The 1.x to 2.x transition and repeated Keras API shuffles mean older tutorials and checkpoints frequently fail to run unmodified against a current install.
- Conversion gaps. TFLite supports a subset of ops, so an otherwise working model — especially one with dynamic shapes or custom layers — can fail at convert time rather than at training time.
- Ecosystem drift. New research code and pretrained weights largely target other frameworks, so reproducing a recent paper often means porting it before you can run inferenceInferenceInference is using a trained model to generate output, as opposed to training — for LLMs, an inherently sequential, token-by-token process with its own performance engineering. at all.
See also: PyTorchPyTorchPyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation., JAXJAXJAX is a Google-associated deep learning framework built around function transformations (grad, jit, vmap), popular in research and large-scale training.
Learn more: Tooling & The Dev Stack · TensorFlow official docs
Mentioned in
Lessons where this comes up in context.
PyTorch
PyTorch is the dominant deep learning framework in both research and production, providing tensor computation, GPU dispatch, and automatic differentiation.
JAX
JAX is a Google-associated deep learning framework built around function transformations (grad, jit, vmap), popular in research and large-scale training.