microgpt

microgpt

安德烈·卡帕西 Andrej Karpathy · Eureka Labs · 2026-02-12 · Karpathy Blog ↗

打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→

摘要 · Abstract

这是关于我的新艺术项目 microgpt 的简要指南,一个 200 行纯 Python 的单文件,无依赖,可训练和推理 GPT。该文件包含了所需的所有算法内容:文档数据集、分词器、自动求导引擎、类似 GPT-2 的神经网络架构、Adam 优化器、训练循环和推理循环。其他一切都只是效率问题。我无法再简化它了。这个脚本是多个项目(micrograd、makemore、nanogpt 等)以及十年来将 LLM 简化至其最本质要素的执念的结晶,我认为它很美🥹。它甚至完美地分布在 3 列中:* 这个 GitHub gist 包含完整源代码:microgpt.py

Andrej Karpathy blog[](https://karpathy.github.io/2026/02/12/microgpt/#) This is a brief guide to my new art project microgpt, a single file of 200 lines of pure Python with no dependencies that trains and inferences a GPT. This file contains the full algorithmic content of what is needed: dataset of documents, tokenizer, autograd engine, a GPT-2-like neural network architecture, the Adam optimizer, training loop, and inference loop. Everything else is just efficiency. I cannot simplify this any further. This script is the culmination of multiple projects (micrograd, makemore, nanogpt, etc.) and a decade-long obsession to simplify LLMs to their bare essentials, and I think it is beautiful 🥹. It even breaks perfectly across 3 columns: * This GitHub gist has the full source code: microgpt.py

核心贡献 · Key contributions

局限 · Limitations

论文章节 · Sections(共 21)

全文 · Full text(逐段中英对照)

概述 Overview

Andrej Karpathy 博客[](https://karpathy.github.io/2026/02/12/microgpt/#)

Andrej Karpathy blog[](https://karpathy.github.io/2026/02/12/microgpt/#)

这是我新艺术项目 microgpt 的简要指南,它是一个 200 行纯 Python 的单一文件,无依赖,可训练和推理一个 GPT。该文件包含了所需的所有算法内容:文档数据集、分词器、自动求导引擎、类似 GPT-2 的神经网络架构、Adam 优化器、训练循环和推理循环。其他一切都只是效率问题。我无法再简化它了。这个脚本是多个项目(micrograd、makemore、nanogpt 等)以及长达十年致力于将 LLM 简化到极致的心血的结晶,我觉得它很美 🥹。它甚至完美地跨三列展示:

This is a brief guide to my new art project microgpt, a single file of 200 lines of pure Python with no dependencies that trains and inferences a GPT. This file contains the full algorithmic content of what is needed: dataset of documents, tokenizer, autograd engine, a GPT-2-like neural network architecture, the Adam optimizer, training loop, and inference loop. Everything else is just efficiency. I cannot simplify this any further. This script is the culmination of multiple projects (micrograd, makemore, nanogpt, etc.) and a decade-long obsession to simplify LLMs to their bare essentials, and I think it is beautiful 🥹. It even breaks perfectly across 3 columns:

* 这个 GitHub gist 包含完整源代码:microgpt.py

* This GitHub gist has the full source code: microgpt.py

* 也可在此网页获取:https://karpathy.ai/microgpt.html

* It’s also available on this web page: https://karpathy.ai/microgpt.html

* 也可作为 Google Colab 笔记本使用

* Also available as a Google Colab notebook

* 新消息:在我的艺术商店 karpathy.art 购买 microgpt 三联画 :)

* NEW: buy microgpt as a triptych on my art store at karpathy.art :)

以下是我引导感兴趣的读者逐步阅读代码的指南。

The following is my guide on stepping an interested reader through the code.

数据集 Dataset

大型语言模型的燃料是文本数据流,可选地分割成一组文档。在生产级应用中,每个文档是一个互联网网页,但对于 microgpt,我们使用一个更简单的示例:32,000 个名字,每行一个。

The fuel of large language models is a stream of text data, optionally separated into a set of documents. In production-grade applications, each document would be an internet web page but for microgpt we use a simpler example of 32,000 names, one per line:

数据集与任务设定 Let there be an input dataset docs: liststr of documents (e.g. a dataset of names)

names_url = 'https://raw.githubusercontent.com/karpathy/makemore/refs/heads/master/names.txt'

names_url = 'https://raw.githubusercontent.com/karpathy/makemore/refs/heads/master/names.txt'

urllib.request.urlretrieve(names_url, 'input.txt')

urllib.request.urlretrieve(names_url, 'input.txt')

docs = [l.strip() for l in open('input.txt').read().strip().split('\n') if l.strip()] # 文档列表(字符串列表)

docs = [l.strip() for l in open('input.txt').read().strip().split('\n') if l.strip()] # list[str] of documents

数据集如下所示。每个名字是一个文档:

The dataset looks like this. Each name is a document:

模型的目标是学习数据中的模式,然后生成共享这些统计模式的新文档。作为预览,在脚本结束时,我们的模型将生成(“幻觉”出!)新的、听起来合理的名字。提前看一下,我们会得到:

The goal of the model is to learn the patterns in the data and then generate similar new documents that share the statistical patterns within. As a preview, by the end of the script our model will generate (“hallucinate”!) new, plausible-sounding names. Skipping ahead, we’ll get:

这看起来没什么,但从 ChatGPT 这类模型的角度看,你与它的对话只是一个有趣的“文档”。当你用提示初始化文档时,模型从自身视角给出的响应只是统计意义上的文档补全。

It doesn’t look like much, but from the perspective of a model like ChatGPT, your conversation with it is just a funny looking “document”. When you initialize the document with your prompt, the model’s response from its perspective is just a statistical document completion.

分词器 Tokenizer

在底层,神经网络处理的是数字而非字符,因此我们需要一种方法将文本转换为整数词元 ID 序列,并能反向转换。像 tiktoken(GPT-4 使用)这样的生产级分词器为了效率而操作字符块,但最简单的分词器只需为数据集中的每个唯一字符分配一个整数:

Under the hood, neural networks work with numbers, not characters, so we need a way to convert text into a sequence of integer token ids and back. Production tokenizers like tiktoken (used by GPT-4) operate on chunks of characters for efficiency, but the simplest possible tokenizer just assigns one integer to each unique character in the dataset:

让分词器将字符串翻译为离散符号并返回 Let there be a Tokenizer to translate strings to discrete symbols and back

uchars = sorted(set(''.join(docs))) # 数据集中的唯一字符成为 token id 0..n-1

uchars = sorted(set(''.join(docs))) # unique characters in the dataset become token ids 0..n-1

BOS = len(uchars) # 特殊序列开始(BOS)标记的 token id

BOS = len(uchars) # token id for the special Beginning of Sequence (BOS) token

vocab_size = len(uchars) + 1 # 总唯一 token 数,+1 用于 BOS

vocab_size = len(uchars) + 1 # total number of unique tokens, +1 is for BOS

在上面的代码中,我们收集了数据集中所有唯一字符(即所有小写字母 a-z),排序后,每个字母通过其索引获得一个 id。注意,整数值本身没有任何意义;每个 token 只是一个独立的离散符号。它们可以是不同的表情符号,而不是 0、1、2。此外,我们创建了一个名为

In the code above, we collect all unique characters across the dataset (which are just all the lowercase letters a-z), sort them, and each letter gets an id by its index. Note that the integer values themselves have no meaning at all; each token is just a separate discrete symbol. Instead of 0, 1, 2 they might as well be different emoji. In addition, we create one more special token called

(序列开始)的特殊 token,它充当分隔符:告诉模型“新文档在此开始/结束”。在后续训练中,每个文档都会包裹上

(Beginning of Sequence), which acts as a delimiter: it tells the model “a new document starts/ends here”. Later during training, each document gets wrapped with

结束。因此,我们最终有 27 个词汇(26 个可能的小写字符 a-z 和 +1 个 BOS token)。

ends it. Therefore, we have a final vocavulary of 27 (26 possible lowercase characters a-z and +1 for the BOS token).

自动微分 Autograd

训练神经网络需要梯度:对于模型中的每个参数,我们需要知道“如果我将这个数字稍微调高一点,损失是上升还是下降,变化多少?”。计算图有许多输入(模型参数和输入词元),但最终汇聚到一个标量输出:损失(我们将在下面准确定义损失)。反向传播从该单个输出开始,沿图向后工作,计算损失相对于每个输入的梯度。它依赖于微积分中的链式法则。在生产中,像 PyTorch 这样的库会自动处理这一点。在这里,我们从一个名为

Training a neural network requires gradients: for each parameter in the model, we need to know “if I nudge this number up a little, does the loss go up or down, and by how much?”. The computation graph has many inputs (the model parameters and the input tokens) but funnels down to a single scalar output: the loss (we’ll define exactly what the loss is below). Backpropagation starts at that single output and works backwards through the graph, computing the gradient of the loss with respect to every input. It relies on the chain rule from calculus. In production, libraries like PyTorch handle this automatically. Here, we implement it from scratch in a single class called

slots = ('data', 'grad', '_children', '_local_grads')

slots = ('data', 'grad', '_children', '_local_grads')

def init(self, data, children=(), local_grads=()):

def init(self, data, children=(), local_grads=()):

self.data = data # 前向传播期间计算的该节点的标量值

self.data = data # scalar value of this node calculated during forward pass

self.grad = 0 # 损失相对于该节点的导数,在反向传播中计算

self.grad = 0 # derivative of the loss w.r.t. this node, calculated in backward pass

self._children = children # 该节点在计算图中的子节点

self._children = children # children of this node in the computation graph

self._local_grads = local_grads # 该节点相对于其子节点的局部导数

self._local_grads = local_grads # local derivative of this node w.r.t. its children

other = other if isinstance(other, Value) else Value(other)

other = other if isinstance(other, Value) else Value(other)

return Value(self.data + other.data, (self, other), (1, 1))

return Value(self.data + other.data, (self, other), (1, 1))

other = other if isinstance(other, Value) else Value(other)

other = other if isinstance(other, Value) else Value(other)

return Value(self.data * other.data, (self, other), (other.data, self.data))

return Value(self.data * other.data, (self, other), (other.data, self.data))

def pow(self, other): return Value(self.data**other, (self,), (other * self.data**(other-1),))

def pow(self, other): return Value(self.dataother, (self,), (other * self.data(other-1),))

def log(self): return Value(math.log(self.data), (self,), (1/self.data,))

def log(self): return Value(math.log(self.data), (self,), (1/self.data,))

def exp(self): return Value(math.exp(self.data), (self,), (math.exp(self.data),))

def exp(self): return Value(math.exp(self.data), (self,), (math.exp(self.data),))

def relu(self): return Value(max(0, self.data), (self,), (float(self.data > 0),))

def relu(self): return Value(max(0, self.data), (self,), (float(self.data > 0),))

def __radd__(self, other): return self + other

def radd(self, other): return self + other

def __sub__(self, other): return self + (-other)

def sub(self, other): return self + (-other)

def __rsub__(self, other): return other + (-self)

def rsub(self, other): return other + (-self)

def __rmul__(self, other): return self * other

def rmul(self, other): return self * other

def __truediv__(self, other): return self * other**-1

def truediv(self, other): return self * other-1

def __rtruediv__(self, other): return other * self**-1

def rtruediv(self, other): return other * self-1

for child, local_grad in zip(v._children, v._local_grads):

for child, local_grad in zip(v._children, v._local_grads):

我意识到这是数学和算法上最密集的部分,我有一个 2.5 小时的视频:micrograd 视频。简而言之,一个

I realize that this is the most mathematically and algorithmically intense part and I have a 2.5 hour video on it: micrograd video. Briefly, a

)并跟踪它是如何计算的。将每个操作视为一个小乐高积木:它接受一些输入,产生一个输出(前向传播),并且知道其输出相对于每个输入的变化方式(局部梯度)。这就是自动微分从每个积木中需要的所有信息。其他一切都只是链式法则,将积木串联起来。

) and tracks how it was computed. Think of each operation as a little lego block: it takes some inputs, produces an output (the forward pass), and it knows how its output would change with respect to each of its inputs (the local gradient). That’s all the information autograd needs from each block. Everything else is just the chain rule, stringing the blocks together.

对象(加法、乘法等),结果是一个新的

objects (add, multiply, etc.), the result is a new

)和该操作的局部导数(

) and the local derivative of that operation (

记录∂(a⋅b)/∂a=b 和∂(a⋅b)/∂b=a。完整的乐高积木集:

records that ∂(a⋅b)∂a=b∂(a⋅b)∂a=b and ∂(a⋅b)∂b=a∂(a⋅b)∂b=a. The full set of lego blocks:

方法以逆拓扑顺序遍历此图(从损失开始,到参数结束),每一步应用链式法则。如果损失是 L,节点 v 有一个子节点 c,局部梯度为∂v/∂c,那么:

method walks this graph in reverse topological order (starting from the loss, ending at the parameters), applying the chain rule at each step. If the loss is L and a node v has a child c with local gradient ∂v∂c, then:

如果你不熟悉微积分,这看起来有点吓人,但这实际上只是直观地乘以两个数字。一种理解方式如下:“如果汽车的速度是自行车的两倍,而自行车的速度是步行者的四倍,那么汽车的速度是步行者的 2×4=8 倍。”链式法则也是同样的思路:沿着路径乘以变化率。

This looks a bit scary if you’re not comfortable with your calculus, but this is literally just multiplying two numbers in an intuitive way. One way to see it looks as follows: “If a car travels twice as fast as a bicycle and the bicycle is four times as fast as a walking man, then the car travels 2 x 4 = 8 times as fast as the man.” The chain rule is the same idea: you multiply the rates of change along the path.

在损失节点处,因为∂L/∂L=1:损失相对于自身的变化率显然是 1。从那里开始,链式法则只是沿着每条路径将局部梯度乘回到参数。

at the loss node, because ∂L∂L=1: the loss’s rate of change with respect to itself is trivially 1. From there, the chain rule just multiplies local gradients along every path back to the parameters.

注意+=(累加,而非赋值)。当一个值在图中被多次使用时(即图分支),梯度沿每个分支独立流回,并且必须求和。这是多变量链式法则的结果:如果 c 通过多条路径对 L 有贡献,则总导数是每条路径贡献的总和。

Note the += (accumulation, not assignment). When a value is used in multiple places in the graph (i.e. the graph branches), gradients flow back along each branch independently and must be summed. This is a consequence of the multivariable chain rule: if c contributes to L through multiple paths, the total derivative is the sum of contributions from each path.

包含∂L/∂v,它告诉我们如果微调该值,最终损失将如何变化。

containing ∂L∂v, which tells us how the final loss would change if we nudged that value.

被使用了两次(图分支),所以其梯度是两条路径的总和:

is used twice (the graph branches), so its gradient is the sum of both paths:

print(a.grad) # 4.0 (dL/da = b + 1 = 3 + 1,通过两条路径)

print(a.grad) # 4.0 (dL/da = b + 1 = 3 + 1, via both paths)

a = torch.tensor(2.0, requires_grad=True)

a = torch.tensor(2.0, requires_grad=True)

b = torch.tensor(3.0, requires_grad=True)

b = torch.tensor(3.0, requires_grad=True)

这是 PyTorch 的

This is the same algorithm that PyTorch’s

运行的相同算法,只是在标量上而不是张量(标量数组)上——算法上相同,但更小更简单,当然效率低得多。

runs, just on scalars instead of tensors (arrays of scalars) - algorithmically identical, significantly smaller and simpler, but of course a lot less efficient.

上面给出的。自动微分计算出,如果

gives us above. Autograd calculated that if

告诉我们关于

is telling us about the local influence of

的局部影响,

would increase by about 4x that (0.004). Similarly,

将增加大约 4 倍(0.004)。类似地,

by about 2x that (0.002). In other words, these gradients tell us the direction (positive or negative depending on the sign), and the steepness (the magnitude) of the influence of each individual input on the final output (the loss). This then allows us to interately nudge the parameters of our neural network to lower the loss, and hence improve its predictions.

参数 Parameters

参数是模型的知识。它们是大量浮点数的集合(包裹在

The parameters are the knowledge of the model. They are a large collection of floating point numbers (wrapped in

中用于自动求导),初始时是随机的,并在训练过程中迭代优化。每个参数的具体作用在我们下面定义模型架构时会更加清晰,但现在我们只需要初始化它们:

for autograd) that start out random and are iteratively optimized during training. The exact role of each parameter will make more sense once we define the model architecture below, but for now we just need to initialize them:

n_head = 4 # 注意力头数

n_head = 4 # number of attention heads

block_size = 16 # 最大序列长度

block_size = 16 # maximum sequence length

head_dim = n_embd // n_head # 每个头的维度

head_dim = n_embd // n_head # dimension of each head

matrix = lambda nout, nin, std=0.08: [[Value(random.gauss(0, std)) for _ in range(nin)] for _ in range(nout)]

matrix = lambda nout, nin, std=0.08: [[Value(random.gauss(0, std)) for _ in range(nin)] for _ in range(nout)]

state_dict = {'wte': matrix(vocab_size, n_embd), 'wpe': matrix(block_size, n_embd), 'lm_head': matrix(vocab_size, n_embd)}

state_dict = {'wte': matrix(vocab_size, n_embd), 'wpe': matrix(block_size, n_embd), 'lm_head': matrix(vocab_size, n_embd)}

state_dict[f'layer{i}.attn_wq'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wq'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wk'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wk'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wv'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wv'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wo'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.attn_wo'] = matrix(n_embd, n_embd)

state_dict[f'layer{i}.mlp_fc1'] = matrix(4 * n_embd, n_embd)

state_dict[f'layer{i}.mlp_fc1'] = matrix(4 * n_embd, n_embd)

state_dict[f'layer{i}.mlp_fc2'] = matrix(n_embd, 4 * n_embd)

state_dict[f'layer{i}.mlp_fc2'] = matrix(n_embd, 4 * n_embd)

params = [p for mat in state_dict.values() for row in mat for p in row]

params = [p for mat in state_dict.values() for row in mat for p in row]

每个参数初始化为从高斯分布中抽取的小随机数。

Each parameter is initialized to a small random number drawn from a Gaussian distribution. The

将它们组织成命名的矩阵(借用 PyTorch 的术语):嵌入表、注意力权重、MLP 权重和最终输出投影。我们还将所有参数展平为一个列表

organizes them into named matrices (borrowing PyTorch’s terminology): embedding tables, attention weights, MLP weights, and a final output projection. We also flatten all parameters into a single list

以便优化器稍后可以遍历它们。在我们的微型模型中,这总共是 4,192 个参数。GPT-2 有 16 亿参数,而现代 LLM 有数千亿参数。

so the optimizer can loop over them later. In our tiny model this comes out to 4,192 parameters. GPT-2 had 1.6 billion, and modern LLMs have hundreds of billions.

架构 Architecture

模型架构是一个无状态函数:它接收一个词元、一个位置、参数以及来自先前位置的缓存键/值,并返回关于模型认为序列中下一个应出现的词元的 logits(得分)。我们遵循 GPT-2,但做了少量简化:使用 RMSNorm 代替 LayerNorm,无偏置,使用 ReLU 代替 GeLU。首先,三个小的辅助函数:

The model architecture is a stateless function: it takes a token, a position, the parameters, and the cached keys/values from previous positions, and returns logits (scores) over what token the model things should come next in the sequence. We follow GPT-2 with minor simplifications: RMSNorm instead of LayerNorm, no biases, and ReLU instead of GeLU. First, three small helper functions:

return [sum(wi * xi for wi, xi in zip(wo, x)) for wo in w]

return [sum(wi * xi for wi, xi in zip(wo, x)) for wo in w]

是一个矩阵-向量乘法。它接收一个向量,并计算该向量与矩阵每一行的点积。这是神经网络的基本构建块:一个学习到的线性变换。

is a matrix-vector multiply. It takes a vector

max_val = max(val.data for val in logits)

, and computes one dot product per row of

exps = [(val - max_val).exp() for val in logits]

. This is the fundamental building block of neural networks: a learned linear transformation.

将原始得分(logits)向量(范围从−∞到+∞)转换为概率分布:所有值都在[0,1]内且总和为 1。我们首先减去最大值以保证数值稳定性(这在数学上不会改变结果,但防止了溢出)。

max_val = max(val.data for val in logits)

(均方根归一化)重新缩放向量,使其值具有单位均方根。这可以防止激活值在网络中传播时增长或缩小,从而稳定训练。它是原始 GPT-2 中使用的 LayerNorm 的一个更简单的变体。

exps = [(val - max_val).exp() for val in logits]

tok_emb = state_dict['wte'][token_id] # 词元嵌入

converts a vector of raw scores (logits), which can range from −∞ to +∞, into a probability distribution: all values end up in [0,1] and sum to 1. We subtract the max first for numerical stability (it doesn’t change the result mathematically, but prevents overflow in

pos_emb = state_dict['wpe'][pos_id] # 位置嵌入

(Root Mean Square Normalization) rescales a vector so its values have unit root-mean-square. This keeps activations from growing or shrinking as they flow through the network, which stabilizes training. It’s a simpler variant of the LayerNorm used in the original GPT-2.

x = [t + p for t, p in zip(tok_emb, pos_emb)] # 联合词元和位置嵌入

tok_emb = state_dict['wte'][token_id] # token embedding

pos_emb = state_dict['wpe'][pos_id] # position embedding

x = [t + p for t, p in zip(tok_emb, pos_emb)] # joint token and position embedding

1) 多头注意力模块 1) Multi-head attention block

q = linear(x, state_dict[f'layer{li}.attn_wq'])

q = linear(x, state_dict[f'layer{li}.attn_wq'])

k = linear(x, state_dict[f'layer{li}.attn_wk'])

k = linear(x, state_dict[f'layer{li}.attn_wk'])

v = linear(x, state_dict[f'layer{li}.attn_wv'])

v = linear(x, state_dict[f'layer{li}.attn_wv'])

k_h = [ki[hs:hs+head_dim] for ki in keys[li]]

k_h = [ki[hs:hs+head_dim] for ki in keys[li]]

v_h = [vi[hs:hs+head_dim] for vi in values[li]]

v_h = [vi[hs:hs+head_dim] for vi in values[li]]

attn_logits = [sum(q_h[j] * k_h[t][j] for j in range(head_dim)) / head_dim0.5 for t in range(len(k_h))]

attn_logits = [sum(q_h[j] * k_h[t][j] for j in range(head_dim)) / head_dim0.5 for t in range(len(k_h))]

head_out = [sum(attn_weights[t] * v_h[t][j] for t in range(len(v_h))) for j in range(head_dim)]

head_out = [sum(attn_weights[t] * v_h[t][j] for t in range(len(v_h))) for j in range(head_dim)]

x = linear(x_attn, state_dict[f'layer{li}.attn_wo'])

x = linear(x_attn, state_dict[f'layer{li}.attn_wo'])

x = [a + b for a, b in zip(x, x_residual)]

x = [a + b for a, b in zip(x, x_residual)]

2) MLP 块 2) MLP block

x = linear(x, state_dict[f'layer{li}.mlp_fc1'])

x = linear(x, state_dict[f'layer{li}.mlp_fc1'])

x = linear(x, state_dict[f'layer{li}.mlp_fc2'])

x = linear(x, state_dict[f'layer{li}.mlp_fc2'])

x = [a + b for a, b in zip(x, x_residual)]

x = [a + b for a, b in zip(x, x_residual)]

logits = linear(x, state_dict['lm_head'])

logits = linear(x, state_dict['lm_head'])

),以及由激活值总结的先前迭代的一些上下文

), and some context from the previous iterations summarized by the activations in

,即 KV 缓存。以下是逐步发生的情况:

, known as the KV Cache. Here’s what happens step by step:

嵌入。神经网络无法直接处理像 5 这样的原始 token ID。它只能处理向量(数字列表)。因此,我们将每个可能的 token 与一个学习到的向量关联起来,并将其作为神经特征输入。token ID 和位置 ID 各自从其对应的嵌入表中查找一行(

Embeddings. The neural network can’t process a raw token id like 5 directly. It can only work with vectors (lists of numbers). So we associate a learned vector with each possible token, and feed that in as its neural signature. The token id and position id each look up a row from their respective embedding tables (

)。这两个向量相加,为模型提供一个编码了 token 是什么以及它在序列中位置的表示。现代 LLM 通常跳过位置嵌入,引入其他基于相对位置的方案,例如 RoPE。

). These two vectors are added together, giving the model a representation that encodes both _what_ the token is and _where_ it is in the sequence. Modern LLMs usually skip the position embedding and introduce other relative-based positioning schemes, e.g. RoPE.

注意力块。当前 token 被投影为三个向量:查询 (Q)、键 (K) 和值 (V)。直观地说,查询表示“我在找什么?”,键表示“我包含什么?”,值表示“如果被选中,我提供什么?”。例如,在名字“emma”中,当模型处于第二个“m”并试图预测下一个词时,它可能会学习一个类似“最近出现了哪些元音?”的查询。前面的“e”会有一个与这个查询匹配得很好的键,因此它获得较高的注意力权重,其值(关于它是元音的信息)流入当前位置。键和值被追加到 KV 缓存中,以便之前的位置可用。每个注意力头计算其查询与所有缓存键的点积(按 √d_head 缩放),应用 softmax 得到注意力权重,并取缓存值的加权和。所有头的输出被拼接并通过

Attention block. The current token is projected into three vectors: a query (Q), a key (K), and a value (V). Intuitively, the query says “what am I looking for?”, the key says “what do I contain?”, and the value says “what do I offer if selected?”. For example, in the name “emma”, when the model is at the second “m” and trying to predict what comes next, it might learn a query like “what vowels appeared recently?” The earlier “e” would have a key that matches this query well, so it gets a high attention weight, and its value (information about being a vowel) flows into the current position. The key and value are appended to the KV cache so previous positions are available. Each attention head computes dot products between its query and all cached keys (scaled by √d h e a d), applies softmax to get attention weights, and takes a weighted sum of the cached values. The outputs of all heads are concatenated and projected through

投影。值得强调的是,注意力块是位置

. It’s worth emphasizing that the Attention block is the exact and only place where a token at position

处 token 进行通信的唯一确切位置。注意力是一种 token 通信机制。

. Attention is a token communication mechanism.

MLP 块。MLP 是“多层感知机”的缩写,它是一个两层前馈网络:投影到嵌入维度的 4 倍,应用 ReLU,再投影回原维度。这是模型在每个位置上进行大部分“思考”的地方。与注意力不同,这个计算完全局限于时间

MLP block. MLP is short for “multilayer perceptron”, it is a two-layer feed-forward network: project up to 4x the embedding dimension, apply ReLU, project back down. This is where the model does most of its “thinking” per position. Unlike attention, this computation is fully local to time

。Transformer 交错进行通信(注意力)和计算(MLP)。

. The Transformer intersperses communication (Attention) with computation (MLP).

残差连接。注意力块和 MLP 块都将其输出加回到其输入(

Residual connections. Both the attention and MLP blocks add their output back to their input (

)。这使得梯度可以直接流过网络,并使更深的模型可训练。

). This lets gradients flow directly through the network and makes deeper models trainable.

输出。最终的隐藏状态通过

Output. The final hidden state is projected to vocabulary size by

投影到词汇表大小,为词汇表中的每个 token 生成一个 logit。在我们的例子中,只有 27 个数字。更高的 logit 意味着模型认为对应的 token 更有可能出现在下一个。

, producing one logit per token in the vocabulary. In our case, that’s just 27 numbers. Higher logit = the model thinks that corresponding token is more likely to come next.

你可能会注意到我们在训练期间使用了 KV 缓存,这并不常见。人们通常将 KV 缓存与推理关联起来。但 KV 缓存概念上始终存在,即使在训练期间也是如此。在生产实现中,它隐藏在高度向量化的注意力计算中,该计算同时处理序列中的所有位置。由于 microgpt 一次处理一个 token(没有批次维度,没有并行时间步),我们显式地构建了 KV 缓存。与典型的推理设置(KV 缓存持有分离的张量)不同,这里的缓存键和值是计算图中的活跃节点,因此我们实际上会通过它们进行反向传播。

You might notice that we’re using a KV cache during training, which is unusual. People typically associate the KV cache with inference only. But the KV cache is conceptually always there, even during training. In production implementations, it’s just hidden inside the highly vectorized attention computation that processes all positions in the sequence simultaneously. Since microgpt processes one token at a time (no batch dimension, no parallel time steps), we build the KV cache explicitly. And unlike the typical inference setting where the KV cache holds detached tensors, here the cached keys and values are live

nodes in the computation graph, so we actually backpropagate through them.

训练循环 Training loop

现在我们将所有部分整合在一起。训练循环反复执行以下步骤:(1) 选取一个文档,(2) 在文档的词元上运行模型前向传播,(3) 计算损失,(4) 反向传播以获取梯度,(5) 更新参数。

Now we wire everything together. The training loop repeatedly: (1) picks a document, (2) runs the model forward over its tokens, (3) computes a loss, (4) backpropagates to get gradients, and (5) updates the parameters.

亚当,受祝福的优化器及其缓冲区 Let there be Adam, the blessed optimizer and its buffers

learning_rate, beta1, beta2, eps_adam = 0.01, 0.85, 0.99, 1e-8

learning_rate, beta1, beta2, eps_adam = 0.01, 0.85, 0.99, 1e-8

m = [0.0] * len(params) # 一阶矩缓冲区

m = [0.0] * len(params) # first moment buffer

v = [0.0] * len(params) # 二阶矩缓冲区

v = [0.0] * len(params) # second moment buffer

序列重复 Repeat in sequence

num_steps = 1000 # 训练步数

num_steps = 1000 # number of training steps

对单个文档进行分词,并在两侧添加 BOS 特殊标记 Take single document, tokenize it, surround it with BOS special token on both sides

tokens = [BOS] + [uchars.index(ch) for ch in doc] + [BOS]

tokens = [BOS] + [uchars.index(ch) for ch in doc] + [BOS]

将词元序列前向传播通过模型,构建直至损失函数的计算图 Forward the token sequence through the model, building up the computation graph all the way to the loss.

keys, values = [[] for _ in range(n_layer)], [[] for _ in range(n_layer)]

keys, values = [[] for _ in range(n_layer)], [[] for _ in range(n_layer)]

token_id, target_id = tokens[pos_id], tokens[pos_id + 1]

token_id, target_id = tokens[pos_id], tokens[pos_id + 1]

logits = gpt(token_id, pos_id, keys, values)

logits = gpt(token_id, pos_id, keys, values)

loss = (1 / n) * sum(losses) # 文档序列上的最终平均损失。愿你的损失值低。

loss = (1 / n) * sum(losses) # final average loss over the document sequence. May yours be low.

Adam 优化器更新 Adam optimizer update: update the model parameters based on the corresponding gradients.

lr_t = learning_rate * (1 - step / num_steps) # 线性学习率衰减

lr_t = learning_rate * (1 - step / num_steps) # linear learning rate decay

m[i] = beta1 * m[i] + (1 - beta1) * p.grad

m[i] = beta1 * m[i] + (1 - beta1) * p.grad

v[i] = beta2 * v[i] + (1 - beta2) * p.grad^2

v[i] = beta2 * v[i] + (1 - beta2) * p.grad 2

p.data -= lr_t * m_hat / (v_hat^0.5 + eps_adam)

p.data -= lr_t * m_hat / (v_hat 0.5 + eps_adam)

print(f"step {step+1:4d} / {num_steps:4d} | loss {loss.data:.4f}")

print(f"step {step+1:4d} / {num_steps:4d} | loss {loss.data:.4f}")

分词。每个训练步骤选取一个文档,并用<|endoftext|>将其包裹。模型的任务是根据前面的词元预测下一个词元。

Tokenization. Each training step picks one document and wraps it with

前向传播与损失。我们逐个将词元输入模型,同时构建 KV 缓存。在每个位置,模型输出 27 个 logits,通过 softmax 转换为概率。每个位置的损失是正确下一个词元的负对数概率:−log p(target)。这称为交叉熵损失。直观上,损失衡量预测错误的程度:模型对实际出现的下一个词元的惊讶程度。如果模型对正确词元赋予概率 1.0,则完全不惊讶,损失为 0;如果赋予接近 0 的概率,则非常惊讶,损失趋于+∞。我们对文档中所有位置的损失取平均,得到单个标量损失。

. The model’s job is to predict each next token given the tokens before it.

通过整个计算图反向传播,从损失一直回溯到 softmax、模型以及每个参数。之后,每个参数的.grad 告诉我们如何改变它以减少损失。

Forward pass and loss. We feed the tokens through the model one at a time, building up the KV cache as we go. At each position, the model outputs 27 logits, which we convert to probabilities via softmax. The loss at each position is the negative log probability of the correct next token: −log p(target). This is called the cross-entropy loss. Intuitively, the loss measures the degree of misprediction: how surprised the model is by what actually comes next. If the model assigns probability 1.0 to the correct token, it is not surprised at all and the loss is 0. If it assigns probability close to 0, the model is very surprised and the loss goes to +∞. We average the per-position losses across the document to get a single scalar loss.

(梯度下降),但 Adam 更智能。它维护每个参数的两个运行平均值:

runs backpropagation through the entire computation graph, from the loss all the way back through softmax, the model, and into every parameter. After this, each parameter’s

跟踪近期梯度的均值(动量,类似滚动的球),而

tells us how to change it to reduce the loss.

跟踪近期梯度平方的均值(自适应每个参数的学习率)。

(gradient descent), but Adam is smarter. It maintains two running averages per parameter:

是偏差校正,用于补偿 m 和 v 初始化为零而需要预热。学习率在训练过程中线性衰减。更新后,我们将.grad 重置为零。

tracks the mean of recent gradients (momentum, like a rolling ball), and

经过 1000 步,损失从约 3.3(27 个词元中随机猜测:−log(1/27)≈3.3)下降到约 2.37。越低越好,最低可能为 0(完美预测),因此仍有改进空间,但模型显然在学习名字的统计模式。

tracks the mean of recent squared gradients (adapting the learning rate per parameter). The

are bias corrections that account for the fact that

are initialized to zero and need a warmup. The learning rate decays linearly over training. After updating, we reset

Over 1,000 steps the loss decreases from around 3.3 (random guessing among 27 tokens: −log(1/27)≈3.3) down to around 2.37. Lower is better, and the lowest possible is 0 (perfect predictions), so there’s still room to improve, but the model is clearly learning the statistical patterns of names.

推理 Inference

训练完成后,我们可以从模型中采样新名称。参数被冻结,我们只需循环运行前向传播,将每个生成的词元作为下一个输入反馈回去:

Once training is done, we can sample new names from the model. The parameters are frozen and we just run the forward pass in a loop, feeding each generated token back as the next input:

temperature = 0.5 # 在 (0, 1] 之间,控制生成文本的“创造性”,从低到高

temperature = 0.5 # in (0, 1], control the "creativity" of generated text, low to high

print("\n--- 推理(新的、幻觉名称) ---")

print("\n--- inference (new, hallucinated names) ---")

keys, values = [[] for _ in range(n_layer)], [[] for _ in range(n_layer)]

keys, values = [[] for _ in range(n_layer)], [[] for _ in range(n_layer)]

logits = gpt(token_id, pos_id, keys, values)

logits = gpt(token_id, pos_id, keys, values)

probs = softmax([l / temperature for l in logits])

probs = softmax([l / temperature for l in logits])

token_id = random.choices(range(vocab_size), weights=[p.data for p in probs])[0]

token_id = random.choices(range(vocab_size), weights=[p.data for p in probs])[0]

print(f"sample {sample_idx+1:2d}: {''.join(sample)}")

print(f"sample {sample_idx+1:2d}: {''.join(sample)}")

词元告诉模型“开始一个新名称”。模型产生 27 个 logits,我们将其转换为概率,并根据这些概率随机采样一个词元。该词元被反馈作为下一个输入,我们重复此过程直到模型产生

token, which tells the model “begin a new name”. The model produces 27 logits, we convert them to probabilities, and we randomly sample one token according to those probabilities. That token gets fed back in as the next input, and we repeat until the model produces

再次(表示“我完成了”)或达到最大序列长度。

again (meaning “I’m done”) or we hit the maximum sequence length.

参数控制随机性。在 softmax 之前,我们将 logits 除以温度。温度为 1.0 时直接从模型学习到的分布中采样。较低的温度(如这里的 0.5)会使分布更尖锐,使模型更保守,更倾向于选择其首选。温度趋近于 0 时,将始终选择最可能的单个词元(贪婪解码)。较高的温度会使分布更平坦,产生更多样化但可能不太连贯的输出。

parameter controls randomness. Before softmax, we divide the logits by the temperature. A temperature of 1.0 samples directly from the model’s learned distribution. Lower temperatures (like 0.5 here) sharpen the distribution, making the model more conservative and likely to pick its top choices. A temperature approaching 0 would always pick the single most likely token (greedy decoding). Higher temperatures flatten the distribution and produce more diverse but potentially less coherent output.

运行它 Run it

你只需要 Python(无需 pip install,无需依赖):

All you need is Python (no pip install, no dependencies):

该脚本在我的 MacBook 上运行大约需要 1 分钟。你会看到每一步打印的损失值:

The script takes about 1 minute to run on my macbook. You’ll see the loss printed at each step:

观察它从约 3.3(随机)下降到约 2.37。这个数字越低,说明网络对序列中下一个词元的预测就越好。训练结束时,训练词元序列的统计模式知识被蒸馏到模型参数中。固定这些参数,我们现在可以生成新的、虚构的名字。你会再次看到:

Watch it go down from ~3.3 (random) toward ~2.37. The lower this number is, the better the network’s predictions already were about what token comes next in the sequence. At the end of training, the knowledge of the stastical patterns of the training token sequences is distilled in the model parameters. Fixing these parameters, we can now generate new, hallucinated names. You’ll see (again):

除了在你的电脑上运行脚本,你也可以尝试直接在这个 Google Colab 笔记本上运行它,并向 Gemini 提问。尝试修改脚本!你可以使用不同的数据集。或者你可以训练更长时间(增加训练步数)或增大模型规模,以获得越来越好的结果。

As an alternative to running the script on your computer, you may try to run it directly on this Google Colab notebook and ask Gemini questions about it. Try playing with the script! You can try a different dataset. Or you can train for longer (increase

) or increase the size of the model to get increasingly better results.

渐进式构建 Progression

为了像剥洋葱一样逐层构建代码,推荐的渐进式构建过程如下:

To see the code built up piece by piece as layers of the onion, the advised progression looks something like this:

我创建了一个名为 build_microgpt.py 的 Gist,在修订历史中你可以看到所有这些版本以及每一步之间的差异。我认为这可能是一种逐步浏览代码库的有用方式,每次只添加一个组件。

I created a Gist called build_microgpt.py, where in the Revisions you can see all of these versions and the diffs between each step. I think this might be one helpful way to step through the code base, where you add one component at a time.

实际内容 Real stuff

microgpt 包含了训练和运行 GPT 的完整算法精髓。但在此与生产级 LLM(如 ChatGPT)之间,有一长串变化。它们都没有改变核心算法和整体布局,但正是这些变化使其能够大规模实际运行。按相同顺序依次介绍:

microgpt contains the complete algorithmic essence of training and running a GPT. But between this and a production LLM like ChatGPT, there is a long list of things that change. None of them alter the core algorithm and the overall layout, but they are what makes it actually work at scale. Walking through the same sections in order:

数据。生产模型不再使用 32K 个短名称,而是在数万亿个互联网文本 token 上训练:网页、书籍、代码等。数据经过去重、质量过滤,并跨领域精心混合。

Data. Instead of 32K short names, production models train on trillions of tokens of internet text: web pages, books, code, etc. The data is deduplicated, filtered for quality, and carefully mixed across domains.

分词器。生产模型不再使用单个字符,而是采用子词分词器,如 BPE(字节对编码),它学习将频繁共现的字符序列合并为单个 token。像“the”这样的常见词成为一个 token,罕见词则被拆分成片段。这提供了约 100K 个 token 的词汇表,并且效率更高,因为模型在每个位置上看到更多内容。

Tokenizer. Instead of single characters, production models use subword tokenizers like BPE (Byte Pair Encoding), which learn to merge frequently co-occurring character sequences into single tokens. Common words like “the” become a single token, rare words get broken into pieces. This gives a vocabulary of ~100K tokens and is much more efficient because the model sees more content per position.

对象。生产系统使用张量(大型多维数字数组),并在每秒执行数十亿次浮点运算的 GPU/TPU 上运行。像 PyTorch 这样的库处理张量的自动求导,而像 FlashAttention 这样的 CUDA 内核融合多个操作以提高速度。数学原理相同,只是对应并行处理的许多标量。

objects in pure Python. Production systems use tensors (large multi-dimensional arrays of numbers) and run on GPUs/TPUs that perform billions of floating point operations per second. Libraries like PyTorch handle the autograd over tensors, and CUDA kernels like FlashAttention fuse multiple operations for speed. The math is identical, just corresponds to many scalars processed in parallel.

架构。microgpt 有 4,192 个参数。GPT-4 类模型有数千亿个参数。总体而言,它是一个非常相似的 Transformer 神经网络,只是更宽(嵌入维度超过 10,000)和更深(超过 100 层)。现代 LLM 还包含一些其他类型的乐高积木并改变其顺序:例如,使用 RoPE(旋转位置嵌入)代替学习的位置嵌入,使用 GQA(分组查询注意力)以减少 KV 缓存大小,使用门控线性激活代替 ReLU,混合专家(MoE)层等。但注意力机制(通信)和 MLP(计算)交错在残差流上的核心结构得到了很好的保留。

Architecture. microgpt has 4,192 parameters. GPT-4 class models have hundreds of billions. Overall it’s a very similar looking Transformer neural network, just much wider (embedding dimensions of 10,000+) and much deeper (100+ layers). Modern LLMs also incorporate a few more types of lego blocks and change their orders around: Examples include RoPE (Rotary Position Embeddings) instead of learned position embeddings, GQA (Grouped Query Attention) to reduce KV cache size, gated linear activations instead of ReLU, Mixture of Experts (MoE) layers, etc. But the core structure of Attention (communication) and MLP (computation) interspersed on a residual stream is well-preserved.

训练。生产训练不再每步一个文档,而是使用大批量(每步数百万个 token)、梯度累积、混合精度(float16/bfloat16)和仔细的超参数调优。训练一个前沿模型需要数千个 GPU 运行数月。

Training. Instead of one document per step, production training uses large batches (millions of tokens per step), gradient accumulation, mixed precision (float16/bfloat16), and careful hyperparameter tuning. Training a frontier model takes thousands of GPUs running for months.

优化。microgpt 使用 Adam 和简单的线性学习率衰减,仅此而已。在大规模下,优化成为一门独立的学科。模型以降低的精度(bfloat16 甚至 fp8)训练,并在大型 GPU 集群上运行以提高效率,这带来了自身的数值挑战。优化器设置(学习率、权重衰减、beta 参数、预热计划、衰减计划)必须精确调整,正确的值取决于模型大小、批量大小和数据集组成。缩放定律(例如 Chinchilla)指导如何在模型大小和训练 token 数量之间分配固定的算力预算。在大规模下,任何细节出错都可能浪费数百万美元的算力,因此团队在投入完整训练运行之前,会进行大量的小规模实验来预测正确的设置。

Optimization. microgpt uses Adam with a simple linear learning rate decay and that’s about it. At scale, optimization becomes its own discipline. Models train in reduced precision (bfloat16 or even fp8) and across large GPU clusters for efficiency, which introduces its own numerical challenges. The optimizer settings (learning rate, weight decay, beta parameters, warmup schedule, decay schedule) must be tuned precisely, and the right values depend on model size, batch size, and dataset composition. Scaling laws (e.g. Chinchilla) guide how to allocate a fixed compute budget between model size and number of training tokens. Getting any of these details wrong at scale can waste millions of dollars of compute, so teams run extensive smaller-scale experiments to predict the right settings before committing to a full training run.

后训练。训练得到的基础模型(称为“预训练”模型)是一个文档补全器,而不是聊天机器人。将其转变为 ChatGPT 分为两个阶段。首先,SFT(监督微调):只需将文档替换为精心策划的对话并继续训练。算法上,没有任何变化。其次,RL(强化学习):模型生成回复,这些回复被评分(由人类、另一个“评判”模型或算法),模型从该反馈中学习。从根本上说,模型仍然在文档上训练,但这些文档现在由来自模型自身的 token 组成。

Post-training. The base model that comes out of training (called the “pretrained” model) is a document completer, not a chatbot. Turning it into ChatGPT happens in two stages. First, SFT (Supervised Fine-Tuning): you simply swap the documents for curated conversations and keep training. Algorithmically, nothing changes. Second, RL (Reinforcement Learning): the model generates responses, they get scored (by humans, another “judge” model, or an algorithm), and the model learns from that feedback. Fundamentally, the model is still training on documents, but those documents are now made up of tokens coming from the model itself.

推理。向数百万用户提供模型服务需要其自身的工程栈:将请求批处理在一起、KV 缓存管理和分页(vLLM 等)、用于加速的推测解码、用于减少内存的量化(以 int8/int4 而非 float16 运行),以及将模型分布到多个 GPU 上。从根本上说,我们仍然在预测序列中的下一个 token,但投入了大量工程来使其更快。

Inference. Serving a model to millions of users requires its own engineering stack: batching requests together, KV cache management and paging (vLLM, etc.), speculative decoding for speed, quantization (running in int8/int4 instead of float16) to reduce memory, and distributing the model across multiple GPUs. Fundamentally, we are still predicting the next token in the sequence but with a lot of engineering spent on making it faster.

所有这些都是重要的工程和研究贡献,但如果你理解了 microgpt,你就理解了算法的精髓。

All of these are important engineering and research contributions but if you understand microgpt, you understand the algorithmic essence.

常见问题 FAQ

模型是否“理解”任何东西?这是一个哲学问题,但从机制上讲:没有魔法发生。模型是一个大的数学函数,将输入词元映射到下一个词元的概率分布。在训练期间,调整参数以使正确的下一个词元更有可能。这是否构成“理解”由你决定,但机制完全包含在上面的 200 行代码中。

Does the model “understand” anything? That’s a philosophical question, but mechanically: no magic is happening. The model is a big math function that maps input tokens to a probability distribution over the next token. During training, the parameters are adjusted to make the correct next token more probable. Whether this constitutes “understanding” is up to you, but the mechanism is fully contained in the 200 lines above.

为什么它能工作?模型有数千个可调参数,优化器每一步都会轻微调整它们以使损失下降。经过许多步,参数稳定在能够捕捉数据统计规律的值上。对于名字,这意味着:名字通常以辅音开头,“qu”往往一起出现,名字很少连续有三个辅音,等等。模型不学习显式规则,它学习一个恰好反映这些规则的概率分布。

Why does it work? The model has thousands of adjustable parameters, and the optimizer nudges them a tiny bit each step to make the loss go down. Over many steps, the parameters settle into values that capture the statistical regularities of the data. For names, this means things like: names often start with consonants, “qu” tends to appear together, names rarely have three consonants in a row, etc. The model doesn’t learn explicit rules, it learns a probability distribution that happens to reflect them.

这与 ChatGPT 有何关系?ChatGPT 是同一个核心循环(预测下一个词元,采样,重复)的巨大规模扩展,并经过后训练使其具有对话能力。当你与它聊天时,系统提示、你的消息和它的回复都只是序列中的词元。模型一次一个词元地完成文档,就像 microgpt 完成一个名字一样。

How is this related to ChatGPT? ChatGPT is this same core loop (predict next token, sample, repeat) scaled up enormously, with post-training to make it conversational. When you chat with it, the system prompt, your message, and its reply are all just tokens in a sequence. The model is completing the document one token at a time, same as microgpt completing a name.

“幻觉”是怎么回事?模型通过从概率分布中采样来生成词元。它没有真理的概念,它只知道哪些序列在统计上根据训练数据是合理的。microgpt“幻觉”出一个像“karia”这样的名字,与 ChatGPT 自信地陈述一个错误事实是同一现象。两者都是听起来合理的补全,但恰好不是真实的。

What’s the deal with “hallucinations”? The model generates tokens by sampling from a probability distribution. It has no concept of truth, it only knows what sequences are statistically plausible given the training data. microgpt “hallucinating” a name like “karia” is the same phenomenon as ChatGPT confidently stating a false fact. Both are plausible-sounding completions that happen not to be real.

为什么这么慢?microgpt 在纯 Python 中一次处理一个标量。一个训练步骤需要几秒钟。同样的数学运算在 GPU 上并行处理数百万个标量,运行速度快几个数量级。

Why is it so slow? microgpt processes one scalar at a time in pure Python. A single training step takes seconds. The same math on a GPU processes millions of scalars in parallel and runs orders of magnitude faster.

我能让它生成更好的名字吗?可以。训练更长时间(增加步数),或使用更大的数据集。这些是在大规模下同样重要的调节旋钮。

Can I make it generate better names? Yes. Train longer (increase

如果我更改数据集会怎样?模型将学习数据中的任何模式。换成城市名、宝可梦名、英语单词或短诗的文件,模型将学习生成这些内容。代码的其余部分无需更改。

), or use a larger dataset. These are the same knobs that matter at scale.

What if I change the dataset? The model will learn whatever patterns are in the data. Swap in a file of city names, Pokemon names, English words, or short poems, and the model will learn to generate those instead. The rest of the code doesn’t need to change.

互动版:图/公式 + 针对本篇提问 →