LLM-Powered Sorting with TrueSkill
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→尽管大语言模型(LLM)在理解和比较概念方面表现出色,但让它们持续对大量数据进行排序仍然非常困难。
The Challenge with LLM SortingEnter TrueSkillImplementationLLM Sorting PromptAn ExampleWhen are you “done”?AlternativesIndividual ScoringEmbedding-based SortingOptimizationSmart Batch SelectionUsing Confidence IntervalsConclusion Thariq Shihipar - 11 February 2025 · 7 min read Large Language Models (LLMs) are remarkably good at understanding and comparing concepts, but getting them to consistently sort large amounts of data is still quite difficult.
LLM 排序的挑战·引入 TrueSkill·实现·LLM 排序提示·示例·何时算“完成”?·替代方案·个体评分·基于嵌入的排序·优化·智能批量选择·使用置信区间·结论
The Challenge with LLM SortingEnter TrueSkillImplementationLLM Sorting PromptAn ExampleWhen are you “done”?AlternativesIndividual ScoringEmbedding-based SortingOptimizationSmart Batch SelectionUsing Confidence IntervalsConclusion
Thariq Shihipar - 2025 年 2 月 11 日·阅读时间 7 分钟
Thariq Shihipar - 11 February 2025 · 7 min read
大型语言模型(LLMs)在理解和比较概念方面非常出色,但让它们持续地对大量数据进行排序仍然相当困难。
Large Language Models (LLMs) are remarkably good at understanding and comparing concepts, but getting them to consistently sort large amounts of data is still quite difficult.
在这篇文章中,我将分享一种技术,该技术结合了 LLMs 的语义理解与 TrueSkill 等 ELO 算法的数学严谨性,以创建稳健且可扩展的语义排序。
In this post, I share a technique that combines the semantic understanding of LLMs with the mathematical rigor of ELO algorithms like TrueSkill to create robust, scalable semantic sorting.
假设您希望 LLM 根据提示创建一个有序列表。例如,您可能对求职者、功能请求、电影推荐或客户反馈进行排序。
Let’s say you’d like a LLM to create an ordered list of items, based on a prompt. For example, you might sort job candidates, feature requests, movie recommendations, or customer feedback.
当您要求 LLM 自行排序一个大型列表时,会遇到几个问题:
When you ask an LLM to sort a large list of items by itself you run into several problems:
1. 上下文窗口限制 - 您一次只能输入有限数量的项目
1. Context window limitations - You can only feed so many items at once
2. 一致性问题 - 模型可能判断 A > B 且 B > C,但另一次却判断 C > A
2. Consistency issues - The model might rank A > B and B > C, but then rank C > A another time
3. 质量下降 - 随着列表增长,注意力机制退化,排序质量往往下降
3. Quality degradation - As the list grows longer, the quality of sorting tends to decrease as attention degrades
4. 令牌限制 - 输出大型排序列表可能触及输出令牌上限
4. Token limits - Getting back large sorted lists can bump up against output token limits
TrueSkill(或任何 ELO 算法),最初由微软为 Xbox Live 匹配开发,为这些挑战提供了优雅的解决方案。TrueSkill 使微软能够比较两个从未一起玩过的玩家,利用他们的游戏历史来估计技能水平。
TrueSkill (or any ELO algorithm), originally developed by Microsoft for Xbox Live matchmaking, provides an elegant solution to these challenges. Trueskill allowed Microsoft to compare two players who had never played before, using the history of their games to estimate their skill level.
在这种背景下,我们可以将每个项目视为一名玩家,它们之间的“游戏”由 LLM 通过在小批次中排序来运行。
In this context, we can think of each item as a player, and the “games” between them are run by the LLM sorting them in a small batch.
因此,我们不是要求 LLM 一次性对所有内容进行排序,而是可以:
So, instead of asking a LLM to sort everything at once, we can:
* 将数据分解为小的、可管理的批次
* Break the data into small, manageable batches
* 让 LLM 对这些小批次进行排序(在它们之间进行“游戏”)
* Have the LLM order these smaller batches _(play a “game” between them)_
* 使用 TrueSkill 根据游戏结果更新项目的评分
* Use TrueSkill to update the ratings of the items based on the results of the games
最后,我们将获得项目的全局排名以及每个项目评分的置信区间,这些信息基于 LLM 的知识。
In the end, we will have a global ranking of the items, and a confidence interval for each item’s rating, informed by the LLMs knowledge.
这可以扩展到非常大的数据集,并且新项目可以增量添加,无需重新排序所有内容。
This scales to very large datasets, and new items can be added incrementally without re-sorting everything.
1. 为数据集中的每个项目初始化一个 TrueSkill 评分
1. Initialize a TrueSkill rating for each item in your dataset
* 抽样一小批项目(例如 10 个)
* Sample a small batch of items (e.g., 10)
* 根据这些结果更新 TrueSkill 评分
* Update TrueSkill ratings based on these results
3. 在足够多的迭代后,按保守评分估计(μ - 3σ)排序
3. After sufficient iterations, sort by the conservative rating estimate (μ - 3σ)
以下是核心循环的简化版本:
Here’s a simplified version of the core loop:
item_id: trueskill.Rating() for item_id in all_items
item_id: trueskill.Rating() for item_id in all_items
batch = random.sample(all_items, batch_size)
batch = random.sample(all_items, batch_size)
teams = [[ratings[item]] for item in sorted_batch]
teams = [[ratings[item]] for item in sorted_batch]
使其良好运作的关键在于为 LLM 排序步骤编写有效的提示词。以下是一个提示模板示例,但您应根据自身用例进行定制。
The key to making this work well is crafting good prompts for the LLM sorting step. Here’s an example prompt template, but you should customize it to your use case.
prompt = f"""您是一位分析{domain}的专家。
prompt = f"""You are an expert at analyzing {domain}.
{"\n".join(f'{item["id"]}. "{item["text"]}"' for item in items)}
{"\n".join(f'{item["id"]}. "{item["text"]}"' for item in items)}
仅返回一个逗号分隔的 ID 列表,按顺序排列,
Return ONLY a comma-separated list of ids sorted in order,
不包含其他文本或说明。例如:“123,456,789”
with no other text or explanation. For example: "123,456,789"
通过要求特定的输出格式并保持任务聚焦,我们可以从 LLM 获得更可靠的结果。
By requesting a specific output format and keeping the task focused, we get more reliable results from the LLM.
假设你有一万多条过去一年收集的客户反馈,你想按“可操作性”对它们进行分类。
Let’s say you have over 10,000 pieces of customer feedback you’ve collected over the past year. And you want to organize it by “actionability”.
你可以为可操作性编写一个你喜欢的提示词,让 LLM 结合技术可行性、潜在影响和请求的清晰度等因素进行综合判断。甚至可能结合你的代码库或产品路线图。
You can write a prompt for actionability that you like, getting the LLM to combine factors like technical feasibility, potential impact, and clarity of the request. Maybe even your codebase or product roadmap.
* "要是能更像 Excel 那样工作就好了"
* "Would be nice if it worked more like Excel"
* "在我打开太多标签页的时候有时候会崩溃"
* "Sometimes crashes when I have too many tabs open"
* "喜欢这个产品,但希望它能更快"
* "Love the product but wish it was faster"
传统的排序方法在这里面临困难是因为:
Traditional sorting methods struggle here because:
* 简单的关键词匹配或嵌入会遗漏含义
* Simple keyword matchin or embedding misses the meaning
* 手动对一万个项目进行分类是难以承受的
* Manual sorting of 10,000 items is overwhelming
* 每条反馈可能因为不同原因具有可操作性
* Each piece of feedback might be actionable for different reasons
* LLM 无法同时处理一万个项目
* LLMs can’t handle 10,000 items at once
系统会将它们分成更小的批次,让 LLM 进行排序,或许像这样:
The system will then break these into smaller batches, and ask the LLM to sort them, perhaps like so:
5. "找不到更改密码的地方"
5. "Can't find where to change my password"
为收敛于一个结果所需的比赛次数随数据集大小和结果分布而变化。幸运的是,TrueSkill 对比赛次数非常稳健,而 LLM 在预测中相当稳定。
The number of matches needed for to converge on a result scales with dataset size and distribution of results. Luckily TrueSkill is very robust to the number of matches played, and LLMs are quite stable in their predictions.
我发现,对于大多数用例,每个项目甚至 1-2 场比赛就足以获得足够好的结果。对于一个包含 10,000 个项目且批次大小为 10 的列表,这可能意味着只需 500 场比赛。
I’ve found that even 1-2 matches per item can be enough to get a good enough result for most use cases. For a list of 10,000 items and a 10 item batch size, this could mean as little as 500 matches.
如果你想要更定量地衡量结果的质量,TrueSkill 会给每个项目一个评分,该评分是技能水平及其不确定性的组合。
If you want a more quantitative measure of how good your results are, TrueSkill gives each item a rating, which is a combination of its skill level and its uncertainty.
* σ (sigma):该评分的不确定性
* σ (sigma): The uncertainty in that rating
Sigma 通常初始化为 8.333,这是一个良好的起点。你可以通过观察 sigma 值的变化来衡量结果的置信度。
Sigma is usually initialized to 8.333, which is a good starting point, you can measure the confidence of your results by looking at how much the sigma values change.
当然,别忘了机器学习的基本法则:检查你的数据,以确保结果符合预期。
And of course course, remember the golden rule of Machine Learning look at your data, to make sure the results are what you expect.
一种常见的排序替代方案是单独给每个项目打分,然后按分数排序。
A common alternative to sorting is to score each item individually, and then sort by score.
如果你有一个清晰的评分函数,这是一个很好的方法,但模型难以一致地给事物打分。
This is a good approach if you have a clear scoring function, but models struggle to score things consistently.
如果你使用 1-10 的评分范围,可以预期出现许多并列和集中在像 7 这样的常见分数附近。如果你使用 1-100 的评分范围,应该预期分数有更多的随机性。
If you score on a range of 1-10, you can expect many ties and clumps around common scores like 7. If you score on a range of 1-100, you should expect more randomness in the scores.
但更重要的是,对于 LLM(以及人类!)来说,做出相对判断(“A 比 B 更有趣”)比绝对判断(“A 的兴趣指数是 73/100”)更容易。
But more importantly, it’s easier for LLMs (and humans!) to make relative judgments (“A is more interesting than B”) than absolute ones (“A is 73/100 interesting”)
鉴于嵌入的低成本,基于嵌入的排序也常被采用。
Given the cheapness of embeddings, embedding-based sorting is often deployed as well.
2. 选择参考点(例如“非常有趣”的嵌入)
2. Choose a reference point (e.g., embedding of “very interesting”)
3. 根据与参考点的余弦相似度对项目排序。
3. Sort items by cosine similarity to this reference point
虽然成本低,但基于嵌入的排序丢失了排序标准的大量细微差别,有时还会受到嵌入模型语义空间的影响(例如蓝色可能比红色更接近可用性概念)。
While cheap, embedding-based sorting loses a lot of the nuance of the sorting criteria and sometimes be affected by the semantic space of the embedding model (e.g. the color blue might be closer to the concept of usability than the color red)
尽管如此,嵌入可与 TrueSkill 结合使用:
That said, embeddings can be useful in conjunction with TrueSkill:
1. 使用嵌入进行初步粗略排序
1. Use embeddings for initial rough sorting
2. 使用嵌入相似性选择要比较的批次
2. Use embedding similarity to choose batches for comparison
3. 结合嵌入和 TrueSkill 得分进行混合排名。
3. Combine embedding and TrueSkill scores for hybrid ranking
与其随机选择批次,我们可以更有策略地选择要比较的项目。一些有效的策略包括:
Rather than selecting batches randomly, we can be strategic about which items to compare. Some effective strategies include:
1. 基于评分的分组:比较评分相似的项目以细化接近的区分
1. Rating-based grouping: Compare items with similar ratings to refine close distinctions
2. 高不确定性配对:优先比较具有高 sigma 值的项目
2. High uncertainty pairs: Prioritize comparisons between items with high sigma values
3. 桥梁比较:偶尔比较不同评分层级中的项目以验证层级边界
3. Bridge comparisons: Occasionally compare items from different rating tiers to verify tier boundaries
4. 新对最大化:优先选择未曾比较过的项目对
4. Novel pair maximization: Prioritize pairs that haven’t been compared before
sigma 值提供了内置的置信度量。我们可以利用它来:
The sigma value provides a built-in confidence measure. We can use this to:
1. 指导批次选择:优先选择 sigma 值高的项目
1. Guide batch selection: Prioritize items with high sigma values
2. 检测收敛:监控 sigma 值何时稳定
2. Detect convergence: Monitor when sigma values stabilize
3. 平衡探索与利用:在批次中混合高和低 sigma 值的项目
3. Balance exploration/exploitation: Mix high and low sigma items in batches
我遇到这个问题几次了,并向几个朋友推荐了这个解决方案,希望它对你有用。如果有用请告诉我!
I’ve run into this problem a few times now, and recommended this solution to a few friends, so I hope it’s useful for you as well. Let me know if it is!
如果你在更复杂的规模上遇到这个问题,请给我发邮件(trq212 AT gmail DOT com),我会看看是否能帮忙。
If you’re running into this problem at a more complicated scale, send me an email (trq212 AT gmail DOT com) and I’ll see if I can help.