LLM-Powered Sorting with TrueSkill

Thariq Shihipar · Anthropic · 2025-02-11

Read on thariq.io

A practical engineering post that solves a specific, common, and unobvious problem: how to sort a large list of items by some semantic quality (interestingness, quality, relevance) when LLMs are bad at ranking long lists directly. Shihipar's answer combines two things you wouldn't normally combine — Microsoft's TrueSkill rating system (the algorithm behind Xbox matchmaking) and pairwise LLM judgments.

The core insight: LLMs are good at relative judgments ("A is more interesting than B") and bad at absolute ones ("rate A from 1 to 10"). And they get worse the longer the list. So you treat each item as a player, each comparison as a match, and let TrueSkill's Bayesian rating algorithm aggregate confidence-weighted scores across many small pairwise comparisons. Each item only needs 1-2 comparisons; the system gives you both a global ranking and a confidence interval (the σ value) for free.

- It's batteries-included: TrueSkill is decades old, well-tested, and gives you uncertainty quantification automatically - It scales: pairwise comparisons in small batches don't run into context limits or position bias - It's robust to LLM noise: Bayesian aggregation absorbs individual judgment errors

> It's easier for LLMs (and humans!) to make relative judgments than absolute ones.

This generalizes well past sorting. Anywhere you'd ask an LLM for a numerical score, ask yourself if pairwise comparison + a rating algorithm would be more robust. (Often yes.)

Most "tips for working with LLMs" content is shallow. This is the opposite: one specific technique, properly explained, with the math and the failure modes, by someone who actually shipped it. Engineers building eval pipelines, recommendation systems, or content-quality scoring will steal this directly.