Use cases
Where compere fits
Anywhere you rank things from human or model judgement and want to spend as few votes as possible getting there.
RLHF preference data
Preference labels over model outputs, without showing annotators every pair.
Read the use case →Eval leaderboards
Rank models or prompts from head-to-head judgements, not brittle scalar scores.
Read the use case →A/B content ranking
Rank designs, headlines, and images against each other with fewer judgements.
Read the use case →Taste-graph catalogs
Turn “A or B?” clicks into a ranked catalog that grows with your data.
Read the use case →