DJbenchBack to the booth

Who reads a room.

Every model gets the same brief and the same crowd. Listeners hear two sets blind and pick one. Ratings move on those votes alone.

0 ballots counted

#ModelRating
01
Gemini 2.5 Pro
Google
1500
02
Gemini 2.5 Flash
Google
1500
03
GPT-OSS 120B
OpenAI
1500
04
DeepSeek R1
DeepSeek
1500
05
Qwen3 235B
Alibaba
1500
06
Claude Opus 5
Anthropic
1500

How the score moves

Elo, the chess rating, applied to sets. Everyone starts at 1500. Beating a highly rated model earns more than beating a struggling one, and a tie moves both toward each other. A model that has not been heard yet sits at its starting value.

What is being measured

Selection and sequencing, not production. No model made any of this music. Each was asked which real records it would play, in what order, for a specific room at a specific hour, and the audio comes from YouTube.

The roster

Gemini 2.5 Progemini-2.5-pro
Gemini 2.5 Flashgemini-2.5-flash
GPT-OSS 120Bopenai/gpt-oss-120b-maas
DeepSeek R1deepseek-ai/deepseek-r1-0528-maas
Qwen3 235Bqwen/qwen3-235b-a22b-instruct-2507-maas
Claude Opus 5claude-opus-5

Google, OpenAI, DeepSeek and Alibaba models run on Vertex AI. The Anthropic set was written by Claude Opus 5 answering the same prompt.