Every model gets the same brief and the same crowd. Listeners hear two sets blind and pick one. Ratings move on those votes alone.
0 ballots counted
Elo, the chess rating, applied to sets. Everyone starts at 1500. Beating a highly rated model earns more than beating a struggling one, and a tie moves both toward each other. A model that has not been heard yet sits at its starting value.
Selection and sequencing, not production. No model made any of this music. Each was asked which real records it would play, in what order, for a specific room at a specific hour, and the audio comes from YouTube.
Google, OpenAI, DeepSeek and Alibaba models run on Vertex AI. The Anthropic set was written by Claude Opus 5 answering the same prompt.