• nednobbins@lemm.ee
    link
    fedilink
    English
    arrow-up
    2
    ·
    2 days ago

    I wouldn’t either but that’s exactly what lmsys.org found.

    That blog post had ratings between 858 and 1169. Those are slightly higher than the average rating of human users on popular chess sites. Their latest leaderboard shows them doing even better.

    https://lmarena.ai/leaderboard has one of the Gemini models with a rating of 1470. That’s pretty good.