Pairwise LLM judges (A/B/tie): budget-aware multi-turn packing, position-bias correction, pseudo-label distillation. Generalized from the 4th-place (gold) solution to Kaggle LMSYS Chatbot Arena.
-
Updated
Jul 24, 2026 - Python
Pairwise LLM judges (A/B/tie): budget-aware multi-turn packing, position-bias correction, pseudo-label distillation. Generalized from the 4th-place (gold) solution to Kaggle LMSYS Chatbot Arena.
📊 Daily auto-updated snapshots of all Arena AI (LMSYS Chatbot Arena) leaderboards — LLM, Vision, Code, Video, Image & more. Structured JSON with historical tracking.
Rank LLM APIs by cost-effectiveness against Arena ELO scores · 按 Arena ELO 性价比对 LLM API 实时定价排名
30 conversational LLM datasets (~7.7M rows) normalized to one unified schema and published as a single HuggingFace dataset with per-source configs.
LLM evaluation arena with blind side-by-side comparison, ELO ratings, and category leaderboards. Like LMSYS Chatbot Arena, but self-hosted.
Add a description, image, and links to the chatbot-arena topic page so that developers can more easily learn about it.
To associate your repository with the chatbot-arena topic, visit your repo's landing page and select "manage topics."