Final Bradley-Terry leaderboards from the CLEF 2026 JOKER lab shared task.
Rankings are fitted on ≈40,000 pairwise battles judged by a 3-judge cross-provider LLM
ensemble (gpt-oss-20b · deepseek-chat · gemini-2.5-flash-lite) under the constraint-aware
pair-pun-v1 prompt, majority-aggregated, with 95% bootstrap confidence intervals.