Threads
ai-toolsai-engineering
Original source

Claude Opus 4 LMArena Benchmark Scores

Summary

@purple.soy posted LMArena (formerly LMSYS Chatbot Arena) benchmark scores for anthropic/claude-opus-4-cumulative, calling it “the best model ever made.”

Scores:

  • Overall Arena Score: 84.7
  • IF Eval (Instruction Following): 88.2
  • Style Control: 83.8
  • Creative Writing, editing: 80.7
  • Math: 77.7
  • Coding: 73.2
  • Intelligence (Hard prompts): 67.4
  • Long Context: 55.4

Notes:

  • Claude Opus 4 is a paid, closed-weight proprietary model from Anthropic
  • Available via Anthropic API, AWS Bedrock, Google Vertex AI
  • The “cumulative” suffix refers to LMArena’s cumulative vote tracking methodology
  • Strong in instruction following and style control; weaker in long context and hard intelligence benchmarks