Best LLMs for Writing in 2026

Aggregated benchmark data across EQ-Bench Creative Writing, LMArena Text, and Artificial Analysis — covering 34 models, updated weekly.

Last updated:  ·  34 models tracked  ·  3 tiers: Premium · Mid-Range · Budget

The short answer

Updated July 2026 To write well, you also need a voice.md

The best LLM for writing right now is Kimi K3 (Moonshot AI), which tops the EQ-Bench Creative Writing leaderboard ahead of Claude Fable 5 and Claude Opus 4.7. For the best quality per dollar, Muse Spark 1.1 (Meta) delivers near-frontier writing at a fraction of the cost. Full ranking across 34 models below.

  1. 1 Kimi K3 EQ Creative 2377
  2. 2 Claude Fable 5 EQ Creative 2091
  3. 3 Claude Opus 4.7 EQ Creative 2047

How we rank models for writing

This ranking combines three independent data sources to give the most complete picture of writing quality across frontier LLMs. No single benchmark captures the full picture — so we aggregate:

Show benchmark details
  • EQ Creative EQ-Bench Creative Writing — specialist benchmark using trained raters to assess narrative quality, emotional depth, prose style, and character voice. Elo scale ~1300–2380. The most relevant signal for marketing copy, long-form content, and creative work.
  • Arena Text LMArena Text — crowd-sourced human preference leaderboard. Broad signal across all text tasks: a model that consistently wins votes is generally pleasant, clear, and useful to read. Elo scale ~1346–1505.
  • EQ General EQ-Bench General — measures emotional intelligence in roleplay scenarios. A proxy for character voice quality and tonal control — useful for brand voice work. Note: high EQ-General does not automatically mean strong creative writing; interpret alongside EQ Creative.
  • Speed Artificial Analysis — median output tokens per second across providers. Matters for iterative draft workflows where waiting costs time. ~75 tokens ≈ 55 words.

Prices are per 1M tokens (input / output) and reflect standard API pricing. A dash (—) means the model has not yet appeared on that leaderboard — never an estimated or interpolated value.

The best AI models for copywriting

Updated weekly  ·  Jul 20, 2026
Model EQ Creative Arena Text EQ General Speed Price / 1M
Kimi K3 Premium
Moonshot AI 📋 Consensus
2,377 1,486
$3.00 $15.00
Claude Fable 5 Premium
Anthropic 📋 Consensus
2,090.8 1,507 2,049.7
$10.00 $50.00
Claude Opus 4.7 Premium
Anthropic 📋 Consensus
2,047 1,503 1,884.3
$5.00 $25.00
Muse Spark 1.1 Mid-Range
Meta 📋 Consensus
2,006.8 1,493
$1.25 $4.25
GPT-5.5 Premium
OpenAI 📋 Consensus
1,931.7 1,475 1,576.9
$5.00 $30.00
GPT-5.4 Premium
OpenAI 📋 Consensus
1,913.1 1,466 1,562.9
Claude Opus 4.8 Premium
Anthropic 📋 Consensus
1,910.5 1,483 2,029.8
$5.00 $25.00
Claude Opus 4.6 Premium
Anthropic 📋 Consensus
1,889.5 1,504 1,717.4 45 t/s
$5.00 $25.00
Claude Sonnet 4.6 Premium
Anthropic 📋 Consensus
1,877.9 1,471 1,714.1 50 t/s
$3.00 $15.00
GLM-5.2 Mid-Range
Zhipu AI 🧠 EQ-Bench
1,741.5 1,575.1
$1.40 $4.40
Claude Sonnet 4.5 Premium
Anthropic 📋 Consensus
1,738.3 1,456 1,511
$3.00 $15.00
Claude Opus 4.5 Premium
Anthropic 📋 Consensus
1,731.9 1,473 1,545.7
$5.00 $25.00
O3 Mid-Range
OpenAI 📋 Consensus
1,730.7 1,431 1,500
$2.00 $8.00
Kimi K2.6 Mid-Range
Moonshot AI 📋 Consensus
1,721.7 1,461 1,561.2
$0.60 $2.50
GPT-5.3 Chat Mid-Range
OpenAI 🧠 EQ-Bench
1,718.2 1,393
Kimi K2 Mid-Range
Moonshot AI 🧠 EQ-Bench
1,662.9 1,562.2 44 t/s
$0.55 $2.20
GPT-5.2 Premium
OpenAI 📋 Consensus
1,650.2 1,435 1,558.8
$1.25 $10.00
GLM-5 Budget
Zhipu AI 📋 Consensus
1,617 1,457 1,526 80 t/s
$0.80 $2.50
Horizon Alpha Mid-Range
Unknown 🧠 EQ-Bench
1,613.9 1,554.4
GLM-5.1 Mid-Range
Zhipu AI 📋 Consensus
1,595.2 1,471 1,566.5
Claude Opus 4 Premium
Anthropic 📋 Consensus
1,578.2 1,424 1,399.8
$15.00 $75.00
Kimi K2.5 Mid-Range
Moonshot AI 📋 Consensus
1,570.4 1,450 1,544.7
Mistral Medium 3 Budget
Mistral AI 🧠 EQ-Bench
1,489.5
$0.40 $2.00
Gemini 3 Pro Mid-Range
Google 📋 Consensus
1,480.2 1,486 1,559.4 80 t/s
$2.00 $12.00
DeepSeek V3.2 Budget
DeepSeek 📋 Consensus
1,480 1,425
$0.28 $0.42
Qwen3-235B Budget
Alibaba 📋 Consensus
1,459 1,375 1,213.6
$0.18 $0.54
Gemini 3.1 Pro Mid-Range
Google 📋 Consensus
1,456.3 1,485 1,537.7
$2.11 $12.66
GPT-4o Premium
OpenAI 📋 Consensus
1,443 1,346 1,393 185 t/s
$2.50 $10.00
GLM-4.7 Budget
Zhipu AI 📋 Consensus
1,395.6 1,442 1,428.3
$0.38 $1.70
MiniMax M2.5 Budget
MiniMax 📋 Consensus
1,393.9 1,391 395 t/s
$0.30 $1.20
Gemini 3 Flash Mid-Range
Google 🏟️ Arena
1,473 250 t/s
$0.50 $3.00
Gemini 3.1 Flash-Lite Budget
Google 🏟️ Arena
1,432
$0.25 $1.50
Grok 4.1 Mid-Range
xAI 🏟️ Arena
1,466 163 t/s
$0.20 $0.50
MiMo-V2.5 Budget
Xiaomi 🏟️ Arena
1,432
$0.40 $2.00

← Scroll to see all columns →

EQ Creative & EQ General: EQ-Bench  ·  Arena Text: LMArena  ·  Speed: Artificial Analysis  ·  Prices per 1M tokens  ·  — = not yet on leaderboard  ·  Click any row for sources

The right model depends on the task

Benchmark leaderboards rank models globally — but the best model for a 2,000-word thought leadership article is not necessarily the best model for a 15-word social media headline. Here's how the leading models split across common writing tasks:

Narrative & long-form

Thought leadership, case studies, email newsletters, ghostwriting. Requires emotional depth, tonal consistency, and the ability to sustain voice across thousands of words.

Best picks: Claude Opus 4.7 · Claude Sonnet 4.6

Structured commercial copy

Product descriptions, landing pages, ad copy, LinkedIn posts. Requires clarity, persuasion structure, and format adherence more than creative flair.

Best picks: GPT-5.5 · Claude Sonnet 4.6

High-volume / fast drafts

Social media scheduling, meta descriptions, bulk content variation. Speed and cost matter more than peak quality; fast iteration wins here.

Best picks: Gemini 3 Flash · Grok 4.1 · Kimi K2

Brand voice & consistency

Any content where staying on-brand is non-negotiable. Requires strong instruction-following, tonal control, and memory of brand guidelines.

Best picks: Claude Opus 4.8 · Claude Sonnet 4.6

Managing this by hand means juggling several API keys, pricing tiers, and a decision tree for every task type. The table above shows where each model wins, so you can match the model to the job instead of forcing one model onto everything.

You picked the model. That's only half the battle.

The model sets the ceiling on quality. What it can't decide is how the writing actually sounds: your voice, your angle, your point of view. That's why output from even the best LLMs comes out generic, gets flagged as AI, and quietly loses reach. A voice.md file is the other half. It captures how you write, so any model turns your ideas into content that's distinct, human, and recognizably yours.

Frequently asked questions

Which LLM is best for creative writing in 2026?

Moonshot AI's new Kimi K3 leads the EQ-Bench Creative Writing leaderboard at an Elo of 2377, with Claude Fable 5 (Anthropic) second at 2091 and Claude Opus 4.7 third at 2047 as of July 2026. Kimi K3 tops the creative benchmark, but Claude Fable 5 still leads the broader boards: LMArena Text at 1507 and EQ General at 2050, where Claude Opus 4.8 is a close second at 2030. These models excel at narrative quality, emotional depth, and character voice — the core skills that separate great writing from generic AI output.

What is EQ-Bench and why does it matter for writing?

EQ-Bench is an independent benchmark that evaluates large language models on emotional intelligence and narrative quality, using a panel of human raters. Its Creative Writing sub-leaderboard specifically measures story quality, emotional resonance, and prose style — making it the most relevant benchmark for marketing copy, long-form content, and creative work. Scores are on an Elo scale where higher is better, typically ranging from ~1300 to ~2380.

What is LMArena Text and how is it different from EQ-Bench?

LMArena Text (formerly LMSYS Chatbot Arena) measures human preference through head-to-head votes: two anonymous models answer the same prompt, and users pick the better response. It's a broad preference signal across all text tasks, not just writing. EQ-Bench Creative Writing is narrower and more specialist — it specifically evaluates narrative and emotional writing quality with trained raters rather than crowd votes.

Which LLM is the best value for writing tasks?

Meta's Muse Spark 1.1 is the standout value pick: an EQ-Bench Creative score of 2007 at just $1.25 input / $4.25 output per 1M tokens — near-frontier writing for a fraction of the premium-tier price. GLM-5.2 (Zhipu AI) is another strong performance-per-dollar option at 1742 Creative for $1.40 input / $4.40 output — roughly 3.5× cheaper than Claude Opus 4.7 with ~85% of its creative writing performance. Kimi K2 by Moonshot AI is the deepest budget pick at 1663 Creative for just $0.55 input / $2.20 output, and Kimi K2.6 scores even higher at 1722 Creative. GLM-5 (Zhipu AI) is another strong value option at $0.80/$2.50 with scores of 1617 EQ Creative and 1526 EQ General.

How often is this ranking updated?

Scores are updated weekly via an automated scraper that fetches the latest data from EQ-Bench and LMArena. Prices are reviewed manually and updated when providers announce changes. The 'Updated weekly' badge in the table header shows the date of the last successful update.

What does 'tokens per second' mean for writing?

Tokens per second (t/s) measures how fast a model outputs text — roughly, 75 tokens equals about 55 words. For writing workflows, speed matters when you need rapid iteration on drafts or real-time dictation-to-copy conversion. MiniMax M2.5 is the fastest tracked model at 395 t/s; Gemini 3 Flash at 250 t/s offers the best speed-to-cost ratio among paid models.

Does the best LLM for writing change depending on the task?

Yes — significantly. Kimi K3 tops EQ Creative (2377) with Claude Fable 5 right behind (2091), while Fable 5 also leads EQ General (2050) — making it especially strong for character voice and emotional tone, with Claude Opus 4.8 close behind at 2030. Claude Sonnet 4.6 remains excellent for structured commercial copy where format consistency matters at a lower price. GPT-5.5 is a strong contender for creative work. Faster models like Gemini 3 Flash or Grok 4.1 suit high-volume, lower-stakes content.

Does picking the best LLM guarantee good writing?

No. The model sets the ceiling on quality, but it does not decide how the writing sounds: your voice, your angle, your point of view. Output from even the top-ranked models often reads as generic and gets flagged as AI, which costs reach. The fix is to give the model your voice. A short voice.md file captures how you write, so any model produces content that is distinct and recognizably yours.