Boredom at Work
Live AI Model IndexBenchmarked via Artificial AnalysisLate 2026 Edition

Frontier LLM Leaderboard

Independent comparison of frontier AI models from Anthropic, OpenAI, Google DeepMind, Meta, and SpaceXAI. We benchmark the Artificial Analysis Intelligence Index, output speed (tokens/sec), first-chunk latency, and cost per task.

58 Index
Top Intelligence
Claude Opus 5.5
249 t/s
Fastest Output Speed
Gemini 3.8 Flash
$0.13
Lowest Cost per Task
MiMo-V2.6-Pro
1M Tokens
Standard Frontier Context
Opus, Sonnet, Astra, Argon
Category Leaders

Best Models by Performance Profile

👑#1 Intelligence

Claude Opus 5.5

Anthropic · 1M Context

Highest quality score on the index (58). Built for multi-layered reasoning, architectural strategy, and zero-compromise accuracy.

Intelligence: 58 · Speed: 93 t/s · $5.98/task
💻Coding Champion

Claude Sonnet 5.5

Anthropic · 1M Context

Leading developer choice worldwide. High intelligence (56) paired with 139 tokens/second for instant code gen and agentic loops.

Intelligence: 56 · Speed: 139 t/s · $2.75/task
⚡Speed King

Gemini 3.8 Flash

Google DeepMind · 1M Context

Record-shattering 249 tokens per second with 0.19s latency. The ultimate engine for real-time voice and high-throughput workflows.

Output: 249 t/s · Latency: 0.19s · $1.24/task
💰Best Value

GPT-6.1 Sol

OpenAI · 1M Context

High intelligence index of 52 at just $0.72 per task. Unbeatable price-to-performance ratio for massive production deployments.

Intelligence: 52 · Cost: $0.72/task · 63 t/s
Artificial Analysis Intelligence Index

Comprehensive Model Comparison

Filter and sort over key intelligence and speed metrics. Real data sourced from Artificial Analysis.
#1

Claude Opus 5.5

Anthropic · Peak Complex Architecture & Hard Reasoning
Proprietary
★ Top Overall Intelligence
Index
58
Speed
93 t/s
Cost
$5.98
Context
1M
The highest intelligence score on the Artificial Analysis index. Excels at intricate multi-step reasoning, architectural planning, and zero-hallucination analysis.Try →
#2

Claude Sonnet 5.5

Anthropic · Full-Stack Coding & Autonomous Terminal Loops
Proprietary
★ Best Developer & Coding Model
Index
56
Speed
139 t/s
Cost
$7.67
Context
1M
Combines near-peak frontier intelligence with blistering 139 tokens/second generation speed. The undisputed industry standard for software engineering.Try →
#3

Claude Fable 5.1

Anthropic · Narrative Synthesis, Law & Complex Prose
Proprietary
★ Creative & Nuance Specialist
Index
53
Speed
68 t/s
Cost
$7.63
Context
1M
Renowned for human-level nuance, creative problem solving, and long-form prose synthesis with a native 1 million token context.Try →
#4

GPT-6 Astra

OpenAI · Multi-Agent Systems & Deep STEM Synthesis
Proprietary
★ Flagship Multimodal Pioneer
Index
53
Speed
54 t/s
Cost
$3.26
Context
1M
OpenAI frontier flagship featuring dynamic thinking tokens, deep vision capabilities, and autonomous tool calling across multi-agent pipelines.Try →
#5

Gemini 4 Argon

Google DeepMind · Enterprise Multimodal Ingestion & Workspace Tasks
Proprietary
★ Multimodal Video & Document Titan
Index
53
Speed
45 t/s
Cost
$1.99
Context
1M
Google next-gen flagship with 53 intelligence score and strong cost efficiency ($1.99 per task). Unrivaled on hours of video and massive PDF archives.Try →
#6

GPT-6.1 Sol

OpenAI · High-Volume Production & Cost-Effective Reasoning
Proprietary
★ Best Intelligence-to-Cost Value
Index
52
Speed
63 t/s
Cost
$0.72
Context
1M
Scores 52 on the Intelligence Index while slashing cost per task to just $0.72. The ideal balance of frontier-tier reasoning and production affordability.Try →
#7

Muse Spark 1.3

Meta · High-Throughput Open Deployments & Real-Time Audio
Open Weights
★ Ultra-Fast Open Ecosystem
Index
48
Speed
152 t/s
Cost
$1.60
Context
1M
Blazing 152 tokens per second with 48 intelligence score. Available across open ecosystems with full 1M context capabilities.Try →
#8

Grok 4.7

SpaceXAI · Real-Time Telemetry & Real-World News Analysis
Proprietary
★ Live Web & Unfiltered Logic
Index
46
Speed
81 t/s
Cost
$3.74
Context
500k
High-intelligence engine trained with real-time web telemetry and unconstrained reasoning, featuring a 500k context window and 81 t/s speed.Try →
#9

MiMo-V2.6-Pro

Xiaomi · Micro-Budget Scaling & Edge Orchestration
Proprietary
★ Lowest Cost Leader
Index
46
Speed
46 t/s
Cost
$0.13
Context
1M
Remarkable efficiency reaching 46 intelligence at an astonishing $0.13 per task cost. A 10x cost reduction for high-volume enterprise pipelines.Try →
#10

Qwen3.8 Max (0902)

Alibaba · Cross-Border Trade, Asian Languages & Complex Math
Proprietary
★ Multilingual & Math Specialist
Index
45
Speed
39 t/s
Cost
$5.41
Context
984k
Top Asian frontier model with near-1M context, exceptional multilingual translation quality, and strong competition mathematics benchmarks.Try →
#11

GLM-5.3

Z AI · API Automation & Structured Workflows
Proprietary
Index
45
Speed
71 t/s
Cost
$2.01
Context
1M
Robust foundation model achieving 45 intelligence index with 1M context window and 71 t/s generation throughput.Try →
#12

Step 5 Preview

StepFun · API Automation & Structured Workflows
Proprietary
Index
44
Speed
86 t/s
Cost
$0.72
Context
1M
Robust foundation model achieving 44 intelligence index with 1M context window and 86 t/s generation throughput.Try →
#13

Kimi K3

Kimi · API Automation & Structured Workflows
Proprietary
Index
44
Speed
34 t/s
Cost
$2.00
Context
1.05M
Robust foundation model achieving 44 intelligence index with 1.05M context window and 34 t/s generation throughput.Try →
#14

GPT-5.6 Terra

OpenAI · API Automation & Structured Workflows
Proprietary
Index
42
Speed
100 t/s
Cost
$1.40
Context
1M
Robust foundation model achieving 42 intelligence index with 1M context window and 100 t/s generation throughput.Try →
#15

GLM-5.3-Flash

Z AI · API Automation & Structured Workflows
Proprietary
Index
42
Speed
54 t/s
Cost
$0.25
Context
1M
Robust foundation model achieving 42 intelligence index with 1M context window and 54 t/s generation throughput.Try →
Market Intelligence

How the Frontier Leaderboard Has Evolved

The artificial intelligence market has transitioned from static benchmark benchmarks to holistic, real-world utility measured across intelligence, execution speed, latency, and operational cost:

  • Anthropic's Dual-Flagship Strategy: Claude Opus 5.5 commands the highest raw intelligence score (58), excelling at complex system architecture and high-stakes reasoning. Meanwhile, Claude Sonnet 5.5 (56 score at 139 tokens/sec) dominates active developer environments like Cursor and VS Code Copilot.
  • OpenAI's Tiered Astra & Sol Models: GPT-6 Astra acts as the reasoning flagship, while GPT-6.1 Sol redefines price-performance by delivering 52 intelligence for just $0.72 per task.
  • Google DeepMind's Speed Supremacy: Gemini 3.8 Flash achieves a historic 249 tokens per second with sub-200ms latency, enabling real-time voice and high-throughput enterprise pipelines. Gemini 4 Argon brings heavy multimodal capability to massive document sets.
  • Global Cost Disruptors: Xiaomi's MiMo-V2.6-Pro delivers 46 intelligence at just $0.13 per task, demonstrating how rapid distillation and quantization are driving token costs to zero.

Understanding the Artificial Analysis Metrics

Intelligence Index

Standardized composite score (0–100) combining reasoning, software coding, mathematical deduction, and factual comprehension across non-memorizable evaluations.

Output Speed (Tokens/s)

Median token generation rate after the first response packet is returned. Speeds above 100 t/s provide an instant, seamless typing experience for interactive human-in-the-loop workflows.

Latency (Time To First Token)

The duration in seconds before the model generates its initial response chunk. Sub-0.5s latency is essential for voice assistants and automated terminal tool feedback loops.

Cost per Task (USD)

Calculates real expenditure across standardized production queries. Comparing $0.13 vs $5.98 highlights how choosing the right tier slashes enterprise operational costs by 95%.

Independent Verification

Multi-Source Benchmarking & Data Privacy

To eliminate provider bias, our index synthesizes intelligence telemetry and evaluation methodologies across three leading research hubs:

Artificial Analysis

Automated latency testing, median token throughput (Tokens/s), and standardized cost-per-task evaluations across global API endpoints.

LMSYS Chatbot Arena

Crowdsourced, double-blind human preference battles measuring real-world conversational helpfulness and subjective nuance (arena.ai).

Collective Brain

European business workflow testing, daily task indexing, and GDPR / AVV data privacy evaluations for enterprise deployment.

Knowledge Base

Frequently Asked Questions

What is the Artificial Analysis Intelligence Index?

The Artificial Analysis Intelligence Index is an aggregate quality score (0 to 100) evaluating foundation models across challenging, non-memorizable benchmarks including complex reasoning, coding benchmarks, and multi-turn instruction following.

Which AI model is #1 on the leaderboard right now?

Claude Opus 5.5 from Anthropic holds the #1 spot with an Intelligence Index of 58, followed closely by Claude Sonnet 5.5 (56), Claude Fable 5.1 (53), GPT-6 Astra (53), and Gemini 4 Argon (53).

Which model is the best for software development and coding?

Claude Sonnet 5.5 is the top choice for developers. With an Intelligence Index of 56 and a generation speed of 139 tokens per second, it offers unmatched precision in repository-scale code refactoring, terminal agent execution, and debugging.

Which AI model is the fastest?

Gemini 3.8 Flash from Google DeepMind is the speed champion, delivering an incredible 249 tokens per second with a Time-To-First-Token (TTFT) latency under 0.20 seconds, making it ideal for real-time voice, copilot completion, and interactive chat.

How are Cost per Task and API token prices calculated?

Cost per Task measures the standardized price (in USD) required to complete a complex real-world query sequence across input tokens, reasoning/thinking tokens, and generated output tokens. Models range from $0.13 (MiMo-V2.6-Pro) to $5.98 (Claude Opus 5.5).

Does kileaderboard.de point to this benchmark page?

Yes, kileaderboard.de and boredom-at-work.com/ki-leaderboard/ route directly to this canonical benchmark index, giving tech leaders, engineers, and creators transparent and verified AI intelligence metrics.

Master Frontier AI in Your Workflow

Learn how to put these top-ranking models to work with our step-by-step master guides and prompt frameworks.