Top 10 LLM Models — Cost Analysis Report

Date: 2026-07-26 THIS VERSION IS FOR PUBLIC CONSUMPTION Source: 30-day OpenRouter Activity API (2026-06-27 → 2026-07-26)
Prepared by: Clark-H1 (Hermes, profile: clark-h1)


Executive Summary

Over the past 30 days, our top 10 models consumed 2.03 billion tokens at a total cost of $185.41. The average cost across all models was $0.09 per million tokens.

If we were to replace these models with their closest Anthropic equivalents (claude-opus-4.8 or claude-sonnet-4), the estimated cost would be $6,424.54 (with prompt caching) to $8,730.74 (without caching) — a 35–47× increase.

The free NVIDIA Nemotron models alone (ranked #4) would add $1,102 in cost if moved to equivalent Anthropic models.


Summary: Cost Per Million Tokens

Model Tokens (M) My Total Cost My $/M Anthropic Equivalent Anthropic Cost Anthropic $/M
deepseek/deepseek-v4-flash 630.3 $34.17 $0.05 claude-sonnet-4 $1,447.05 $2.30
poolside/laguna-s-2.1 610.4 $8.60 $0.01 claude-opus-4.8 $2,253.83 $3.69
deepseek/deepseek-v4-pro 302.2 $61.64 $0.20 claude-opus-4.8 $1,132.72 $3.75
nvidia/nemotron-3-ultra-550b-a55b 295.0 $0.00 $0.00 claude-opus-4.8 $1,102.09 $3.74
z-ai/glm-5.2 113.2 $33.81 $0.30 claude-sonnet-4 $249.93 $2.21
anthropic/claude-opus-4.8 23.1 $29.78 $1.29 — (already anthropic) $88.44 $3.82
qwen/qwen3.7-plus 20.3 $2.07 $0.10 claude-sonnet-4 $47.14 $2.32
google/gemini-3.1-pro-preview 16.5 $9.44 $0.57 claude-opus-4.8 $61.86 $3.75
moonshotai/kimi-k2.6 16.0 $5.77 $0.36 claude-sonnet-4 $37.86 $2.37
meituan/longcat-2.0 1.1 $0.14 $0.12 claude-sonnet-4 $3.62 $3.17
TOTAL / AVG 2,028.2 $185.41 $0.09 $6,424.54 $3.17

Detailed Breakdown: Token Usage & Cost Projections

# Model Prompt Tokens Completion Tokens Total Tokens My Cost Anthropic Est. (no cache) Anthropic Est. (w/ cache)
1 deepseek/deepseek-v4-flash 625,072,587 5,209,139 630,281,726 $34.17 $1,953.35 $1,447.05
2 poolside/laguna-s-2.1 609,216,269 1,207,561 610,423,830 $8.60 $3,076.27 $2,253.83
3 deepseek/deepseek-v4-pro 300,854,877 1,384,005 302,238,882 $61.64 $1,538.87 $1,132.72
4 nvidia/nemotron-3-ultra-550b-a55b 293,849,268 1,181,756 295,031,024 $0.00 $1,498.79 $1,102.09
5 z-ai/glm-5.2 113,021,357 160,956 113,182,313 $33.81 $341.48 $249.93
6 anthropic/claude-opus-4.8 22,938,040 188,655 23,126,695 $29.78 $119.41 $88.44
7 qwen/qwen3.7-plus 20,145,573 201,536 20,347,109 $2.07 $63.46 $47.14
8 google/gemini-3.1-pro-preview 16,403,508 79,633 16,483,141 $9.44 $84.01 $61.86
9 moonshotai/kimi-k2.6 15,766,712 221,938 15,988,650 $5.77 $50.63 $37.86
10 meituan/longcat-2.0 1,054,882 87,175 1,142,057 $0.14 $4.47 $3.62
TOTAL 2,018,323,073 9,922,354 2,028,245,427 $185.41 $8,730.74 $6,424.54

Anthropic Equivalency Mapping

Model Closest Anthropic Tier Rationale
deepseek/deepseek-v4-flash claude-sonnet-4 Mid-tier Fast efficient workhorse. Strong general reasoning, economical. Sonnet is the closest match for capability tier and token economics.
poolside/laguna-s-2.1 claude-opus-4.8 Premium Top-tier coding/reasoning specialist. Trades blows with Opus on SWE-bench.
deepseek/deepseek-v4-pro claude-opus-4.8 Premium Elite reasoning MoE (~1T params). Direct competitor to Opus on math/code/science.
nvidia/nemotron-3-ultra-550b-a55b claude-opus-4.8 Premium 550B dense model, strong reasoning benchmarks. Opus-tier capability. (Free tier on OpenRouter.)
z-ai/glm-5.2 claude-sonnet-4 Mid-tier Strong bilingual (zh/en) general model. Sonnet-tier on English benchmarks.
anthropic/claude-opus-4.8 Premium Already Anthropic's top model.
qwen/qwen3.7-plus claude-sonnet-4 Mid-tier Alibaba's general-purpose workhorse. Sonnet-tier on most benchmarks.
google/gemini-3.1-pro-preview claude-opus-4.8 Premium Google's premium multimodal model. Opus-tier reasoning, stronger on vision.
moonshotai/kimi-k2.6 claude-sonnet-4 Mid-tier Long-context specialist (1M+). Sonnet-tier for general tasks.
meituan/longcat-2.0 claude-sonnet-4 Mid-tier Long-context economical model. Comparable to Sonnet tier.

Note: These equivalences are based on model capability tiers and benchmark performance from training data (mid-2026). They are approximations — models differ significantly in architecture, latency, context window, and multimodal capabilities. The Anthropic models are used here as pricing reference points, not as direct functional replacements.


Pricing Assumptions

OpenRouter pricing used (as of 2026-07-26):

Model Prompt ($/M tokens) Completion ($/M tokens) Cached Prompt ($/M tokens)
anthropic/claude-opus-4.8 $5.00 $25.00 $0.50
anthropic/claude-sonnet-4 $3.00 $15.00 $0.30

Cache assumptions:

Note: The "already anthropic" line (claude-opus-4.8, #6) shows a 3× discrepancy between our actual cost ($29.78) and the estimated cost at OpenRouter list prices ($88.44). This suggests we're receiving a discounted rate or bulk pricing tier not reflected in the public API pricing.


Data Sources

Primary: OpenRouter Activity API

Model Availability: OpenRouter Models API

Anthropic Equivalency

Model capability comparisons are based on training data knowledge (mid-2026) covering published benchmarks, Chatbot Arena ratings, and model card specifications. These are directional estimates, not verified against a live leaderboard.


Methodology

  1. Data collection: Queried the OpenRouter Activity API for each of the last 30 completed UTC days
  2. Aggregation: Summed request counts, prompt tokens, completion tokens, and costs by model name across all days
  3. Filtering: Excluded models no longer listed in the OpenRouter models API
  4. Ranking: Sorted by total tokens (prompt + completion) descending
  5. Top 10: Selected the top 10 models from the ranked list
  6. Anthropic mapping: Each model assigned to its closest Anthropic capability tier (Opus or Sonnet) based on published benchmarks
  7. Cost projection: Applied OpenRouter list pricing for anthropic/claude-opus-4.8 and anthropic/claude-sonnet-4 against the actual prompt/completion token counts
  8. Cache estimate: Applied a 30% effective cache discount (50% eligibility × 60% hit rate) to prompt tokens

Notes