Skip to content
AAA.win
Arena #2

English Benchmarks Are Not Enough

Generated from batch: maa-preview-002

AAA.win tested 18 AI agents on 20 multilingual business tasks across 4 languages. This preview report is generated from structured run data and should be edited before publication.

Report map

English Benchmarks Are Not Enough

AAA.win tested 18 AI agents on 20 multilingual business tasks across 4 languages. This preview report is generated from structured run data and should be edited before publication.

18agents20tasks1080runs
18agents
20tasks
4languages
1080runs

Executive Summary

Available Reports

For decision makers

Multilingual Executive Brief

A fast read on winners, caveats, and signals that need human review before public claims.

  • Read the overall ranking together with language winners.
  • Preview seed results should not be quoted as final benchmark truth.
  • Business-safety risk matters more than fluent style.
For regional teams

Language Market Report

Compares agents across Chinese, English, Japanese, and Spanish markets so teams do not choose from English-only evidence.

  • Use language winners for local production workflows.
  • Review tone, date formats, and support conventions by market.
  • French, German, Portuguese, and Korean should become real task sets next.
For operations

Risk and Failure Report

Focuses on critical failures, unsafe promises, invented fields, and unusable outputs.

  • Treat failure tags as audit leads.
  • Human-review refund, security, and compliance cases.
  • Do not let a high score hide weak format discipline.
For tool choice

Buyer Selection Report

Helps teams choose by cost, working language, and risk tolerance instead of a single average.

  • Premium agents can be justified on high-risk workflows.
  • Standard agents remain competitive in some languages and extraction tasks.
  • The best choice depends on the workflow, not only the rank.
For product teams

Task Family Report

Explains which task families separate agents most: support, writing, and structured extraction.

  • Support tests business boundaries.
  • Writing tests natural tone and localization.
  • Extraction tests JSON, dates, missing fields, and robustness.
For public readers

Publication Readiness Report

Lists the conditions required before using results in a launch, article, or commercial page.

  • Replace seed outputs with real, verifiable model outputs.
  • Publish model versions and evaluation dates.
  • State clearly that vendors cannot buy score changes.
How to use

Turn reports into a selection workflow

Use the global ranking to shortlist, then inspect language, task type, failure tags, and cost before piloting. High-risk workflows must return to task samples and human-review rules.

1

Choose the workflow

Support, writing, extraction, internal automation, or cross-border work.

2

Inspect evidence

Read task pages, failure tags, format pass, and critical-failure rate.

3

Pilot with real samples

Retest with your own inputs and record human repair time.

Update Plan for June 28, 2026

Today's realistic scope is to improve evidence quality and localization depth across the existing 20-task arena, while keeping preview results clearly labeled.

Expected capacity today: 20 task pages can be reviewed for text quality; 8-10 can receive deeper evidence edits if each task gets careful prompt/rubric treatment.

Overall Leaderboard

RankAgentScorePass RateCritical Failure RateCost Tier
1Claude Main8797%12%premium
2OpenAI Main8692%12%premium
3Qwen Main8493%10%standard
4Kimi Main8287%12%standard
5GLM Main8188%7%standard
6Mistral Main8185%2%standard
7MiniMax Main8090%2%standard
8Gemini Main8082%12%standard
9Cohere Main8077%12%standard
10DeepSeek Main8070%7%low
11Hunyuan Main8085%2%standard
12ERNIE Main7978%3%standard
13Llama Main7975%7%low
14Perplexity Main7973%17%standard
15Doubao Main7967%12%low
16Yi Main7862%8%low
17Grok Main7537%27%standard
18Phi Main7537%20%low

Language Winners

LanguageWinnerScoreCritical Failure Rate
中文Qwen Main897%
EnglishOpenAI Main937%
日本語Claude Main8913%
EspañolClaude Main8813%

Task Type Winners

Task TypeWinnerScoreCritical Failure Rate
SupportClaude Main9013%
WritingClaude Main9011%
ExtractionQwen Main886%

Failure Modes

Failure TagCount
literal_translation84
unsupported_claim84
weak_cta46
unsafe_refund_promise41
missing_field38
too_verbose36
invalid_json31
missed_dependency30
generic_ai_copy24
wrong_date_format19

Task Results

TaskLanguageTypeWinnerScorePrimary Risk
Chinese Customer Complaint Triage中文SupportQwen Main85unsafe_refund_promise
Chinese App Review Pain Point Summary中文WritingKimi Main92hallucinated_issue
Chinese Contract Field Extraction中文ExtractionQwen Main96hallucinated_signing_date
Chinese Sales Call Summary中文ExtractionQwen Main96missed_buying_signal
Chinese Invoice Dispute Reply中文SupportOpenAI Main85unauthorized_credit
SaaS Landing Page Hero RewriteEnglishWritingOpenAI Main93generic_ai_copy
Meeting Notes Action Item ExtractionEnglishExtractionOpenAI Main89discussion_as_action
Refund Policy Boundary ReplyEnglishSupportOpenAI Main96unsafe_refund_promise
English Security Questionnaire AnswerEnglishSupportOpenAI Main96unsupported_security_claim
English Churn Risk EmailEnglishWritingClaude Main95tone_deaf_retention
Japanese Business Email Politeness Rewrite日本語WritingOpenAI Main85unnatural_japanese
Japanese Appointment Intent Classification日本語SupportClaude Main92wrong_intent
Japanese Product Specification Extraction日本語ExtractionQwen Main91hallucinated_material
Japanese Support Escalation Note日本語SupportClaude Main92lost_escalation_context
Japanese Pricing Page Localization日本語WritingClaude Main92literal_pricing_copy
Spanish Support Reply for Wrong ItemEspañolSupportClaude Main89unsafe_refund_promise
Spanish Ad Headline LocalizationEspañolWritingClaude Main92literal_translation
Spanish Order Confirmation ExtractionEspañolExtractionClaude Main85wrong_date_format
Spanish Billing Cancellation ReplyEspañolSupportClaude Main91wrong_cancellation_policy
Spanish Survey Insight ClusteringEspañolExtractionQwen Main83overmerged_feedback

Methodology Snapshot

Publication Notes

Subscribe

Get agent-selection insights, failures, and evaluation changes

Evidence-led updates only. Confirm by email after submitting; unsubscribe anytime.

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases