AAA.win tested 18 AI agents on 20 multilingual business tasks across 4 languages. This preview report is generated from structured run data and should be edited before publication.
Report map
English Benchmarks Are Not Enough
AAA.win tested 18 AI agents on 20 multilingual business tasks across 4 languages. This preview report is generated from structured run data and should be edited before publication.
18agents20tasks1080runs
18agents
20tasks
4languages
1080runs
Executive Summary
Overall winner: Claude Main with an average score of 87.
Lowest critical-failure rate among top agents: Qwen Main.
Most common failure mode: literal_translation.
Strong overall performance does not imply winning every language or task type.
Available Reports
For decision makers
Multilingual Executive Brief
A fast read on winners, caveats, and signals that need human review before public claims.
Read the overall ranking together with language winners.
Preview seed results should not be quoted as final benchmark truth.
Business-safety risk matters more than fluent style.
For regional teams
Language Market Report
Compares agents across Chinese, English, Japanese, and Spanish markets so teams do not choose from English-only evidence.
Use language winners for local production workflows.
Review tone, date formats, and support conventions by market.
French, German, Portuguese, and Korean should become real task sets next.
For operations
Risk and Failure Report
Focuses on critical failures, unsafe promises, invented fields, and unusable outputs.
Treat failure tags as audit leads.
Human-review refund, security, and compliance cases.
Do not let a high score hide weak format discipline.
For tool choice
Buyer Selection Report
Helps teams choose by cost, working language, and risk tolerance instead of a single average.
Premium agents can be justified on high-risk workflows.
Standard agents remain competitive in some languages and extraction tasks.
The best choice depends on the workflow, not only the rank.
For product teams
Task Family Report
Explains which task families separate agents most: support, writing, and structured extraction.
Support tests business boundaries.
Writing tests natural tone and localization.
Extraction tests JSON, dates, missing fields, and robustness.
For public readers
Publication Readiness Report
Lists the conditions required before using results in a launch, article, or commercial page.
Replace seed outputs with real, verifiable model outputs.
Publish model versions and evaluation dates.
State clearly that vendors cannot buy score changes.
How to use
Turn reports into a selection workflow
Use the global ranking to shortlist, then inspect language, task type, failure tags, and cost before piloting. High-risk workflows must return to task samples and human-review rules.
1
Choose the workflow
Support, writing, extraction, internal automation, or cross-border work.
2
Inspect evidence
Read task pages, failure tags, format pass, and critical-failure rate.
3
Pilot with real samples
Retest with your own inputs and record human repair time.
Update Plan for June 28, 2026
Today's realistic scope is to improve evidence quality and localization depth across the existing 20-task arena, while keeping preview results clearly labeled.
Morning: localize the report page and verify every language route.
Midday: enrich 8-10 task pages with clearer prompt summaries, rubrics, and risk notes.
Afternoon: review the remaining 10-12 task pages for wording consistency and broken links.
Evening: regenerate the report, rerun checks, and deploy one stable release.
Expected capacity today: 20 task pages can be reviewed for text quality; 8-10 can receive deeper evidence edits if each task gets careful prompt/rubric treatment.
Overall Leaderboard
Rank
Agent
Score
Pass Rate
Critical Failure Rate
Cost Tier
1
Claude Main
87
97%
12%
premium
2
OpenAI Main
86
92%
12%
premium
3
Qwen Main
84
93%
10%
standard
4
Kimi Main
82
87%
12%
standard
5
GLM Main
81
88%
7%
standard
6
Mistral Main
81
85%
2%
standard
7
MiniMax Main
80
90%
2%
standard
8
Gemini Main
80
82%
12%
standard
9
Cohere Main
80
77%
12%
standard
10
DeepSeek Main
80
70%
7%
low
11
Hunyuan Main
80
85%
2%
standard
12
ERNIE Main
79
78%
3%
standard
13
Llama Main
79
75%
7%
low
14
Perplexity Main
79
73%
17%
standard
15
Doubao Main
79
67%
12%
low
16
Yi Main
78
62%
8%
low
17
Grok Main
75
37%
27%
standard
18
Phi Main
75
37%
20%
low
Language Winners
Language
Winner
Score
Critical Failure Rate
中文
Qwen Main
89
7%
English
OpenAI Main
93
7%
日本語
Claude Main
89
13%
Español
Claude Main
88
13%
Task Type Winners
Task Type
Winner
Score
Critical Failure Rate
Support
Claude Main
90
13%
Writing
Claude Main
90
11%
Extraction
Qwen Main
88
6%
Failure Modes
Failure Tag
Count
literal_translation
84
unsupported_claim
84
weak_cta
46
unsafe_refund_promise
41
missing_field
38
too_verbose
36
invalid_json
31
missed_dependency
30
generic_ai_copy
24
wrong_date_format
19
Task Results
Task
Language
Type
Winner
Score
Primary Risk
Chinese Customer Complaint Triage
中文
Support
Qwen Main
85
unsafe_refund_promise
Chinese App Review Pain Point Summary
中文
Writing
Kimi Main
92
hallucinated_issue
Chinese Contract Field Extraction
中文
Extraction
Qwen Main
96
hallucinated_signing_date
Chinese Sales Call Summary
中文
Extraction
Qwen Main
96
missed_buying_signal
Chinese Invoice Dispute Reply
中文
Support
OpenAI Main
85
unauthorized_credit
SaaS Landing Page Hero Rewrite
English
Writing
OpenAI Main
93
generic_ai_copy
Meeting Notes Action Item Extraction
English
Extraction
OpenAI Main
89
discussion_as_action
Refund Policy Boundary Reply
English
Support
OpenAI Main
96
unsafe_refund_promise
English Security Questionnaire Answer
English
Support
OpenAI Main
96
unsupported_security_claim
English Churn Risk Email
English
Writing
Claude Main
95
tone_deaf_retention
Japanese Business Email Politeness Rewrite
日本語
Writing
OpenAI Main
85
unnatural_japanese
Japanese Appointment Intent Classification
日本語
Support
Claude Main
92
wrong_intent
Japanese Product Specification Extraction
日本語
Extraction
Qwen Main
91
hallucinated_material
Japanese Support Escalation Note
日本語
Support
Claude Main
92
lost_escalation_context
Japanese Pricing Page Localization
日本語
Writing
Claude Main
92
literal_pricing_copy
Spanish Support Reply for Wrong Item
Español
Support
Claude Main
89
unsafe_refund_promise
Spanish Ad Headline Localization
Español
Writing
Claude Main
92
literal_translation
Spanish Order Confirmation Extraction
Español
Extraction
Claude Main
85
wrong_date_format
Spanish Billing Cancellation Reply
Español
Support
Claude Main
91
wrong_cancellation_policy
Spanish Survey Insight Clustering
Español
Extraction
Qwen Main
83
overmerged_feedback
Methodology Snapshot
Runs per task-agent pair: 3
Tools enabled: false
Web browsing enabled: false
Memory enabled: false
Scores are computed from five dimensions: task success, language fit, instruction following, business safety, and output reliability.
Publication Notes
Replace preview seed outputs with exact model outputs before public claims.
Human-review all critical business-safety failures.
Confirm model versions, pricing dates, and evaluation dates.
Keep the vendor policy visible: vendors cannot pay to change scores.
Subscribe
Get agent-selection insights, failures, and evaluation changes
Evidence-led updates only. Confirm by email after submitting; unsubscribe anytime.