Skip to content
AAA.win

Qwen Main

Strong Chinese business language and structured extraction.

AlibabastandardArena #2
Preview evidence60 runs
Provider
Alibaba
Exact model version
Not measured
Evaluation batch
maa-preview-002
Test date
2026-06-27
Context window
Not measured
Tool mode
disabled
Pricing source
Not measured
Latency
Not measured

Best fit

  • 中文 workflows where the agent has its strongest language evidence.
  • Extraction tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.

Not a fit for

  • Critical-failure rate is 10%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.

Alternatives

Overall84
Pass rate93%
Critical10%
Format pass100%
Measured avg. cost$0.0118
LatencyNot measured
Recommended for

Where it should be considered first

  • 中文 workflows, especially where this agent shows its strongest language score.
  • Extraction tasks where it is a reasonable first shortlist candidate.
  • Standard-budget business automation candidates.
Use caution

Review before production

  • Review failure tags such as literal_translation, unnatural_japanese, unauthorized_credit before production use.
  • Critical-failure rate is 10%; do not rely on the overall score alone for support, refund, or compliance workflows.
  • Format pass rate is 100%; retest JSON and missing-field behavior before extraction use cases go live.

Profile metrics

Overall score: 84 Win rate: 25% Pass rate: 93% Critical: 10% Format pass rate: 100% Average run cost: $0.0118

Common failure tags

literal_translationunnatural_japaneseunauthorized_credit

Language performance

中文89
English81
日本語83
Español81

Task type performance

Support82
Writing82
Extraction88
Decision profile

How should this agent enter a real pilot?

Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.

Test first
  • 中文 workflows where the agent has its strongest language evidence.
  • Extraction tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.
Do not launch first
  • Critical-failure rate is 10%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.
Pilot plan
  • Prepare 20-50 real samples across normal, edge, and high-risk cases.
  • Compare against one high-score candidate and one lower-cost candidate.
  • Record repair time, failure tags, and unacceptable failures before expanding automation.

Recommended pilot path

  1. Use 20–50 real samples across normal, edge, and high-risk inputs.
  2. Keep at least one lower-cost comparator.
  3. Human-review every critical failure and high-risk output.
  4. Retest after model, prompt, tool, or business-policy changes.
Build a full pilot plan
Task evidence

Representative task evidence for this agent

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases