Scoring
Each run is scored on task success, language fit, instruction following, business safety, and output reliability.
AAA.win evaluates specific agents on specific multilingual business tasks under documented conditions.
This page explains how AAA.win turns multilingual business work into reviewable agent evaluations: how tasks are selected, how runs are controlled, how scores should be read, and where the current preview should not be over-claimed.
This page explains how AAA.win turns multilingual business work into reviewable agent evaluations: how tasks are selected, how runs are controlled, how scores should be read, and where the current preview should not be over-claimed.
Each run is scored on task success, language fit, instruction following, business safety, and output reliability.
Each agent runs each task 3 times with tools, browsing, and memory disabled.
A critical failure is unsafe, misleading, unusable, or structurally invalid in a real business workflow.
Vendors cannot pay to change scores. Sponsored placements, if introduced later, will be labeled separately.
Tasks come from recurring business workflows such as support, writing, and structured extraction.
Each task keeps metadata for language, task type, difficulty, and primary risk.
Repeated runs reveal stability instead of rewarding one lucky output.
Each run receives a score, critical-failure flag, format check, failure tags, and judge notes.
Reports are generated from structured data, then reviewed before strong public claims.
The output must complete the task, not merely sound fluent.
Language evaluation includes tone, format, politeness, and market convention.
Many failures come from ignoring schema, fields, or role limits.
Unsafe refunds, invented security claims, and false contract data matter more than style.
Automation needs stable structure and honest handling of missing information.
Buyers should weigh score against cost, risk level, and workflow importance.
Failure tags explain what went wrong and where readers should inspect the raw output.
observed seed runs
observed seed runs
observed seed runs
observed seed runs
observed seed runs
observed seed runs
The public preview uses maa-preview-002.
Each task-agent pair runs 3 times.
Tools, browsing, and memory are disabled in the current batch.
Real task sets currently cover Chinese, English, Japanese, and Spanish.
The current arena is a useful comparison framework, not a permanent universal model ranking.
Reports should show batch ID, date, tool settings, and runs per pair.
Scores should trace back to prompts, outputs, tags, and notes.
Old results should keep dates and be replaced by new batches when conditions change.
Sponsored placements, if introduced later, must be clearly separated from ranking.