Skip to content
AAA.win

Claude Main

Strong writing and safety boundaries, especially in support tasks.

AnthropicpremiumArena #2
Preview evidence60 runs
Provider
Anthropic
Exact model version
Not measured
Evaluation batch
maa-preview-002
Test date
2026-06-27
Context window
Not measured
Tool mode
disabled
Pricing source
Not measured
Latency
Not measured

Best fit

  • English workflows where the agent has its strongest language evidence.
  • Support tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.

Not a fit for

  • Critical-failure rate is 12%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.

Alternatives

Overall87
Pass rate97%
Critical12%
Format pass100%
Measured avg. cost$0.0247
LatencyNot measured
Recommended for

Where it should be considered first

  • English workflows, especially where this agent shows its strongest language score.
  • Support tasks where it is a reasonable first shortlist candidate.
  • Quality-first tests for higher-value or higher-risk workflows.
Use caution

Review before production

  • Review failure tags such as too_verbose, overly_humble, unsafe_refund_promise before production use.
  • Critical-failure rate is 12%; do not rely on the overall score alone for support, refund, or compliance workflows.
  • Format pass rate is 100%; retest JSON and missing-field behavior before extraction use cases go live.

Profile metrics

Overall score: 87 Win rate: 55% Pass rate: 97% Critical: 12% Format pass rate: 100% Average run cost: $0.0247

Common failure tags

too_verboseoverly_humbleunsafe_refund_promise

Language performance

中文81
English92
日本語89
Español88

Task type performance

Support90
Writing90
Extraction82
Decision profile

How should this agent enter a real pilot?

Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.

Test first
  • English workflows where the agent has its strongest language evidence.
  • Support tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.
Do not launch first
  • Critical-failure rate is 12%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.
Pilot plan
  • Prepare 20-50 real samples across normal, edge, and high-risk cases.
  • Compare against one high-score candidate and one lower-cost candidate.
  • Record repair time, failure tags, and unacceptable failures before expanding automation.

Recommended pilot path

  1. Use 20–50 real samples across normal, edge, and high-risk inputs.
  2. Keep at least one lower-cost comparator.
  3. Human-review every critical failure and high-risk output.
  4. Retest after model, prompt, tool, or business-policy changes.
Build a full pilot plan
Task evidence

Representative task evidence for this agent

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases