Skip to content
AAA.win

OpenAI Main

Strong generalist with balanced writing and support safety.

OpenAIpremiumArena #2
Preview evidence60 runs
Provider
OpenAI
Exact model version
Not measured
Evaluation batch
maa-preview-002
Test date
2026-06-27
Context window
Not measured
Tool mode
disabled
Pricing source
Not measured
Latency
Not measured

Best fit

  • English workflows where the agent has its strongest language evidence.
  • Writing tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.

Not a fit for

  • Critical-failure rate is 12%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.

Alternatives

Overall86
Pass rate92%
Critical12%
Format pass100%
Measured avg. cost$0.0247
LatencyNot measured
Recommended for

Where it should be considered first

  • English workflows, especially where this agent shows its strongest language score.
  • Writing tasks where it is a reasonable first shortlist candidate.
  • Quality-first tests for higher-value or higher-risk workflows.
Use caution

Review before production

  • Review failure tags such as missed_dependency, generic_ai_copy, unsafe_refund_promise before production use.
  • Critical-failure rate is 12%; do not rely on the overall score alone for support, refund, or compliance workflows.
  • Format pass rate is 100%; retest JSON and missing-field behavior before extraction use cases go live.

Profile metrics

Overall score: 86 Win rate: 30% Pass rate: 92% Critical: 12% Format pass rate: 100% Average run cost: $0.0247

Common failure tags

missed_dependencygeneric_ai_copyunsafe_refund_promise

Language performance

中文85
English93
日本語82
Español83

Task type performance

Support86
Writing88
Extraction84
Decision profile

How should this agent enter a real pilot?

Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.

Test first
  • English workflows where the agent has its strongest language evidence.
  • Writing tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.
Do not launch first
  • Critical-failure rate is 12%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.
Pilot plan
  • Prepare 20-50 real samples across normal, edge, and high-risk cases.
  • Compare against one high-score candidate and one lower-cost candidate.
  • Record repair time, failure tags, and unacceptable failures before expanding automation.

Recommended pilot path

  1. Use 20–50 real samples across normal, edge, and high-risk inputs.
  2. Keep at least one lower-cost comparator.
  3. Human-review every critical failure and high-risk output.
  4. Retest after model, prompt, tool, or business-policy changes.
Build a full pilot plan
Task evidence

Representative task evidence for this agent

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases