Skip to content
AAA.win

Grok Main

Fast outputs with higher variance on business constraints.

xAIstandardArena #2
Preview evidence60 runs
Provider
xAI
Exact model version
Not measured
Evaluation batch
maa-preview-002
Test date
2026-06-27
Context window
Not measured
Tool mode
disabled
Pricing source
Not measured
Latency
Not measured

Best fit

  • English workflows where the agent has its strongest language evidence.
  • Writing tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.

Not a fit for

  • Critical-failure rate is 27%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 78%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.

Alternatives

Overall75
Pass rate37%
Critical27%
Format pass78%
Measured avg. cost$0.0121
LatencyNot measured
Recommended for

Where it should be considered first

  • English workflows, especially where this agent shows its strongest language score.
  • Writing tasks where it is a reasonable first shortlist candidate.
  • Standard-budget business automation candidates.
Use caution

Review before production

  • Review failure tags such as unsafe_refund_promise, unsupported_claim, invalid_json before production use.
  • Critical-failure rate is 27%; do not rely on the overall score alone for support, refund, or compliance workflows.
  • Format pass rate is 78%; retest JSON and missing-field behavior before extraction use cases go live.

Profile metrics

Overall score: 75 Win rate: 0% Pass rate: 37% Critical: 27% Format pass rate: 78% Average run cost: $0.0121

Common failure tags

unsafe_refund_promiseunsupported_claiminvalid_json

Language performance

中文74
English79
日本語74
Español74

Task type performance

Support75
Writing77
Extraction75
Decision profile

How should this agent enter a real pilot?

Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.

Test first
  • English workflows where the agent has its strongest language evidence.
  • Writing tasks as a first shortlist candidate against at least two alternatives.
  • Quality-sensitive business pilots.
Do not launch first
  • Critical-failure rate is 27%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 78%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.
Pilot plan
  • Prepare 20-50 real samples across normal, edge, and high-risk cases.
  • Compare against one high-score candidate and one lower-cost candidate.
  • Record repair time, failure tags, and unacceptable failures before expanding automation.

Recommended pilot path

  1. Use 20–50 real samples across normal, edge, and high-risk inputs.
  2. Keep at least one lower-cost comparator.
  3. Human-review every critical failure and high-risk output.
  4. Retest after model, prompt, tool, or business-policy changes.
Build a full pilot plan
Task evidence

Representative task evidence for this agent

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases