Skip to content
AAA.win

DeepSeek Main

Best value profile for structured extraction and classification.

DeepSeeklowArena #2
Preview evidence60 runs
Provider
DeepSeek
Exact model version
Not measured
Evaluation batch
maa-preview-002
Test date
2026-06-27
Context window
Not measured
Tool mode
disabled
Pricing source
Not measured
Latency
Not measured

Best fit

  • 中文 workflows where the agent has its strongest language evidence.
  • Extraction tasks as a first shortlist candidate against at least two alternatives.
  • Budget-sensitive, lower-risk internal pilots.

Not a fit for

  • Critical-failure rate is 7%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.

Alternatives

Overall80
Pass rate70%
Critical7%
Format pass100%
Measured avg. cost$0.0050
LatencyNot measured
Recommended for

Where it should be considered first

  • 中文 workflows, especially where this agent shows its strongest language score.
  • Extraction tasks where it is a reasonable first shortlist candidate.
  • Budget-sensitive automation pilots.
Use caution

Review before production

  • Review failure tags such as weak_cta, missing_field, hallucinated_issue before production use.
  • Critical-failure rate is 7%; do not rely on the overall score alone for support, refund, or compliance workflows.
  • Format pass rate is 100%; retest JSON and missing-field behavior before extraction use cases go live.

Profile metrics

Overall score: 80 Win rate: 5% Pass rate: 70% Critical: 7% Format pass rate: 100% Average run cost: $0.0050

Common failure tags

weak_ctamissing_fieldhallucinated_issue

Language performance

中文82
English78
日本語79
Español80

Task type performance

Support79
Writing78
Extraction83
Decision profile

How should this agent enter a real pilot?

Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.

Test first
  • 中文 workflows where the agent has its strongest language evidence.
  • Extraction tasks as a first shortlist candidate against at least two alternatives.
  • Budget-sensitive, lower-risk internal pilots.
Do not launch first
  • Critical-failure rate is 7%; do not fully automate refund, compliance, safety, or legal work first.
  • Format pass is 100%; retest JSON, dates, and missing fields before extraction launch.
  • Do not treat the current score as permanent after model, pricing, or tool changes.
Pilot plan
  • Prepare 20-50 real samples across normal, edge, and high-risk cases.
  • Compare against one high-score candidate and one lower-cost candidate.
  • Record repair time, failure tags, and unacceptable failures before expanding automation.

Recommended pilot path

  1. Use 20–50 real samples across normal, edge, and high-risk inputs.
  2. Keep at least one lower-cost comparator.
  3. Human-review every critical failure and high-risk output.
  4. Retest after model, prompt, tool, or business-policy changes.
Build a full pilot plan
Task evidence

Representative task evidence for this agent

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases