Where it should be considered first
- English workflows, especially where this agent shows its strongest language score.
- Support tasks where it is a reasonable first shortlist candidate.
- Quality-first tests for higher-value or higher-risk workflows.
Strong writing and safety boundaries, especially in support tasks.
Overall score: 87 Win rate: 55% Pass rate: 97% Critical: 12% Format pass rate: 100% Average run cost: $0.0247
| 中文 | 81 |
| English | 92 |
| 日本語 | 89 |
| Español | 88 |
| Support | 90 |
| Writing | 90 |
| Extraction | 82 |
Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.