Where it should be considered first
- English workflows, especially where this agent shows its strongest language score.
- Writing tasks where it is a reasonable first shortlist candidate.
- Quality-first tests for higher-value or higher-risk workflows.
Strong generalist with balanced writing and support safety.
Overall score: 86 Win rate: 30% Pass rate: 92% Critical: 12% Format pass rate: 100% Average run cost: $0.0247
| 中文 | 85 |
| English | 93 |
| 日本語 | 82 |
| Español | 83 |
| Support | 86 |
| Writing | 88 |
| Extraction | 84 |
Translate the score into practical selection guidance: fit, no-fit, pilot path, and task evidence.