Full task input
Extract pain points, counts, severity, representative comments, and three product suggestions.
Can the agent summarize messy Chinese app reviews without inventing pain points?
Extract pain points, counts, severity, representative comments, and three product suggestions.
Must merge similar issues, count accurately, cite evidence, and avoid unsupported suggestions.
Extract pain points, counts, severity, representative comments, and three product suggestions.
Must merge similar issues, count accurately, cite evidence, and avoid unsupported suggestions.
| Kimi Main | 92 | 0% critical |
| OpenAI Main | 89 | 0% critical |
| Qwen Main | 85 | 0% critical |
| GLM Main | 85 | 0% critical |
| Doubao Main | 85 | 0% critical |
| Claude Main | 83 | 0% critical |
| MiniMax Main | 83 | 0% critical |
| Gemini Main | 80 | 0% critical |
| Perplexity Main | 80 | 0% critical |
| ERNIE Main | 80 | 0% critical |
| Hunyuan Main | 80 | 0% critical |
| Yi Main | 80 | 0% critical |
| DeepSeek Main | 79 | 33% critical |
| Llama Main | 79 | 0% critical |
| Mistral Main | 79 | 0% critical |
| Grok Main | 77 | 33% critical |
| Cohere Main | 77 | 0% critical |
| Phi Main | 75 | 0% critical |
Kimi Main currently leads this task. When reading samples, focus on task completion, whether the output respects business boundaries such as hallucinated_issue, and whether the structure can move into a workflow. The most visible failure tag is unsupported_claim.
Preview seed output. Replace with exact model output for real evaluations.
Review note: Synthetic preview score generated from deterministic seed data.
Preview seed output. Replace with exact model output for real evaluations.
Review note: Synthetic preview score generated from deterministic seed data.
Downloads use the same run objects as this page.