Full task input
Rewrite a rough SaaS description into H1, subheadline, three value props, and CTA.
Can the agent turn vague SaaS marketing copy into specific, credible hero copy?
Rewrite a rough SaaS description into H1, subheadline, three value props, and CTA.
Must avoid buzzwords and unsupported claims while keeping multilingual ticket triage specific.
Rewrite a rough SaaS description into H1, subheadline, three value props, and CTA.
Must avoid buzzwords and unsupported claims while keeping multilingual ticket triage specific.
| OpenAI Main | 93 | 0% critical |
| Claude Main | 92 | 0% critical |
| Llama Main | 83 | 0% critical |
| Mistral Main | 83 | 0% critical |
| Cohere Main | 83 | 0% critical |
| Perplexity Main | 83 | 0% critical |
| GLM Main | 83 | 0% critical |
| MiniMax Main | 83 | 0% critical |
| Gemini Main | 80 | 0% critical |
| Qwen Main | 80 | 0% critical |
| Kimi Main | 80 | 0% critical |
| Doubao Main | 80 | 0% critical |
| Grok Main | 79 | 33% critical |
| ERNIE Main | 79 | 0% critical |
| Hunyuan Main | 79 | 0% critical |
| Yi Main | 77 | 0% critical |
| Phi Main | 76 | 0% critical |
| DeepSeek Main | 75 | 33% critical |
OpenAI Main currently leads this task. When reading samples, focus on task completion, whether the output respects business boundaries such as generic_ai_copy, and whether the structure can move into a workflow. The most visible failure tag is literal_translation.
Preview seed output. Replace with exact model output for real evaluations.
Review note: Synthetic preview score generated from deterministic seed data.
Preview seed output. Replace with exact model output for real evaluations.
Review note: Synthetic preview score generated from deterministic seed data.
Downloads use the same run objects as this page.