Skip to content
AAA.win

Japanese Appointment Intent Classification

Can the agent classify short Japanese appointment messages into stable intent labels?

日本語Supportwrong_intent
Preview evidencemaa-preview-002

Full task input

Classify messages as booking, cancellation, reschedule, pricing_question, or other.

Scoring rubric

Must use only allowed labels and include short reasons.
Primary risk
wrong_intent
Temperature
0.2
Tools
Disabled
Web browsing
Disabled
Memory
Disabled

Agent prompt summary

Classify messages as booking, cancellation, reschedule, pricing_question, or other.

Rubric summary

Must use only allowed labels and include short reasons.

Task leaderboard

Claude Main9233% critical
Qwen Main8133% critical
Mistral Main800% critical
GLM Main800% critical
MiniMax Main800% critical
OpenAI Main7933% critical
Gemini Main7933% critical
DeepSeek Main790% critical
Llama Main790% critical
Perplexity Main790% critical
Kimi Main7933% critical
ERNIE Main790% critical
Hunyuan Main790% critical
Cohere Main7733% critical
Yi Main7733% critical
Doubao Main7633% critical
Grok Main7333% critical
Phi Main730% critical

Common failure tags

wrong_intentunsupported_claimliteral_translationweak_ctamissed_dependencytoo_verbose

Sample outputs and scoring notes

Claude Main currently leads this task. When reading samples, focus on task completion, whether the output respects business boundaries such as wrong_intent, and whether the structure can move into a workflow. The most visible failure tag is wrong_intent.

Claude Main

maa-preview-002.japanese-appointment-intent-classification.claude-main.02

Preview
Batch: maa-preview-002Model version: Not measuredTemperature: 0.2Cost: $0.0244Latency: Not measuredHuman review: Not reviewed

Raw output (untruncated)

Preview seed output. Replace with exact model output for real evaluations.
Valid JSONRequired fields presentNo critical failure
Run score100
task success
5/5
language fit
5/5
instruction following
5/5
business safety
5/5
output reliability
5/5

Review note: Synthetic preview score generated from deterministic seed data.

Grok Main

maa-preview-002.japanese-appointment-intent-classification.grok-main.02

Preview
Batch: maa-preview-002Model version: Not measuredTemperature: 0.2Cost: $0.0124Latency: Not measuredHuman review: Not reviewed

Raw output (untruncated)

Preview seed output. Replace with exact model output for real evaluations.
Valid JSONRequired fields missingCritical failure
wrong_intentunsupported_claim
Run score72
task success
4/5
language fit
4/5
instruction following
4/5
business safety
3/5
output reliability
3/5

Review note: Synthetic preview score generated from deterministic seed data.

Downloads use the same run objects as this page.

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases