Skip to content
AAA.win

Meeting Notes Action Item Extraction

Can the agent distinguish real action items from general meeting discussion?

EnglishExtractiondiscussion_as_action
Preview evidencemaa-preview-002

Full task input

Extract owner, deadline, task, and risk from English beta launch meeting notes.

Scoring rubric

Must use unclear when owner/deadline is absent and avoid turning discussions into tasks.
Primary risk
discussion_as_action
Temperature
0.2
Tools
Disabled
Web browsing
Disabled
Memory
Disabled

Agent prompt summary

Extract owner, deadline, task, and risk from English beta launch meeting notes.

Rubric summary

Must use unclear when owner/deadline is absent and avoid turning discussions into tasks.

Task leaderboard

OpenAI Main890% critical
Gemini Main850% critical
Claude Main830% critical
Qwen Main830% critical
Mistral Main830% critical
Cohere Main830% critical
GLM Main830% critical
Grok Main800% critical
DeepSeek Main800% critical
Llama Main800% critical
Kimi Main800% critical
Hunyuan Main800% critical
Perplexity Main7933% critical
ERNIE Main790% critical
MiniMax Main790% critical
Doubao Main790% critical
Yi Main7733% critical
Phi Main760% critical

Common failure tags

unsupported_claimdiscussion_as_actionliteral_translationgeneric_ai_copyunsafe_refund_promisecitation_overreach

Sample outputs and scoring notes

OpenAI Main currently leads this task. When reading samples, focus on task completion, whether the output respects business boundaries such as discussion_as_action, and whether the structure can move into a workflow. The most visible failure tag is unsupported_claim.

OpenAI Main

maa-preview-002.meeting-notes-action-items.openai-main.02

Preview
Batch: maa-preview-002Model version: Not measuredTemperature: 0.2Cost: $0.0245Latency: Not measuredHuman review: Not reviewed

Raw output (untruncated)

Preview seed output. Replace with exact model output for real evaluations.
Valid JSONRequired fields presentNo critical failure
Run score100
task success
5/5
language fit
5/5
instruction following
5/5
business safety
5/5
output reliability
5/5

Review note: Synthetic preview score generated from deterministic seed data.

Perplexity Main

maa-preview-002.meeting-notes-action-items.perplexity-main.01

Preview
Batch: maa-preview-002Model version: Not measuredTemperature: 0.2Cost: $0.0118Latency: Not measuredHuman review: Not reviewed

Raw output (untruncated)

Preview seed output. Replace with exact model output for real evaluations.
Valid JSONRequired fields presentCritical failure
discussion_as_actioncitation_overreach
Run score76
task success
4/5
language fit
4/5
instruction following
4/5
business safety
3/5
output reliability
4/5

Review note: Synthetic preview score generated from deterministic seed data.

Downloads use the same run objects as this page.

Version · v4.3.6-indexnow-delivery-verified

Latest releases

IndexNow production delivery verification

Verified production IndexNow receipt with a new key: Microsoft's external proof returned HTTP 200, single-URL GET and POST both returned HTTP 200, and eight batches covering 386 canonical URLs each returned HTTP 200 with zero retries. A second run found the inventory unchanged and sent no request.

IndexNow proof-origin protocol fix

Production Bing validation exposed HTTP 500 at the external root proof because TLS termination was followed by an internal rewrite using the wrong protocol. This release corrects that rewrite boundary, but remains unverified until the external proof returns HTTP 200, the automatic submission runs, a later canary returns HTTP 200, and the idempotent rerun passes.

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

View all releases