Skip to content
AAA.win

benchmark protocol

Structured document extraction benchmark protocol

A proposed protocol for testing field accuracy, evidence spans, abstention, and schema compliance on business documents.

Preview — no resultsVerified 2026-08-112 official or primary sources

Protocol

  1. Define every field, allowed format, and missing-value behavior before testing.
  2. Create a legally usable, representative document set with a human-reviewed answer key.
  3. Require a source page or section and minimal evidence span for each extracted value.
  4. Validate schemas mechanically, then blind-review semantic correctness.
  5. Report performance by document type and field instead of only one average.

Planned measurements

  • Exact and normalized field accuracy
  • Unsupported-value rate
  • Correct abstention rate
  • Source-span support rate
  • Valid-schema rate

Publication gate

The page remains preview and noindex until every gate is met.

  • Document rights and privacy review complete
  • Answer key reviewed by a qualified domain owner
  • Ambiguous fields separated from wrong fields
  • Every public claim linked to inspectable evidence
  • No legal interpretation claim inferred from extraction accuracy

Why results are not publishable yet

These are real missing run artifacts and external review conditions—not completed evidence.

  • No rights-cleared, reproducible business-document set and field-level answer key with stable identifiers and expected-answer records has been published.
  • No candidate run has frozen the product surface, model identifier, prompt, settings, tools, region, runtime, dependency versions, and hardware or service environment.
  • No completed run log records sample size, exclusions, failures, interventions, start date, end date, and the exact configuration used for every candidate.
  • No versioned scoring rubric and scoring implementation have been published with examples that another reviewer can reproduce.
  • No blind or independent human-review record signed by the domain owner who established the answer key has resolved disagreements and critical-failure labels.
  • No inspectable input-to-output evidence bundle exposes failed, incomplete, excluded, and successful cases rather than only selected examples.

Sources checked

Open the original pages before relying on a time-sensitive product decision.

  1. NIST AI Risk Management FrameworkNIST
  2. NIST Generative AI ProfileNIST
Version · v4.3.4-indexnow-root-proof

Latest releases

IndexNow root-proof request compatibility

After v4.3.3, the exact root-level {key}.txt proof returned HTTP 200, but a full request carrying keyLocation still returned HTTP 403; a minimal homepage request with the same production key and no keyLocation returned HTTP 202. This release omits that field, while the automatic full run and idempotent rerun remain deployment checks.

IndexNow key-proof compatibility

Changed IndexNow verification to the official root-level {key}.txt convention after the first production notification returned HTTP 403; revalidation remains pending, while the website, sitemaps, and Bing sitemap processing are unaffected.

Bing sitemap discovery and IndexNow change notifications

Prepared canonical sitemap discovery for Bing and added automatic IndexNow change notifications while keeping segmented sitemaps authoritative and making no claim that a notified URL has been crawled or indexed.

View all releases