Protocol
- Define every field, allowed format, and missing-value behavior before testing.
- Create a legally usable, representative document set with a human-reviewed answer key.
- Require a source page or section and minimal evidence span for each extracted value.
- Validate schemas mechanically, then blind-review semantic correctness.
- Report performance by document type and field instead of only one average.
Planned measurements
- Exact and normalized field accuracy
- Unsupported-value rate
- Correct abstention rate
- Source-span support rate
- Valid-schema rate
Publication gate
The page remains preview and noindex until every gate is met.
- Document rights and privacy review complete
- Answer key reviewed by a qualified domain owner
- Ambiguous fields separated from wrong fields
- Every public claim linked to inspectable evidence
- No legal interpretation claim inferred from extraction accuracy
Why results are not publishable yet
These are real missing run artifacts and external review conditions—not completed evidence.
- No rights-cleared, reproducible business-document set and field-level answer key with stable identifiers and expected-answer records has been published.
- No candidate run has frozen the product surface, model identifier, prompt, settings, tools, region, runtime, dependency versions, and hardware or service environment.
- No completed run log records sample size, exclusions, failures, interventions, start date, end date, and the exact configuration used for every candidate.
- No versioned scoring rubric and scoring implementation have been published with examples that another reviewer can reproduce.
- No blind or independent human-review record signed by the domain owner who established the answer key has resolved disagreements and critical-failure labels.
- No inspectable input-to-output evidence bundle exposes failed, incomplete, excluded, and successful cases rather than only selected examples.
Sources checked
Open the original pages before relying on a time-sensitive product decision.