AI BENCHMARK REPORT

PARSING ACCURACY MATTERS.

ParseHook dynamically routes emails through a curated selection of enterprise AI models. This benchmark evaluates 19 models on 304 multilingual emails to measure JSON validity, field accuracy, and latency.

Our Multi-Model engine guarantees a minimum of 99% field accuracy on standard formats. Models are anonymized to protect our proprietary routing logic and prevent model-specific reverse engineering.

METHODOLOGY

  • Dataset: 304 synthetic emails across 9 categories (Alerts, Invoices, Leads, Orders, Receipts, Shipping, etc.).
  • Languages: English, German, Spanish, French.
  • Evaluation: Automated Python scripts measuring JSON schema validity and exact field accuracy.
  • Note on PDFs: PDF attachment parsing is tested separately and operates near flawlessly. This benchmark focuses on email body parsing.

ROUTING ARCHITECTURE

When an email arrives, it is first sanitized by our Python pre-processor to strip threads and signatures without using AI. The payload then enters our Semantic Routing Engine:

  1. Keyword & Whitelist Check: Identifies context (e.g., specific keywords trigger the German Invoice profile).
  2. Specialist Routing: Routes to the highest-scoring model for that specific context and region.
  3. Dynamic Fallback Chain: If the specialist fails or no keyword matches, it cascades to our Primary Allrounders until a valid JSON payload is achieved.

ANONYMIZATION POLICY: Model names are intentionally obfuscated. This protects our proprietary routing logic, prevents reverse engineering of our pipeline, and shields our infrastructure from targeted rate-limiting or manipulation.

RankModel AliasScoreLatencyJSON ValidField Acc.Role / Profile
1Model Alpha-196.02190 ms100.0%99.5%Primary Allrounder. Highest stability for complex layouts.
2Model Beta-294.41449 ms100.0%100.0%Rapid Parser. Exceptional on short-form invoices.
3Model Gamma-392.61501 ms100.0%100.0%Secondary Allrounder. Strong on nested table data.
4Model Delta-480.85598 ms100.0%100.0%Specialist for Spanish and Portuguese shipping alerts.
5Model Epsilon-577.01419 ms100.0%100.0%Deep Context Fallback. High accuracy, slightly higher latency.
6Model Zeta-674.23068 ms100.0%100.0%Specialist for formatted e-commerce order confirmations.
7Model Eta-772.38102 ms98.7%98.8%Niche expert for Portuguese purchase orders and non-standard layouts.
8Model Theta-871.98501 ms100.0%100.0%Fallback for massive email threads and historical stripping.
9Model Iota-971.95389 ms100.0%100.0%Technical routing expert for server alerts and stack traces.
10Model Kappa-1068.53111 ms100.0%100.0%Specialist for German invoices and VAT extraction.
11Model Lambda-1168.04111 ms100.0%100.0%Fallback for payment receipts and transactional data.
12Model Mu-1253.36002 ms100.0%98.7%Long-form lead capture. Occasional JSON drift on unstructured text.
13Model Nu-1352.92295 ms100.0%100.0%High-speed fallback. Used when primary chain faces rate limits.
14Model Xi-1450.63227 ms100.0%100.0%Fallback for French receipts. Struggles with nested arrays.
15Model Omicron-1547.24098 ms100.0%100.0%Last-resort fallback. Minimal context window, strict formatting required.
16Model Pi-160.01720 ms100.0%99.8%Disqualified: API instability and strict regional data-passing restrictions.
17Model Rho-170.02030 ms100.0%100.0%Disqualified: Low availability during peak hours and hallucination risk on null fields.
18Model Sigma-180.05390 ms82.9%99.1%Disqualified: High JSON validation failure rate (>15%) breaking backend pipelines.
19Model Tau-190.03489 ms100.0%100.0%Disqualified: Vendor data-retention policies conflict with our No-AI-Training standard.

SCORE CALCULATION

The Model Score is a weighted average designed to identify models that deliver perfect structural data at optimal processing speeds. Cost is intentionally omitted from public display to focus purely on quality and backend stability.

PHASE 1: DUAL HARD-CUTO

A model must pass two non-negotiable quality gates to even enter the scoring race:

  • JSON Validity ≥ 98% (Prevents backend pipeline crashes)
  • Field Accuracy ≥ 95% (Ensures data integrity for the user)

If a model fails either, its score is instantly zeroed and it is disqualified from production routing.

PHASE 2: THE PERFORMANCE RACE

Qualified models are ranked by a composite of structural accuracy and latency tiers:

  • < 3,000 ms latency → 100 Pts (Lightning Fast)
  • 3,000 - 6,000 ms latency → 70 Pts (Standard Webhook)
  • 6,000 - 10,000 ms latency → 40 Pts (Medium)
  • > 10,000 ms latency → 0 Pts (Timeout Risk)

GUARANTEED MIN 99% ACCURACY

Stop fixing broken parsing rules. Let AI route your emails dynamically.

If we fail to hit 99% accuracy on standard formats, you get your paid plan money back. Read the full guarantee.