PARSING ACCURACY MATTERS.
ParseHook dynamically routes emails through a curated selection of enterprise AI models. This benchmark evaluates 19 models on 304 multilingual emails to measure JSON validity, field accuracy, and latency.
Our Multi-Model engine guarantees a minimum of 99% field accuracy on standard formats. Models are anonymized to protect our proprietary routing logic and prevent model-specific reverse engineering.
METHODOLOGY
- Dataset: 304 synthetic emails across 9 categories (Alerts, Invoices, Leads, Orders, Receipts, Shipping, etc.).
- Languages: English, German, Spanish, French.
- Evaluation: Automated Python scripts measuring JSON schema validity and exact field accuracy.
- Note on PDFs: PDF attachment parsing is tested separately and operates near flawlessly. This benchmark focuses on email body parsing.
ROUTING ARCHITECTURE
When an email arrives, it is first sanitized by our Python pre-processor to strip threads and signatures without using AI. The payload then enters our Semantic Routing Engine:
- Keyword & Whitelist Check: Identifies context (e.g., specific keywords trigger the German Invoice profile).
- Specialist Routing: Routes to the highest-scoring model for that specific context and region.
- Dynamic Fallback Chain: If the specialist fails or no keyword matches, it cascades to our Primary Allrounders until a valid JSON payload is achieved.
ANONYMIZATION POLICY: Model names are intentionally obfuscated. This protects our proprietary routing logic, prevents reverse engineering of our pipeline, and shields our infrastructure from targeted rate-limiting or manipulation.
| Rank | Model Alias | Score | Latency | JSON Valid | Field Acc. | Role / Profile |
|---|---|---|---|---|---|---|
| 1 | Model Alpha-1 | 96.0 | 2190 ms | 100.0% | 99.5% | Primary Allrounder. Highest stability for complex layouts. |
| 2 | Model Beta-2 | 94.4 | 1449 ms | 100.0% | 100.0% | Rapid Parser. Exceptional on short-form invoices. |
| 3 | Model Gamma-3 | 92.6 | 1501 ms | 100.0% | 100.0% | Secondary Allrounder. Strong on nested table data. |
| 4 | Model Delta-4 | 80.8 | 5598 ms | 100.0% | 100.0% | Specialist for Spanish and Portuguese shipping alerts. |
| 5 | Model Epsilon-5 | 77.0 | 1419 ms | 100.0% | 100.0% | Deep Context Fallback. High accuracy, slightly higher latency. |
| 6 | Model Zeta-6 | 74.2 | 3068 ms | 100.0% | 100.0% | Specialist for formatted e-commerce order confirmations. |
| 7 | Model Eta-7 | 72.3 | 8102 ms | 98.7% | 98.8% | Niche expert for Portuguese purchase orders and non-standard layouts. |
| 8 | Model Theta-8 | 71.9 | 8501 ms | 100.0% | 100.0% | Fallback for massive email threads and historical stripping. |
| 9 | Model Iota-9 | 71.9 | 5389 ms | 100.0% | 100.0% | Technical routing expert for server alerts and stack traces. |
| 10 | Model Kappa-10 | 68.5 | 3111 ms | 100.0% | 100.0% | Specialist for German invoices and VAT extraction. |
| 11 | Model Lambda-11 | 68.0 | 4111 ms | 100.0% | 100.0% | Fallback for payment receipts and transactional data. |
| 12 | Model Mu-12 | 53.3 | 6002 ms | 100.0% | 98.7% | Long-form lead capture. Occasional JSON drift on unstructured text. |
| 13 | Model Nu-13 | 52.9 | 2295 ms | 100.0% | 100.0% | High-speed fallback. Used when primary chain faces rate limits. |
| 14 | Model Xi-14 | 50.6 | 3227 ms | 100.0% | 100.0% | Fallback for French receipts. Struggles with nested arrays. |
| 15 | Model Omicron-15 | 47.2 | 4098 ms | 100.0% | 100.0% | Last-resort fallback. Minimal context window, strict formatting required. |
| 16 | Model Pi-16 | 0.0 | 1720 ms | 100.0% | 99.8% | Disqualified: API instability and strict regional data-passing restrictions. |
| 17 | Model Rho-17 | 0.0 | 2030 ms | 100.0% | 100.0% | Disqualified: Low availability during peak hours and hallucination risk on null fields. |
| 18 | Model Sigma-18 | 0.0 | 5390 ms | 82.9% | 99.1% | Disqualified: High JSON validation failure rate (>15%) breaking backend pipelines. |
| 19 | Model Tau-19 | 0.0 | 3489 ms | 100.0% | 100.0% | Disqualified: Vendor data-retention policies conflict with our No-AI-Training standard. |
SCORE CALCULATION
The Model Score is a weighted average designed to identify models that deliver perfect structural data at optimal processing speeds. Cost is intentionally omitted from public display to focus purely on quality and backend stability.
PHASE 1: DUAL HARD-CUTO
A model must pass two non-negotiable quality gates to even enter the scoring race:
- JSON Validity ≥ 98% (Prevents backend pipeline crashes)
- Field Accuracy ≥ 95% (Ensures data integrity for the user)
If a model fails either, its score is instantly zeroed and it is disqualified from production routing.
PHASE 2: THE PERFORMANCE RACE
Qualified models are ranked by a composite of structural accuracy and latency tiers:
- < 3,000 ms latency → 100 Pts (Lightning Fast)
- 3,000 - 6,000 ms latency → 70 Pts (Standard Webhook)
- 6,000 - 10,000 ms latency → 40 Pts (Medium)
- > 10,000 ms latency → 0 Pts (Timeout Risk)
GUARANTEED MIN 99% ACCURACY
Stop fixing broken parsing rules. Let AI route your emails dynamically.
If we fail to hit 99% accuracy on standard formats, you get your paid plan money back. Read the full guarantee.