# Encoder versus decoder deployment benchmarks
Research checked on 13 September 2026, supplementing [Document-level financial event extraction](<2026-09-13 A Review of document-level financial event extraction.md>). Named entity recognition (NER), large language model (LLM), and financial event extraction (FinEE) are used below.
**Small encoders have measured inference advantages on adjacent tasks.** This search did not locate a benchmark that jointly controls financial multi-event record quality, hardware, and complete serving cost for a small encoder versus a fine-tuned 1–8B generative decoder. The studies below support a narrower efficiency claim. Their measurements were not independently reproduced here.
## Same-hardware evidence
### Just Pass Twice — NER
[Just Pass Twice, §4.4 and Table 3][1] reports A100, batch-size-one timings on CrossNER-Politics:
| Model | Inference method | Reported elapsed seconds |
| :--- | :--- | ---: |
| GLiNER-L | Encoder span classification | 33.3 |
| UniNER-7B | Autoregressive extraction | 1,970.2 |
| JPT-4B | Qwen3 token classification | 89.7 |
| JPT-8B | Qwen3 token classification | 146.2 |
Calculated from these timings, GLiNER-L is approximately 59× faster than UniNER, but only 2.7×/4.4× faster than JPT. These are benchmark-run totals, not per-document latencies. JPT caches entity-definition embeddings and avoids generation.
**Quality caveat:** Table 3 is labeled CrossNER-Politics, but its GLiNER/UniNER F1 values match the seven-dataset averages in Table 1 rather than its Politics column. This inconsistency prevents treating Table 3 as a clean quality–cost frontier. Baseline quality scores are also imported from prior work. Precision, serving optimizations, and matched training are not fully controlled in the reported comparison.
The original [GLiNER paper, Table 1][2] labels GLiNER-L as 0.3B parameters. Treat this as the paper's reported count, rather than a verified total deployed parameter count.
### GLiNER Guard — safety moderation
[GLiNER Guard, Tables 1–2][3] reports batch-size-one inference on an A100 80GB:
| Model | Parameters | Seconds/request | Average safety F1 |
| :--- | ---: | ---: | ---: |
| GLiGuard bi-encoder | 145M | 0.019 | 74.3 |
| WildGuard | 7B | 0.744 | 80.9 |
| YuFeng-XGuard | 8B | 0.051 | 86.4 |
The encoder has approximately 48–55× fewer parameters. Calculated latency advantages range from 2.7× to 39×, with lower average quality. This is a production-oriented technical report with a short configured context; decoder inference settings are not sufficiently detailed to attribute the entire gap to architecture.
Its dynamic-batching P50/P95/P99 experiments compare encoders with encoders, not decoders. Do not transfer that concurrency evidence to the cross-architecture comparison.
### Power Hungry Processing — energy
[Luccioni et al., FAccT 2024, §§3.2 and 4.2][4] compares task-specific BERT-family models with BLOOMz decoders and Flan-T5 encoder–decoders on identical classification and question-answering samples. It uses 1,000 inputs, ten repetitions, sequential inference, and the same A100 environment. Specialized discriminative models consume less energy on these tasks.
This is useful controlled operational evidence, but model training and accuracy differ. Measurements include idle power from the other GPUs on the eight-GPU node. The older models and unbatched Transformers runtime do not establish current optimized serving ratios. Flan-T5 results must not be presented as encoder-only evidence; energy and emissions are not monetary operating cost.
## Same-workload evidence with different deployment environments
[GLiNER-Relex, §5.1][5] compares local DeBERTa-v3-large-based extraction on an L4, batch size one, with GPT-5-mini through its API. For 50 FineWeb documents averaging 288 words and a schema of six entity types and 50 relations, mean latency is 0.9 versus 64 seconds. The approximately 71× ratio includes network and default reasoning overhead. Quality results come from separate benchmarks; the timed corpus does not establish matched quality. Relation triples are closer to FinEE than classification, but do not test complete financial event records.
[GLiNER2, §3.3 and Table 4][6] reports CPU classification latency of 130–208 ms versus 358–463 ms through the GPT-4o API as label counts increase from five to fifty. Its approximately 2.6× reported speedup is an operational comparison, not a same-hardware experiment. Its separate NER quality results should not be combined with classification timings into an extraction quality–latency claim.
## Implication for the financial review
Suggested interpretation: controlled experiments on adjacent classification and extraction tasks provide empirical support for the inference-efficiency advantage of small specialized encoders over some larger generative decoders. The size of the advantage depends on the comparator and execution method; it does not establish equal complete-record quality or a specific deployment saving for financial document extraction.
A useful local experiment should distinguish three candidates: a small encoder with structured inference, a 1–8B autoregressive extractor, and a decoder backbone with discriminative heads. Compare identical documents and output requirements, quality thresholds, truncation, precision, optimized runtimes, and concurrency. Include preprocessing, event grouping, validation, peak memory, latency percentiles, and cost per accepted record.
## References
1. [Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER — ACL 2026][1]
2. [GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer — NAACL 2024][2]
3. [GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy — technical report, 2026][3]
4. [Power Hungry Processing: Watts Driving the Cost of AI Deployment? — FAccT 2024][4]
5. [GLiNER-Relex: A Unified Framework for Joint Named Entity Recognition and Relation Extraction — preprint, 2026][5]
6. [GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface — 2025][6]
[1]: https://aclanthology.org/2026.acl-long.526.pdf
[2]: https://aclanthology.org/2024.naacl-long.300.pdf
[3]: https://arxiv.org/html/2605.05277v1
[4]: https://arxiv.org/html/2311.16863v3
[5]: https://arxiv.org/html/2605.10108v1
[6]: https://arxiv.org/html/2507.18546v1