Skip to content

Competitive benchmarks ​

Where does Verbora stand against the Rust ecosystem? This page reports like-for-like, version-pinned measurements — same input, equivalent result, and a public loss whenever another library is faster.

Read by capability, not by one global score. A tokenizer, a stemmer and a string-distance function solve different problems. Each section therefore shows the measured workload, the result and its limits.
CoverageWhat it means
614 measurements15 benchmarked modules, every figure traced to results/results.json
14 modules with a fair Verbora-vs-Rust comparisonevery module below except POS tagging, where the comparison is withdrawn (see below)
2 modules with no Rust peer at allPhonetic Index/Neighbors and sentence analysis — documented on their own feature pages instead
1 comparison withdrawnPOS tagging — verbora-tagger 0.3.0 ships no lexicon, so the competitor figures stand alone

Start with Phonetics, Trie or Language detection. Exact harnesses and raw data are linked from each section; reproduction instructions are at the end.

Methodology and audit details

Benchmark methodology ​

CPUIntel(R) Core(TM) i9-14900KF (32 threads)
Memory125 GiB
OSLinux 7.0.11-76070011-generic
rustc1.97.1, --release (opt-level = 3, lto = "thin", codegen-units = 16) — identical [profile.release]/[profile.bench] to the main Verbora workspace, not a tuned profile for this audit
Node.jsv25.9.0 (used only for the three JavaScript-library-only modules linked above, not for the tables on this page)
Verbora treeTagged bench-2026-08-22. The tag is the durable anchor: a branch commit does not survive a squash merge, and this page once cited a hash that resolves nowhere. Within it, every module except Phonetics was measured at 80c302b and Phonetics re-measured at 0313eae — one commit later, touching scripts/ and site/ only, no crates/ change, so this is the same phonetics code measured a second time, not a different implementation. Crate versions: verbora-tagger 0.3.0, verbora-wordnet 0.3.0, every other crate measured on this page 0.2.0
DatasetsShared word/name/pair lists from benches/data/*.json (tools/bench-data/generate.py, one generator read by every implementation); the 13-language, 4-tier UDHR corpus for language-detection accuracy (sourced below); the 2,438-word AFINN-111/AFINN-165 intersection for sentiment; a real Princeton WordNet 3.1 dict/ distribution for WordNet
WarmupCriterion's own warmup phase before every measured sample (400 ms–1 s per group; see below)
SamplesCriterion's default 100 per benchmark, reduced for the most expensive groups: 30 for language/script detection and transliteration, 20 for spellcheck batch-correction and distance-2 corrections, 15 for POS-tagging cold start (rust-bert's model load), 10 for WordNet's open/cold groups (real-file I/O)
MetricMedian, per Criterion's own robust-statistics estimate — not mean, per this project's own PRIMARY METRIC policy
Threads1 (single-threaded) for every benchmark on this page — no parallel API is exercised anywhere in this audit; see Thread counts
Sourcebenchmarks/competitive/rust-competitors/benches/*.rs (one file per module), raw Criterion output under benchmarks/competitive/results/raw/, joined into results/results.json
DateEvery measurement on this page is stamped 2026-08-22. Both commits inside the tag were cut the same day and differ only outside crates/, so this is one campaign's figures, not a mixed-vintage set

Every number on this page is read directly from results.json's saved median_ns values; the relative-speedup figures are computed from those same values at page-generation time — none is retyped from memory or rounded inconsistently. See Reproducing these numbers for the exact commands that regenerate all of it from a clean checkout.

One machine, one run. Per the project's own environment policy (results/metadata.json's own note): "a single dedicated run on one physical machine with no other significant workload active" — CPU affinity pinning, thermal monitoring and containerization were evaluated and deliberately not built, since the spec marks all three optional. Treat exact figures as one machine's numbers and the ratios/orders-of-magnitude as the more portable signal, exactly as String distance results already asks of its own numbers.

Thread counts ​

No benchmark on this page compares a parallel implementation against a sequential one. Every call measured here — Verbora's and every competitor's — runs on a single thread; where Verbora exposes a parallel-feature batch API (phonetics' par_encode_batch, language detection's par_detect_batch), that comparison is sequential-vs-sequential too, because no competitor in this audit exposes an equivalent batch-parallel API to compare against. Verbora's own sequential-vs-parallel numbers, with thread counts disclosed, live on the Parallelism page instead — a different question from this one.

How to read these tables ​

  • Time (median) is Criterion's median estimate for one call of the benchmarked operation (or, where noted, one call over a fixed-size batch — the input-size column says which).
  • Throughput is 1 ÷ median time — calls of this exact operation per second at the measured latency. It is a call-rate figure, not a per-item/per-token throughput, because this audit does not have a verified items-per-call count for every operation; where a benchmark already operates over a fixed-size batch (e.g. "1024 words"), throughput is still batches/second, not words/second — read the input-size column together with it.
  • Relative is each row's time divided by the fastest row's time in that same table, labeled slower above 1.00×. Rows are ordered by time, not by library — Verbora is not pinned to the top row, and several tables on this page show it losing.
  • Where a comparison is explicitly not a fair like-for-like ratio (a different dictionary, a different corpus), no Relative column is shown at all, and the table says so — see Spellcheck for the cases this applies to.
  • A figure is published only while the code and the measurement behind it are current. Nothing on this page is filled in with an estimate, an interpolation, or an inference from a prior run's direction of change.

How these numbers were audited ​

This page's structure — which comparisons are fair, which are narrowed, and why — was cleared by an independent fairness audit that read every benchmark file and correctness test in this workspace, re-ran the Rust suite and the language-accuracy report itself, and cross-referenced every disclosed loss against docs/PERFORMANCE_GAPS.md. Its verdict: every comparison retained on this page is FAIR — same input, genuinely equivalent (or honestly narrowed and labeled) semantics, black_box on every call's input and output, correctness-before-performance tests that were run and passed, and version pins of =x.y.z on every third-party crate. Items the audit flagged as borderline rather than unfair are called out inline where they occur: the normalizers' accented-input case (Normalizers) and the WhatlangDetector wrapper-overhead check (Language detection), neither of which is a ranked "X beats Y" comparison in the first place.

The byte-exact phonetics encoder table in Phonetics carries a stronger correctness check than the rest of the page: byte-exact output equality with the competitor itself, asserted in tests/phonetics_correctness.rs and independently re-verified by an adversarial audit that differentially fuzzed 104,114 inputs per encoder against rphonetic with zero mismatches, proved every documented divergence exactly as narrow as each module's own documentation claims, and mutation-tested the correctness suites.

Results by capability ​

Per this project's OVERALL SCORE policy, there is no combined ranking anywhere on this page. Tokenization, stemming, spell-correction and every other capability below are different workloads solving different problems — mixing their numbers into one score would hide more than it reveals.

Distance ​

Verbora's edit-distance functions (docs/COMPETITIVE_BENCHMARKS.md§1.8) against strsim 0.11.1 and rapidfuzz 0.5.0 — the Rust ecosystem's de-facto-standard string-similarity crate (strsim: ~990M downloads) and the tightest single-crate algorithm match found. Both index by Unicode scalar value, exactly as Verbora does, and are restricted to ASCII input here in any case. Unlike the heuristic encoders elsewhere on this page, Levenshtein, Damerau-Levenshtein, Hamming, Jaro and Jaro-Winkler are exact, well-specified integer/float functions with a single correct answer per input. Their equivalence is nonetheless pinned by a runtime correctness suite, not by algorithm research alone: every timed implementation is asserted against Verbora on the shared corpus, on the near-identical and planted-needle derived shapes, and on a seeded randomized sweep, before any timing number on this page is accepted.

rapidfuzz implements Myers/Hyyrö bit-parallel Levenshtein (O(nm/64)). Verbora matches that algorithmic class with a single-word bit-vector fast path plus a multi-word block extension (Hyyrö's 2003 generalisation, verified independently line-by-line, then adversarially fuzz- and mutation-tested against the trusted scalar DP before being trusted for anything). The kernels' pattern-match (Peq) tables are flat/packed bit tables rather than a hash map, and the single-word gate covers 1–64 units. Bit-parallelism extends beyond plain Levenshtein too: restricted-Damerau/OSA kernels (Hyyrö's 2003 transposition extension of Myers, single-word and multi-word block, gated to unit costs), and Jaro/Jaro-Winkler match-flagging kernels in Verbora's own greedy orientation. Unrestricted Damerau-Levenshtein has no bit-vector formulation, so it runs Zhao–Sahni's linear-space algorithm instead of a full cost+parent matrix: three rows and a last-occurrence table, kept on the stack for short operands, with a table-free stack matrix for byte operands of at most 8 units. Common-prefix/ suffix trimming runs ahead of the unrestricted-Damerau and OSA kernels alike. Every kernel is parity-verified by differential tests against the retained scalar implementations, plus an independent adversarial audit with mutation testing, before being trusted; Hamming has no bit-parallel kernel.

Plain Levenshtein beats all six Rust competitors at every size from 4 to 1024 characters (Verbora is 1.15× faster than the closest competitor at 1024 characters, 1.83× at 16). Restricted Damerau/OSA beats every competitor at every size, and Jaro/Jaro-Winkler wins every size too. Unrestricted Damerau-Levenshtein is the one metric in this section with a genuine size-dependent crossover — see its own section below. See PERFORMANCE_GAPS.md entry 26 for the mechanism and verification story these kernels build on.

Levenshtein ​

stringmetrics 2.2.2 joins as a fourth full-equivalence (Yes/Yes) competitor here and in Hamming below — char-indexed like strsim/ rapidfuzz, just without a Damerau-Levenshtein implementation at all (its damerau module is commented out of both mod and pub use in the published crate, confirmed by reading stringmetrics-2.2.2's own source — not merely unused, genuinely uncompiled).

LibraryVersionLanguageTime (median, 1024 chars)ThroughputRelative
Verbora0.2.0Rust27.72 µs36.1K/s1.00×
rapidfuzz0.5.0Rust31.85 µs31.4K/s1.15× slower
strsim0.11.1Rust648.83 µs1.5K/s23.41× slower
stringmetrics2.2.2Rust973.08 µs1.0K/s35.10× slower

Random pairs — two independently generated strings, so the edit script is long and every implementation does its full work:

InputVerborarapidfuzzstrsimstringmetrics
410.9 ns38.1 ns18.3 ns25.8 ns
1641.2 ns75.5 ns182.7 ns176.5 ns
64160.8 ns258.6 ns2.89 µs3.13 µs
2562.20 µs3.37 µs41.38 µs56.62 µs
102427.72 µs31.85 µs648.83 µs973.08 µs

Near-identical pairs (d = 1) — the shape spell-checking and deduplication actually feed a distance metric, and the one where the implementations diverge most:

InputVerborarapidfuzzstrsimstringmetrics
410.5 ns21.0 ns18.1 ns16.3 ns
1611.7 ns40.1 ns177.3 ns19.5 ns
6413.7 ns138.9 ns2.90 µs48.7 ns
25620.0 ns277.1 ns42.43 µs166.5 ns
102441.5 ns868.8 ns660.07 µs593.6 ns

Verbora is the fastest implementation at every size and both shapes — against all four char-indexed competitors here and the two byte-level ones below. But the two tables answer different questions, and the result only has its real shape if you read both.

On random pairs the lead over rapidfuzz — the closest competitor and the only other bit-parallel implementation here — runs 3.50× at 4 characters, 1.83× at 16, 1.61× at 64, 1.53× at 256 and 1.15× at 1024. It is widest at small sizes, where the flat [u64; 256] Peq tables and the 1–64-unit single-word gate keep per-call setup low, and narrowest at 1024, where both sides run the same class of multi-word block algorithm. Against strsim the win is 1.68× (4) up to 23.41× (1024), against stringmetrics 2.37× up to 35.10× — neither scalar design has a bit-vector formulation to close with.

On near-identical pairs the margin moves the other way: it widens with length, from roughly 1.5× to 16× against the best char-indexed competitor at each size. Common prefixes and suffixes are trimmed before the kernel runs, so Verbora's cost rises only from 10.5 ns to 41.5 ns across a 256-fold increase in input length, while strsim — which evaluates the full matrix regardless of how similar the inputs are — rises to 660 µs, a factor of roughly 16,000.

The random rows are intentionally retained as their own workload. They do not show the common-affix optimization: independent strings have almost no affix to remove. The same competitive harness also measures a 1,024-unit pair with one central substitution (1024-near):

LibraryMedian (1024-near)Relative to Verbora
Verbora41.5 ns1.00×
editdistancek192.3 ns4.64× slower
stringmetrics593.6 ns14.32× slower
rapidfuzz868.8 ns20.96× slower
triple_accel525.17 µs12,655× slower
strsim660.07 µs15,904× slower

Two different designs sit behind those numbers, and the spread separates them cleanly. editdistancek is the only competitor built for this shape — its bounded-edit algorithm stops early when the distance is small, which is why it lands two orders of magnitude ahead of the full-matrix implementations. triple_accel and strsim evaluate the whole matrix whether the inputs differ by one character or by a thousand, so similarity buys them nothing.

Verbora reaches 41.5 ns by removing the common prefix and suffix before the kernel sees anything: on a 1,024-unit pair differing in one position, the kernel runs over what is left, not over the pair. The cost is therefore set by the size of the difference rather than the size of the input — 10.5 ns at 4 units, 41.5 ns at 1,024.

Competitive shape suite ​

The competitive harness also has a dedicated levenshtein_edge_shapes group. It uses the exact same lowercase-ASCII, unit-cost inputs for every implementation, so both char-indexed and byte-indexed competitors remain directly comparable. It keeps the shape-sensitive results separate from the random-size table above:

CaseOperandsVerboraFastest competitorSlowest competitor
near/10241,024 vs. 1,024; one central substitution42.2 nseditdistancek, 186.0 ns (4.40×)strsim, 626.73 µs (14,834×)
disjoint/10241,024 vs. 1,024; no character overlap1.20 µsrapidfuzz, 28.45 µs (23.70×)editdistancek, 1.40 ms (1,167×)
late-overlap/65x1000065 vs. 10,000; overlap only at the end1.94 µsrapidfuzz, 44.72 µs (23.00×)editdistancek, 51.12 ms (26,296×)

Verbora wins all three shapes against all five competitors. This is deliberately a separate group, rather than an extra median in the table above: random pairs, near pairs and disjoint alphabets exercise different valid algorithmic shortcuts, and blending them would hide that trade-off. The accompanying correctness test checks that all six implementations return the same distance on every timed shape. Reproduce it with:

bash
cd benchmarks/competitive/rust-competitors
cargo test --test distance_correctness levenshtein_competitors_agree_on_the_timed_edge_shapes
cargo bench --bench distance -- levenshtein_edge_shapes

Damerau–Levenshtein (unrestricted) ​

damerau_levenshtein computes canonical unrestricted Damerau-Levenshtein — the Lowrance–Wagner distance, where a transposed pair of adjacent characters costs one edit and may be edited again afterwards. It is a true metric: symmetric, and satisfying the triangle inequality. strsim's damerau_levenshtein and rapidfuzz's distance::damerau_levenshtein compute the same function, so agreement between the three is a property of the algorithm rather than of the corpus they were compared on — and it is checked that way: 202,000 randomized pairs, zero divergences, across alphabets of 2, 3, 4 and 26 letters, lengths 1–25 including unequal and empty operands, and mutation chains of up to eight edits over a binary alphabet (the shape that actually separates unrestricted Damerau from OSA). Reproduce it with:

bash
cd benchmarks/competitive/rust-competitors
cargo test --release --test distance_correctness \
  unrestricted_damerau_agrees_with_both_competitors_over_a_wide_randomized_sweep

Verbora runs Zhao–Sahni's linear-space formulation of that recurrence: three rows plus a last-occurrence table rather than a full cost+parent matrix, on the stack for short operands and a single allocation beyond, with a table-free stack matrix for byte operands of at most 8 units. Because the distance is symmetric, the shorter operand becomes the column operand, which is what sets the row width and therefore the kernel's whole memory footprint. Operands are stripped of their maximal common prefix and suffix first — a genuine algorithmic reduction rather than a micro-optimisation, since a pair differing in one interior position collapses to work proportional to the differing region alone. Weighted (non-unit-cost) calls take the full-matrix path instead, which evaluates the same recurrence with per-operation weights and doubles as the differential oracle for every fast path.

A genuine size-dependent crossover on random pairs, and a clean win on near-identical pairs. No bit-vector formulation exists for the unrestricted recurrence, so Verbora's linear-space scalar algorithm competes directly against `strsim` and `rapidfuzz`'s own scalar implementations — and the three are close enough on random input that the ranking flips with size.
InputVerborarapidfuzzstrsim
427.6 ns74.1 ns60.3 ns
16260.3 ns522.5 ns510.1 ns
644.63 µs8.17 µs7.76 µs
256134.80 µs135.57 µs133.86 µs
10242.50 ms2.22 ms2.03 ms

Verbora wins at 4, 16 and 64 characters (1.8×–2.7× faster than the closer of the two competitors), is essentially tied with strsim at 256 (1.007× slower — within this run's noise floor), and loses at 1024 characters (1.23× slower than strsim, 1.13× slower than rapidfuzz). On the near-identical shape, where common-prefix/suffix trimming collapses the work to the size of the actual edit rather than the size of the operand, Verbora wins outright at every size, including 1024 (47.0 ns vs. rapidfuzz's 852.2 ns, 18.1× faster, and strsim's 2.08 ms, over 44,000× slower on this shape since it evaluates the full matrix regardless of similarity).

Damerau–Levenshtein (restricted / OSA) ​

LibraryVersionLanguageTime (median, 1024 chars)ThroughputRelative
Verbora0.2.0Rust35.97 µs27.8K/s1.00×
rapidfuzz0.5.0Rust43.28 µs23.1K/s1.20× slower
strsim0.11.1Rust2.39 ms419.3/s66.32× slower
Input sizeVerborarapidfuzzstrsim
416.6 ns27.9 ns69.4 ns
1648.8 ns78.9 ns306.2 ns
64179.8 ns255.2 ns4.35 µs
2562.59 µs3.02 µs138.38 µs
102435.97 µs43.28 µs2.39 ms

Verbora is the fastest at every size. Restricted Damerau's one-transposition-back reach needs more state than plain Levenshtein's two-row shape, so it gets its own bit-parallel kernels implementing Hyyrö's 2003 transposition extension of Myers' algorithm — a single-word kernel plus a multi-word block generalisation, gated to unit-cost options, with the scalar three-row DP retained for every non-unit-cost call and as the differential-test oracle. Against rapidfuzz, the only other bit-parallel OSA here: 1.69× faster at 4 characters, 1.62× at 16, 1.42× at 64, 1.17× at 256, 1.20× at 1024. Against strsim's scalar implementation the margin runs 4.19× (4 chars) up to 66.3× (1024). triple_accel's byte-level rdamerau is covered in the byte-level subsection below.

Hamming ​

LibraryVersionLanguageTime (median, 1024 chars)ThroughputRelative
Verbora0.2.0Rust22.6 ns44.3M/s1.00×
stringmetrics2.2.2Rust580.2 ns1.72M/s25.67× slower
strsim0.11.1Rust580.4 ns1.72M/s25.68× slower
rapidfuzz0.5.0Rust620.1 ns1.61M/s27.44× slower

Verbora wins Hamming against every char-indexed competitor here from 16 characters up — a fixed per-call setup cost (still small in absolute terms, nanoseconds) makes Verbora the slowest of the four specifically at 4 characters, the one exception, before it pulls ahead and stays ahead. No bit-parallel state to build on either side, so this stays an apples-to-apples scalar-vs-scalar race outside that one small-input crossover. (triple_accel's genuinely SIMD-accelerated Hamming is a different story — see the byte-level subsection below, where Verbora loses decisively instead.)

Input sizeVerborarapidfuzzstrsimstringmetrics
43.8 ns5.7 ns4.6 ns2.4 ns
163.8 ns18.2 ns13.4 ns16.1 ns
644.6 ns50.4 ns43.4 ns42.1 ns
2567.4 ns162.5 ns143.4 ns142.6 ns
102422.6 ns620.1 ns580.4 ns580.2 ns

Jaro / Jaro–Winkler ​

LibraryVersionLanguageTime (median, 1024 chars)ThroughputRelative
Verbora0.2.0Rust8.55 µs116.9K/s1.00×
rapidfuzz0.5.0Rust15.35 µs65.2K/s1.79× slower
strsim0.11.1Rust332.31 µs3.0K/s38.85× slower

Verbora beats both competitors at every size. rapidfuzz's Jaro/Jaro-Winkler is bit-parallelized (rapidfuzz-0.5.0/src/distance/jaro.rs); Verbora matches that with its own bit-parallel match-flagging kernels (word-sized plus multi-word block) in its own greedy match orientation, with the scalar loop retained for inputs of at most 16 units and as the differential-test oracle, and the fractional-transposition semantics preserved exactly. Against rapidfuzz on Jaro-Winkler: 2.09× faster at 4 characters, 1.14× at 16, 2.04× at 64, 1.17× at 256, 1.79× at 1024.

Input sizeVerborarapidfuzzstrsim
413.0 ns32.3 ns27.2 ns
1674.6 ns84.9 ns155.6 ns
64136.8 ns278.9 ns1.46 µs
2561.86 µs2.18 µs24.10 µs
10248.55 µs15.35 µs332.31 µs

Sørensen–Dice against a Rust crate, and plain Jaro against a JavaScript library, are not benchmarked (plain Jaro against the Rust competitors above is — see the table just before this paragraph) — the matrix records both as narrowed/no-fair-competitor per docs/COMPETITIVE_BENCHMARKS.md §1.8. See benches/distance.rs's own module doc comment for the full accounting of every row this module benchmarks and every row it deliberately does not.

triple_accel and editdistancek — byte-level, ASCII-only-fair ​

Kept separate from every table above rather than merged into them: triple_accel and editdistancek operate on raw &[u8] bytes, not Unicode scalars like Verbora/strsim/rapidfuzz/stringmetrics — numerically identical to the char-indexed approach on the ASCII-only corpus this whole module shares, but genuinely different on non-ASCII input, which the research matrix marks Partial/Selected cases rather than the full Yes/Yes equivalence every row above carries. triple_accel is genuinely SIMD-accelerated (AVX2/SSE4.1); editdistancek is a Myers-style banded/diagonal algorithm over isize buffers. Byte-identical correctness against Verbora verified in tests/distance_correctness.rs before any number below was trusted — with one deliberate exception. The restricted-Damerau equality is asserted only on the shared corpus and not on randomized input, because triple_accel's restricted Damerau carries a real, independently-confirmed upstream defect (see Upstream bugs found) that the corpus happens never to trigger. The timing row stays: both sides do the same shape of work on the pairs actually benchmarked. The correctness claim is scoped to match.

Levenshtein — the same bit-vector kernels above win here too, by wider margins than against rapidfuzz/strsim:

LibraryVersionTime (median, 1024 chars)Relative
Verbora0.2.027.72 µs1.00×
triple_accel0.4.0527.10 µs19.01× slower
editdistancek1.0.21.07 ms38.46× slower

Random pairs:

Input sizeVerboratriple_acceleditdistancek
410.9 ns102.5 ns46.3 ns
1641.2 ns339.2 ns370.0 ns
64160.8 ns1.67 µs4.93 µs
2562.20 µs35.43 µs78.63 µs
102427.72 µs527.10 µs1.07 ms

Near-identical pairs (d = 1), where editdistancek's banded algorithm is at its strongest and triple_accel's SIMD full-matrix pass gains nothing:

Input sizeVerboratriple_acceleditdistancek
410.5 ns101.9 ns22.3 ns
1611.7 ns336.3 ns23.5 ns
6413.7 ns1.64 µs73.5 ns
25620.0 ns37.33 µs104.9 ns
102441.5 ns525.17 µs192.3 ns

Verbora wins at every size and both shapes, outright. The near-identical rows are worth reading beside the random ones, because they separate the two competitors rather than confirming the first table. triple_accel is unmoved by similarity — it runs the same vectorized full matrix either way, so 1024 costs it ~525 µs regardless — while editdistancek's banded algorithm is built for exactly this case and drops to 192.3 ns. Verbora is faster still at 41.5 ns, but against a competitor doing the right thing rather than one doing the wrong thing quickly.

Restricted Damerau-Levenshtein — Verbora's OSA bit-parallel kernels win at every size, against triple_accel's byte-level rdamerau (35.97 µs vs. triple_accel's 782.29 µs at 1024 characters, 21.75× faster).

Hamming — the one byte-level case where Verbora loses: triple_accel's Hamming is a vectorized XOR-and-popcount over the whole string with no data-dependent branching, versus Verbora's scalar per-position comparison loop. triple_accel is 2.05× faster at 1024 characters (11.0 ns vs. Verbora's 22.6 ns), a gap that widens with length: Verbora is actually faster at 16 characters (3.8 ns vs. 4.5 ns) and only slightly behind at 4 and 64, before triple_accel's vectorization advantage compounds at longer input.


Tokenizers ​

WordTokenizer and SentenceTokenizer both perform full UAX #29 segmentation, and both are now measured against their Rust rivals on that current implementation.

Word tokenization — a genuine size-dependent crossover ​

WordTokenizer against tantivy 0.26.1's SimpleTokenizer and Hugging Face tokenizers 0.23.1's Whitespace pre-tokenizer, called in isolation (never through HF's full BPE pipeline). The workload is narrowed to punctuation-free ASCII text, and boundary-exact agreement (not just token-count agreement) is proved against the current implementation in tests/tokenizers_correctness.rs before any timing is trusted. A verbora-lazy variant (an iterator that yields tokens without collecting them into a Vec first) runs alongside the default, allocating Vec-returning WordTokenizer.

LibraryVersionLanguageTime (median, 311,023 B)ThroughputRelative
tantivy0.26.1Rust488.18 µs2.0K/s1.00×
Verbora0.2.0Rust488.67 µs2.0K/s1.00× slower
Verbora (lazy)0.2.0Rust491.36 µs2.0K/s1.01× slower
huggingface0.23.1Rust8.70 ms114.9/s17.82× slower
Input (bytes)VerboraVerbora (lazy)tantivyhuggingface
123211.9 ns159.9 ns128.3 ns2.15 µs
566790.9 ns631.2 ns563.6 ns8.94 µs
1,1871.56 µs1.20 µs1.15 µs18.63 µs
4,7515.83 µs4.90 µs4.64 µs78.48 µs
9,70912.57 µs11.86 µs11.16 µs169.84 µs
38,76460.04 µs59.32 µs60.52 µs981.43 µs
77,684120.95 µs119.16 µs127.50 µs1.87 ms
311,023488.67 µs491.36 µs488.18 µs8.70 ms

tantivy's SimpleTokenizer wins at small-to-medium input (up to ~1.4× at the shortest text), and Verbora's lazy iterator overtakes it from roughly 38,000 bytes up, with the two effectively tied at 311,023 bytes. Both beat Hugging Face's Whitespace pre-tokenizer by more than an order of magnitude at every size — that pre-tokenizer is not the crate's fast path; it exists as one stage of a full BPE pipeline, and calling it standalone times overhead a real HF user would not pay in isolation. Full UAX #29 word segmentation does strictly more work per boundary than a character-class scan, so this crossover is the honest cost of that correctness: Verbora wins on long input despite doing more per byte, and loses on short input where per-call setup dominates. See PERFORMANCE_GAPS.md entry 4.

Sentence tokenization ​

SentenceTokenizer against segtok 0.1.5, on the narrowed plain-declarative-sentence domain both sides agree on (no abbreviations/URIs/digits/quotes/brackets), with boundary-exact agreement proven in the same correctness test before any timing is trusted.

LibraryVersionLanguageTime (median, 474,752 B)ThroughputRelative
Verbora (lazy)0.2.0Rust5.39 ms185.7/s1.00×
Verbora0.2.0Rust5.42 ms184.5/s1.01× slower
segtok0.1.5Rust104.91 ms9.5/s19.48× slower

Verbora wins at every size measured, by roughly 14.7×–20.8× depending on input length, whether or not the lazy iterator is used. SentenceTokenizer is built directly on split_sentence_bound_indices() with no placeholder mask, no unmask pass and no trimming.

The unicode-segmentation pairing does not appear here, and for a reason worth stating plainly: SentenceTokenizer is built on unicode-segmentation and WordTokenizer is str::unicode_words(), so timing either against its own dependency measures Verbora against the primitive it delegates to, which is a wrapper-overhead question rather than a competitive one. Those rows live in the *_wrapper_overhead groups below instead, and are never reported as Verbora beating or losing to unicode-segmentation.

Wrapper overhead — not ranked comparisons ​

GroupVerbora (default)Verbora (lazy)Primitive(s)
Word tokenization, 311,023 B526.31 µs486.69 µsunicode-words 492.65 µs, unicode-bounds 2.46 ms
Sentence tokenization, 474,752 B5.42 ms5.41 msunicode-bounds 5.48 ms, unicode-sentences 5.57 ms

Both wrappers cost roughly 0–8% over the bare unicode-segmentation call they make, well within this run's measurement noise at most sizes — the default (Vec-collecting) WordTokenizer pays the most at short input (up to ~1.5× the lazy iterator's cost at 123 bytes, from the extra allocation), converging with the lazy variant as input grows. Neither row is a rival-implementation comparison: unicode-words/unicode-bounds/ unicode-sentences are the exact primitives WordTokenizer/ SentenceTokenizer call.


N-Grams ​

Character n-gram generation with frequency counting (docs/COMPETITIVE_BENCHMARKS.md§1.2) against ngrammatic 0.7.0's Ngram/NgramBuilder — the character n-gram + frequency-count generator its headline Corpus/search fuzzy-matching feature is itself built on. Only that generator is benchmarked here: Corpus/search solves a different problem (fuzzy corpus search) with no Verbora equivalent, and is not compared. Both sides pad with arity - 1 copies of the same character and slide an identical window over every word in the shared 20,000-word list; byte-identical (gram, count) output is proven at arity 2 and arity 3 in tests/ngrams_correctness.rs before any number below was trusted.

Verbora wins both arities. verbora-ngrams's ngrams yields borrowed sub-slices of the caller's own slice, with nothing allocated per window.
LibraryVersionLanguageTime (median, bigrams, 20,000 words)ThroughputRelative
Verbora0.2.0Rust7.87 ms127.0/s1.00×
ngrammatic0.7.0Rust10.32 ms96.9/s1.31× slower
LibraryVersionLanguageTime (median, trigrams, 20,000 words)ThroughputRelative
Verbora0.2.0Rust10.12 ms98.8/s1.00×
ngrammatic0.7.0Rust11.16 ms89.6/s1.10× slower

Verbora wins bigram generation by 1.31× and trigram generation by 1.10× — a clean win at both arities. Both implementations do the same conceptual work — pad, slide a window, fold into a (gram, count) map — over the same input, so the residual reflects each side's small-string-accumulation strategy rather than an algorithmic difference: ngrammatic accumulates directly into a HashMap<SmolStr, usize>, whose small-string optimization skips a heap allocation for any gram that fits inline, while Verbora's borrowed-window ngrams() engine now avoids building an owned Vec/String per gram altogether.


Stemmers ​

Nine canonical Snowball-algorithm languages (de es fr it nl no pt ru sv) against two independent Snowball-to-Rust ports: rust-stemmers 1.2.0 (the official Snowball compiler's own output, and the Rust ecosystem's de-facto Snowball crate) and snowball_stemmers_rs 1.0.1 (a second, independently-generated port from the same compiler, published by the original SymSpell author). English is benchmarked separately against nltk-porter 0.1.0 and porter-stemmer 0.1.2: both competitors' "English" is a documented different algorithm from rust-stemmers' Snowball Porter2 and matches Verbora's original-1980 Porter instead (docs/COMPETITIVE_BENCHMARKS.md§1.3). Byte-exact agreement on every benchmarked word verified in tests/stemmers_correctness.rs — including several real, narrow divergences found and excluded from the benchmarked domain rather than hidden (Russian ё→е folding, Dutch's sticky cross-call state, porter-stemmer's single isolated "sky"→"ski" bug).

Verbora's suffix matching is the Snowball runtime's own find_among/find_among_b binary search (crates/verbora-stemmers/src/among.rs): a table sorted by reversed scalar sequence, common_i/common_j prefix tracking so no unit is compared twice, and substring_i-style links so one search replaces a whole guarded else-if chain. Ten of the eleven language modules route through it — de, en, es, fr, it, nl, no, pt, ru, sv. That is the same algorithm both competitors get from the official Snowball compiler.

LanguageVerbora (median, 1024 chars)rust-stemmerssnowball_stemmers_rsVerbora advantage (best competitor)
German (de)97.36 µs137.30 µs187.35 µs1.41×
Spanish (es)73.80 µs130.04 µs94.02 µs1.27×
French (fr)109.19 µs192.25 µs156.98 µs1.44×
Italian (it)94.49 µs228.51 µs196.43 µs2.08×
Dutch (nl)93.21 µs236.45 µs97.52 µs1.05×
Norwegian (no)45.02 µs48.61 µs33.51 µs1.34× slower
Portuguese (pt)115.05 µs174.15 µs94.45 µs1.22× slower
Russian (ru)122.61 µs82.82 µs81.02 µs1.51× slower
Swedish (sv)53.53 µs52.27 µs41.25 µs1.30× slower

Verbora wins five of nine languages outright (German, Spanish, French, Italian, Dutch) and loses four (Norwegian, Portuguese, Russian, Swedish) — to snowball_stemmers_rs in three of those four cases and to rust-stemmers in the fourth (Russian). Neither competitor is uniformly faster than the other across the four Verbora loses; this is a genuine per-language split rather than a single systematic gap.

snowball_stemmers_rs — a second, independently-generated Snowball port ​

Languages are never averaged together, so this is a second, independent data point rather than a repeat of the rust-stemmers comparison. What it establishes without reference to any measurement is a correctness finding: snowball_stemmers_rs's russian.sbl carries the same ё→е fold Verbora's stemmer does, so Russian agrees 100% byte-exact including ёлка, where rust-stemmers does not. Dutch needs Algorithm::DutchPorter specifically; the crate's plainly-named Algorithm::Dutch is actually Kraaij–Pohlmann, a different, non-canonical stemmer, confirmed by reading the crate's own algorithm list rather than assumed from the name.

English — nltk-porter and porter-stemmer ​

Two independent original-1980-Porter ports, since rust-stemmers' own "English" is Snowball Porter2, a different algorithm (excluded from the Snowball comparison above). Verbora's English module routes through among.rs like the other ten.

LibraryVersionLanguageTime (median, 1024 chars)ThroughputRelative
Verbora0.2.0Rust211.44 µs4.7K/s1.00×
porter-stemmer0.1.2Rust318.60 µs3.1K/s1.51× slower
nltk-porter0.1.0Rust1.87 ms534.6/s8.85× slower

Verbora wins at every size measured against both competitors, by 1.10×–1.50× against porter-stemmer and 7.82×–10.54× against nltk-porter.

The correctness finding stands on its own: porter-stemmer operates on grapheme clusters rather than Unicode scalar values, an architectural difference that turns out not to matter on this plain-ASCII corpus — 63 of 64 benchmarked words agree byte-exact, the one mismatch a real, isolated porter-stemmer bug ("sky"→"ski"), unrelated to graphemes and excluded from the sample.

Japanese — lindera-analysis ​

Verbora's StemmerJa (trailing katakana U+30FC drop, minimum 4 Unicode scalar values) against lindera-analysis's JapaneseKatakanaStemTokenFilter, min = 3 (the filter's own default) — verified to reproduce Verbora's >= 4-scalar threshold exactly on the shared word list before any number below was trusted.

Input sizeVerboralindera-analysisFaster
426.7 ns454.4 nsVerbora, 17.00×
10248.29 µs138.25 µsVerbora, 16.68×

A clean, decisive win at every size, 15.8×–20.3× depending on length — Verbora's algorithm borrows and allocates nothing, while lindera-analysis's Vec<Token>-batch filter API always allocates at least the Vec, on top of running through a full dictionary-backed tokenizer pipeline Verbora's purpose-built stemmer doesn't need.

Indonesian — sastrawi ​

Verbora's StemmerId (Sastrawi/Nazief–Adriani) against sastrawi (iDevoid/rust-sastrawi) 0.1.1 — genuine shared lineage, not a coincidence: both implement the same published PHP Sastrawi algorithm, and both dictionaries hold exactly 29,932 root words, confirmed directly. Real correctness gaps found in sastrawi during verification, excluded from the benchmarked sample rather than hidden: no hyphenated-reduplication/ compound-plural handling at all, and only a single (not iterated-up-to-3×) prefix-stripping pass — 13 of 16 benchmarked words still agree byte-exact.

Input sizeVerborasastrawiFaster
41.08 µs879.9 nssastrawi, 1.23×
1024631.63 µs617.35 µssastrawi, 1.02×
A real, narrow loss on per-word time (1.02×–1.15× across the sizes measured). sastrawi's own one-time Dictionary::new() + Stemmer::new() construction cost is real but paid once; Verbora's StemmerId::new() is a zero-sized unit struct backed entirely by compiled-in static data, needing no runtime construction at all — a trade-off the per-word numbers above don't capture.

CarryStemmerFr (French, Carry variant — a distinct 3-pass suffix-table algorithm, not standard Snowball French) has no fair Rust competitor: no crate in the ecosystem implements it, confirmed by checking that rust-stemmers' Algorithm::French is standard Snowball French rather than Carry.


Normalizers ​

remove_diacritics (docs/COMPETITIVE_BENCHMARKS.md§1.4) against diacritics 0.2.2, chosen over five other Rust candidates as the closest semantic match — case-preserving, non-decomposing table lookup, not NFD-based like unaccent and not forced-lowercasing like secular. Byte-exact agreement on ASCII, precomposed accented Latin, Cyrillic rejection and the shared ß/ẞ/ſ/İ/ı/œ table quirks verified in tests/normalizers_correctness.rs before any timing was trusted — one real divergence found and deliberately excluded from the benchmarked domain: diacritics silently strips standalone Unicode combining marks, which remove_diacritics never decomposes and so leaves untouched.

Pure-ASCII input ​

LibraryVersionLanguageTime (median, 1024 B)ThroughputRelative
Verbora0.2.0Rust39.6 ns25.27M/s1.00×
diacritics0.2.2Rust10.42 µs96.0K/s263.28× slower
Input sizeVerboradiacriticsVerbora vs. diacritics
42.9 ns95.5 ns32.78× faster
166.3 ns170.5 ns27.20× faster
648.9 ns700.0 ns78.44× faster
25623.4 ns2.66 µs113.92× faster
102439.6 ns10.42 µs263.28× faster

Verbora's one-line s.is_ascii() fast path returns Cow::Borrowed immediately — no scan, no allocation — while diacritics::remove_diacritics folds through its char match unconditionally, even on input it will not change.

Accented (working) input ​

Input sizeVerboradiacriticsVerbora vs. diacritics
43.05 µs492.1 ns6.21× slower
1612.90 µs1.89 µs6.83× slower
6448.15 µs7.91 µs6.09× slower
256201.98 µs31.60 µs6.39× slower
1024746.16 µs121.39 µs6.15× slower
A real, disclosed loss on accented input, consistent across every size measured. Verbora's remove_diacritics decomposes accented characters (its ASCII fast path above is the payoff of that design on input that needs no decomposition); diacritics uses a direct table lookup with no decomposition step, which is faster whenever the input actually contains accented characters to remove. See PERFORMANCE_GAPS.md for this module's tracked entries.

Inflectors ​

OrdinalInflector::nth — English ordinal suffixing (1st/2nd/3rd/…) — against ordinal 0.4.0 (docs/COMPETITIVE_BENCHMARKS.md§1.5), the single full-equivalence (Yes) row in the whole Inflectors group. Two real divergences were found and excluded from the benchmarked domain before any timing was trusted, in tests/inflectors_correctness.rs: negative integers (different rounding conventions), and a real bug in ordinal 0.4.0 itself: its teens exception uses n % 20 where it needs n % 100, misformatting 12% of non-negative integers (31.to_ordinal_string() returns "31th", not "31st"). The benchmarked domain verifiably avoids every affected value.

LibraryVersionLanguageTime (median, 1024-int batch)ThroughputRelative
Verbora0.2.0Rust14.56 µs68.7K/s1.00×
ordinal0.4.0Rust25.89 µs38.6K/s1.78× slower
Input sizeVerboraordinalVerbora vs. ordinal
453.7 ns97.7 ns1.82× faster
16223.2 ns419.4 ns1.88× faster
64838.7 ns1.58 µs1.88× faster
2563.72 µs6.94 µs1.87× faster
102414.56 µs25.89 µs1.78× faster

A flat ~1.8× regardless of batch size: both sides are O(1) per integer, so this is a genuine constant-factor difference — ordinal formats through Display/format!, Verbora writes digits directly into a pre-sized buffer.

NounInflector (pluralize/singularize) against inflector 0.11.4 and pluralizer 0.5.0:

OperationVerbora (median, 1024 words)inflectorpluralizer
pluralize226.97 µs463.95 µs (2.04× slower)603.21 µs (2.66× slower)
singularize311.85 µs735.69 µs (2.36× slower)701.67 µs (2.25× slower)

Verbora wins both operations against both competitors at every size measured (4 to 1024 words), by roughly 1.3×–2.7× depending on operation and size.


Trie ​

Generic prefix-search throughput. Verbora's trie keys one node per Unicode scalar and enumerates in ascending scalar order, which for well-formed Rust strings is <str as Ord>; competitors key by byte or by nybble and order their own way, so the comparison is timed on equal work rather than asserted element-for-element.

The competitors are trie-rs 0.4.2 (highest download count of any competitor in the whole audit: 5.9M) and qp-trie 0.8.2 (docs/COMPETITIVE_BENCHMARKS.md§1.18). Order-blind set-equality of every operation's result proven in tests/trie_correctness.rs before any timing was trusted; "build" is timed as push-then-compile for trie-rs's LOUDS architecture, matching how that crate is actually used. insert_all clamps the reservation it takes from an iterator's size_hint at 4,096 nodes and reaches the rest by amortised growth — a size_hint is a hint, not a bound, and an iterator that overstates it could otherwise turn a bulk load into an unbounded allocation.

LibraryVersionLanguageTime (median, 20,000-word build, random)ThroughputRelative
Verbora0.2.0Rust2.13 ms469.0/s1.00×
qp-trie0.8.2Rust3.05 ms328.1/s1.43× slower
trie-rs0.4.2Rust10.95 ms91.3/s5.14× slower

Verbora wins build against both competitors (1.43×–5.14× at random keys; similar margins at prefix-heavy and sorted-key shapes) and wins every operation against trie-rs by roughly two orders of magnitude. Against qp-trie specifically, the read path inverts:

OperationVerboraqp-trieVerdict
contains (hit, 20K words)469.74 µs855.31 µsVerbora 1.82× faster
contains (miss, 20K words)417.62 µs772.24 µsVerbora 1.85× faster
common_prefix_search256.74 µs— (not implemented)—
predictive_search (1-char prefix)286.3 ns118.76 µsVerbora 414.9× faster
predictive_search (empty prefix, all 20K)1.15 ms123.88 µsqp-trie 9.28× faster
A split read path. Verbora's arena trie wins contains and single-character predictive_search outright against qp-trie; only full-corpus enumeration (empty-prefix predictive_search) favours qp-trie's path-compressed radix structure, which stores each key whole at its leaf and pays no per-scalar reconstruction cost when enumerating everything. See PERFORMANCE_GAPS.md entry 2 for the underlying mechanism.

common_prefix_search has no qp-trie competitor: that crate implements no "stored words that are prefixes of a query" operation at all.

fast_radix_trie — a path-compressed radix map, and Verbora's own FrozenTrie answer ​

fast_radix_trie 1.2.0 is a path-compressed radix map created 2025-10-30. It uses unsafe internally (dynamically-sized nodes via raw pointers, miri-tested per its own docs); Verbora's own trie has zero unsafe anywhere.

OperationVerborafast_radix_trieVerdict
build (random)2.13 ms2.20 msVerbora 1.03× faster
contains (hit, sorted keys)559.74 µs1.02 msVerbora 1.82× faster
contains (miss, first-char probe)49.45 µs175.53 µsVerbora 3.55× faster
predictive_search (1-char prefix)286.3 ns500.20 µsVerbora 1,748× faster
predictive_search (empty prefix, all 20K)1.15 ms—— (not part of this group's competitor set)

Verbora's arena trie wins build and every measured operation against fast_radix_trie — including single-character prefix enumeration, despite fast_radix_trie's own path compression.

Verbora's own answer: FrozenTrie.Trie::freeze() is a safe-Rust (zero unsafe), path-compressed, read-only representation built once from a Trie, offered for the realistic autocomplete shape where enumeration dominates. See Trie for its own usage and characteristics.

fst — a frozen finite-state transducer ​

fst 0.4.7 (Andrew Gallant's) is architecturally nothing like a trie: a finite-state transducer built once from sorted input, queried via a Streamer, never mutated again.

OperationVerborafstVerdict
build (random)2.13 ms6.80 msVerbora 3.19× faster
contains (hit)469.74 µs1.75 msVerbora 3.74× faster
predictive_search (1-char prefix)286.3 ns1.84 msVerbora 6,432× faster

fst's own Levenshtein automaton (via its levenshtein feature) answers the same fuzzy-candidate question as verbora_spellcheck::FuzzyIndex — a genuine double crossover on both construction and query:

WordsConstruction: FuzzyIndexConstruction: fstQuery: FuzzyIndexQuery: fst
10024.07 µs60.73 µs281.48 µs (d1)5.22 ms (d1)
1,000759.04 µs471.54 µs4.00 ms (d1)12.95 ms (d1)
10,00011.45 ms3.89 ms24.78 ms (d1)14.78 ms (d1)
20,00026.90 ms6.98 ms40.62 ms (d1)20.30 ms (d1)

FuzzyIndex wins construction small, fst wins from 1,000 words up; the distance-1 query race is closer than before but keeps the same shape — FuzzyIndex wins small corpora, fst overtakes as the corpus grows large. fst is also one of three crates in this audit found to carry a real, independently-confirmed upstream defect — see Upstream bugs found.


Phonetics ​

Against rphonetic 3.0.6 — the one actively-maintained Rust crate covering the same phonetic-algorithm families, in the Apache commons-codec lineage, in a single crate (docs/COMPETITIVE_BENCHMARKS.md§1.6). Twelve encoder families are measured, in three regimes with three different equivalence claims, all verified in tests/phonetics_correctness.rs before any number was trusted:

  • Three variant encoders — throughput-only (Partial, never Yes).SoundEx, Metaphone and DoubleMetaphone implement Verbora's own documented variants (condense-before-drop Soundex, a documented Metaphone stage-ordering quirk), rphonetic the textbook originals — byte-exact output is never asserted, only that both sides do the same shape of work.
  • Eight byte-exact encoders — full output equivalence (Yes).Cologne, Nysiis, Caverphone1/Caverphone2, Phonex, RefinedSoundex, MatchRatingApproach and the branching DaitchMokotoff produce output byte-identical to rphonetic's on every input rphonetic handles without panicking — a stronger claim than any Partial row on this page carries.
  • BeiderMorse — throughput-only, Partial, and by far the heaviest encoder in the module. Covered in its own subsection below.

The three variant encoders — throughput only ​

rphonetic's Metaphone/Double Metaphone default to a 4-character max code length; both are reconfigured to Some(32) here to match Verbora's real default of 32 — independently verified by test to actually change rphonetic's output length, not silently still capped at 4.

Algorithm1 name10,000 names100,000 names
Soundex6.44× faster3.27× faster3.31× faster
Metaphone2.28× faster1.55× faster1.53× faster
Double Metaphone2.77× faster1.94× faster1.88× faster

Verbora wins all three algorithms at every benchmarked size.

LibraryVersionLanguageTime (median, 100,000 names)ThroughputRelative
Verbora0.2.0Rust5.04 ms198.2/s1.00×
rphonetic3.0.6Rust7.72 ms129.5/s1.53× slower
Metaphone — a clean sweep at every size. Verbora's Metaphone runs as a single skip-gated driver, fused from the original 21 ordered whole-string rewrite stages, over per-thread pooled scratch: letter-mask gates decide which rules can possibly fire on a given word, window edits plus fused rules replace whole-string rewrites, and the pipeline's two scratch buffers are reused across calls — an ASCII token folds lowercase directly into pooled scratch, so a steady-state call's only allocation is the returned code. rphonetic's Metaphone is a single indexed forward scan (O(n)); Verbora wins anyway: 2.28× at a single name (32.8 ns vs. 74.9 ns), 1.55× at 10,000 names (487.15 vs. 756.74 µs), 1.53× at 100,000 (5.04 vs. 7.72 ms). See PERFORMANCE_GAPS.md entry 6 for the full mechanism.

Full per-algorithm, per-size data for these groups: results/results.json (module "phonetics").

The eight byte-exact encoders ​

Verification for this table goes beyond the shape-parity check above, because here byte-exact equality is the claim: tests/phonetics_correctness.rs asserts identical output over the shared 653-name corpus plus per-algorithm extras for every encoder; for MatchRatingApproach it additionally checks the real MRA match decision (compare, not just the code) over every ordered pair of corpus names (~426K pairs); and Daitch–Mokotoff is checked three ways at once — the pipe-joined process string, the codes vector, and the first-branch code against rphonetic's non-branching encode. An independent adversarial audit then differentially fuzzed 104,114 inputs per encoder against rphonetic with zero mismatches, proved every documented divergence exactly as narrow as claimed, and mutation-tested the suites. The only divergences are inputs on which rphonetic itself panics — see Upstream bugs found; those input shapes are excluded from the benchmark domain per this page's fairness pattern, and the ASCII-only shared corpus never reaches them anyway. Configuration is identical on both sides: Nysiis runs strict (the commons-codec default), Phonex at its default max code length of 4.

Verbora is faster in all 24 cells (Verbora vs. rphonetic, Criterion medians):

Encoder1 name10,000 names100,000 names
Cologne17.0 ns vs. 66.8 ns (3.93×)318.24 µs vs. 765.23 µs (2.40×)3.32 ms vs. 7.24 ms (2.18×)
NYSIIS30.6 ns vs. 224.0 ns (7.31×)274.32 µs vs. 1.89 ms (6.90×)2.67 ms vs. 19.01 ms (7.12×)
Caverphone 1.0169.4 ns vs. 924.9 ns (5.46×)1.92 ms vs. 9.98 ms (5.19×)19.07 ms vs. 101.27 ms (5.31×)
Caverphone 2.0152.8 ns vs. 803.8 ns (5.26×)1.68 ms vs. 9.17 ms (5.47×)17.20 ms vs. 92.66 ms (5.39×)
Phonex30.4 ns vs. 133.0 ns (4.38×)337.95 µs vs. 1.20 ms (3.56×)3.64 ms vs. 12.42 ms (3.41×)
Refined Soundex18.5 ns vs. 116.9 ns (6.31×)151.28 µs vs. 973.22 µs (6.43×)1.48 ms vs. 9.76 ms (6.60×)
Match Rating Approach47.4 ns vs. 526.1 ns (11.10×)460.10 µs vs. 5.39 ms (11.72×)4.54 ms vs. 53.99 ms (11.90×)
Daitch–Mokotoff (branching)152.7 ns vs. 380.1 ns (2.49×)1.72 ms vs. 3.93 ms (2.29×)15.94 ms vs. 36.89 ms (2.31×)

The mechanism is consistent across all eight groups rather than one trick: Verbora's encoders run single-pass scans over one reused buffer, with one heap allocation per call for the returned code (the branching Daitch–Mokotoff adds a small branch list), against static compiled-in rule tables. rphonetic's implementations allocate intermediate Strings as they go (Caverphone's rewrite cascade is one freshly allocated String per step there) and, for Daitch–Mokotoff, parse the rules text with a nom grammar at builder time and walk a BTreeMap per lookup where Verbora indexes a pre-sorted static array. The margins range from 2.29× (Daitch–Mokotoff at 10,000 names — the one algorithm where both sides spend most of their time in the same branching walk) to 11.90× (Match Rating at 100,000).

The Daitch–Mokotoff row compares Verbora's branching DaitchMokotoff::process against rphonetic's own pipe-joined soundex() — output-format identical, which is what makes byte-exact equality the claim here rather than shape parity. Four rphonetic Daitch–Mokotoff behavioral quirks are reproduced deliberately for byte-parity and documented in crates/verbora-phonetics/src/daitch_mokotoff.rs's own module documentation. Raw Criterion estimates for these eight groups live in the same Criterion tree cargo bench writes — see Reproducing these numbers.

BeiderMorse — throughput only, and the heaviest encoder in the module ​

BeiderMorse (19-language auto-detecting Beider–Morse phonetic matching, with no equivalent in Verbora's other reference points) against rphonetic's BeiderMorseBuilder (features = ["embedded_bm"]). This is a coverage asymmetry, not a fully equivalent comparison: rphonetic's embedded_bm feature ships only the "any"/"common" rule files per NameType, not the full per-language corpus, so its ConfigFiles::default() can never resolve a specific guessed language and always falls back to "any" — Verbora's full 18/10/5-language (Generic/Ashkenazi/Sephardic) LangGuesser auto-detection has no equivalent on the rphonetic side to compare against. Throughput only; output equivalence is never asserted, since both sides are textbook-derived but independently implemented.

Read this table's scale column before comparing it to any other row on this page. BeiderMorse is dramatically heavier per call than every other encoder in this module — a single name costs Verbora 7.24 µs here, against 32.8 ns for Metaphone and 17.0 ns for Cologne on the same one-name input (roughly 220× and 426× respectively). This module's own benchmark caps BeiderMorse's sweep at 1,000 names rather than the 100,000 every other encoder reaches, because the algorithm's own per-name cost makes a 100,000-name sweep impractically slow on either side.
NamesVerborarphoneticVerbora advantage
17.24 µs17.66 µs2.44×
100881.35 µs1.66 ms1.88×
1,0009.71 ms20.60 ms2.12×

Verbora wins at every scale measured, by 1.68×–2.44× depending on size — narrower and noisier than the byte-exact encoders' margins above, consistent with both sides doing substantially more per-call work here (rule-table lookups across an auto-detected language guess) than the single-pass scans the rest of this module runs. Two additional input shapes are measured at 100 names: compound_100 (hyphenated/compound surnames, 2.24 ms vs. 8.16 ms, 3.65× faster) and prefixed_100 (surnames with a common prefix, 2.28 ms vs. 6.00 ms, 2.63× faster) — both still clear Verbora wins, at wider margins than the plain 100-name row, since rphonetic's per-call allocation cost scales with the extra rule-table backtracking these shapes trigger more than Verbora's does.

A second Double Metaphone implementation — C++, not Rust ​

Every other competitor on this page is a Rust crate. pixelglow/double_metaphone is different: a header-only C++11 implementation of Lawrence Philips' Double Metaphone algorithm, vendored into the workspace and compiled by build.rs, then called through a thin extern "C" shim — the only non-Cargo, non-Rust competitor benchmarked anywhere on this page. Every measured call crosses the Rust/C++ boundary once (one CString construction, one FFI call, two bounded buffer copies, one UTF-8 validation on the way back), and that cost is measured as part of the number, not subtracted out. This comparison did not run as part of this campaign; the figures below are carried forward from the last run that measured it and are not re-verified against the current commit.

Correctness here is Partial, not byte-exact: 584 of 653 real English surnames (89.4%) produce identical primary and secondary keys on both sides. The one confirmed, dominant rule difference: Verbora silences a trailing S whenever it is preceded by A or I, while the C++ library only silences a trailing S in the narrower pattern where I or Y is immediately followed by S then L — island, isle, carlisle. Both sides agree Isle should lose its S; they disagree on names like Davis, which keeps its S on the C++ side but loses it on Verbora's, encoding the same way Isle does. Neither reading is more correct; Double Metaphone's own published algorithm write-up never fully disambiguates this case.

LibraryVersionLanguageTime (median, 653 names)ThroughputRelative
Verbora0.2.0Rust47.23 µs21.2K/s1.00×
pixelglow/double_metaphone79dd226 (2014)C++1185.30 µs11.7K/s1.81× slower

Language detection ​

Statistical language detection over free text (docs/COMPETITIVE_BENCHMARKS.md§1.9) against lingua 1.8.0 (built with from_languages(), restricted to the 21-language overlap with Verbora, never its default 75) and whichlang 0.1.1 (13-language overlap, and — disclosed explicitly, not folded silently into the accuracy numbers — it cannot abstain: detect_language always returns a guess). A widely-used JavaScript NLP library has no general statistical language-detection module (verified from source, not assumed), so it does not appear here.

Verbora now ships three detector strategies, and two of them beat whichlang outright. WhatlangDetector remains the crate's default (best accuracy). HashedLinearDetector (a zero-allocation, stack-only linear model, opt-in behind the fast-language-detection feature) and FallbackDetector<HashedLinearDetector, WhatlangDetector> (the fast model as primary, deferring to WhatlangDetector only where it declines to judge) are both new, and both are measured here alongside the default.

Speed, by input length (English) ​

TierHashedLinearDetectorFallbackDetectorWhatlangDetector (default)whichlanglingua
short word (~6 B)44.2 ns23.55 µs23.28 µs62.6 ns48.36 µs
short phrase (~30 B)150.2 ns152.3 ns38.08 µs203.5 ns90.10 µs
sentence (~140 B)308.6 ns312.6 ns32.16 µs531.4 ns199.29 µs
paragraph (~500 B)2.91 µs2.80 µs103.35 µs6.00 µs608.82 µs
LibraryVersionLanguageTime (median, paragraph)ThroughputRelative
FallbackDetector⟨Hashed, Whatlang⟩0.2.0Rust2.80 µs357.4K/s1.00×
HashedLinearDetector0.2.0Rust2.91 µs343.6K/s1.04× slower
whichlang0.1.1Rust6.00 µs166.7K/s2.14× slower
WhatlangDetector (default)0.2.0Rust103.35 µs9.7K/s36.94× slower
lingua1.8.0Rust608.82 µs1.6K/s217.4× slower
Only at the single-word tier does the story flip.HashedLinearDetector alone is fastest everywhere, including word tier (44.2 ns, still beating whichlang's 62.6 ns). But FallbackDetector costs 23.55 µs at word tier — almost the full price of the default detector — because that is exactly the tier where its accuracy is weakest and it defers to WhatlangDetector most often. See the accuracy table below for why that trade exists.

Speed, by language (sentence tier) ​

LanguageHashedLinearDetectorwhichlanglinguaFaster
German299.2 ns523.2 ns208.53 µsVerbora, 1.75×
English305.7 ns504.8 ns174.03 µsVerbora, 1.65×
Spanish742.4 ns1.36 µs183.16 µsVerbora, 1.83×
French352.6 ns565.6 ns237.21 µsVerbora, 1.60×
Hindi3.06 µs527.8 ns5.25 µswhichlang, 5.80×
Italian335.0 ns561.3 ns236.51 µsVerbora, 1.68×
Japanese416.4 ns295.3 ns5.86 µswhichlang, 1.41×
Dutch316.7 ns577.2 ns259.56 µsVerbora, 1.82×
Portuguese330.6 ns585.1 ns249.21 µsVerbora, 1.77×
Russian2.35 µs403.2 ns73.81 µswhichlang, 5.83×
Swedish308.2 ns478.1 ns253.87 µsVerbora, 1.55×
Vietnamese793.9 ns467.0 ns222.65 µswhichlang, 1.70×
Chinese267.9 ns142.9 ns3.00 µswhichlang, 1.88×

HashedLinearDetector wins 9 of 13 languages — every Latin-script language in this set — and loses the four where a hashed-bucket linear model gets less signal per byte: Hindi, Japanese, Russian and Vietnamese. whichlang's own hand-tuned per-language feature weights hold an edge on those four.

Accuracy ​

13 languages (the triple overlap all detectors can be scored on identically) × 4 length tiers, sourced from the OHCHR UDHR Translation Project (public-domain UN text; full sourcing and per-tier extraction rule in datasets/README.md). Reproduced with cargo run --release --example language_accuracy, and re-scored as an executed test in crates/verbora-language/tests/default_detector.rs.

Detectorshort wordshort phrasesentenceparagraphOverall
lingua (21-language restricted)92.3% (12/13)100%100%100%98.1% (51/52)
WhatlangDetector (default) / FallbackDetector76.9% (10/13)100%100%100%94.2% (49/52)
whichlang (13-language, cannot abstain)69.2% (9/13)100%100%100%92.3% (48/52)
HashedLinearDetector53.8% (7/13)92.3% (12/13)100%100%86.5% (45/52)

FallbackDetector scores identically to the default WhatlangDetector on every tier — that equivalence is the point of composing it, not a coincidence — and it beats whichlang on both accuracy (94.2% vs. 92.3%) and speed (every tier except the single-word one, where it defers to the slower default). HashedLinearDetector alone trades 4 of 52 correct answers, concentrated entirely in the two hardest, shortest tiers, for being the fastest detector in this whole audit at every tier including that one. This is why the fastest detector is not the default: shipping it unqualified would understate exactly the cost that makes it fast.

WhatlangDetector wrapper overhead — not a ranked comparison ​

Isolates the cost of Verbora's own wrapper around whatlang::Detector — it is not "Verbora vs. whatlang," because WhatlangDetector literally constructs a whatlang::Detector and calls .detect() on it. The noisy ratios include a tier where the wrapper measured faster than the bare call it makes, which is structurally impossible as a real effect. Read as noise from a shared benchmark machine, not a finding.
TierVerbora (WhatlangDetector)whatlang (raw crate)Ratio
short word23.54 µs23.93 µs0.98×
short phrase39.94 µs39.93 µs1.00×
sentence30.32 µs30.28 µs1.00×
paragraph103.60 µs99.32 µs1.04×

Script detection ​

Verbora's detect_script against whatlang::detect_script 0.18.0 (docs/COMPETITIVE_BENCHMARKS.md§1.10) — a real, public, standalone function doing the same conceptual work (per-codepoint Unicode-range classification, majority vote), just over a wider set (25 scripts vs. Verbora's 10). A widely-used JavaScript NLP library has no script-detection module at all (verified from source).

TierVerborawhatlangVerbora advantage
short word8.8 ns37.0 ns4.2×
short phrase14.1 ns60.6 ns4.3×
sentence27.1 ns121.4 ns4.5×
paragraph250.0 ns1.12 µs4.5×
LibraryVersionLanguageTime (median, paragraph)ThroughputRelative
Verbora0.2.0Rust250.0 ns4.00M/s1.00×
whatlang0.18.0Rust1.12 µs889.4K/s4.50× slower

Verbora wins at every length tested, by a fairly steady ~4.2×–4.5×. By language (sentence tier), Verbora is faster in 9 of 13 and loses 4:

LanguageVerborawhatlangFaster
German43.0 ns132.3 nsVerbora, 3.08×
English27.3 ns126.1 nsVerbora, 4.62×
Spanish87.3 ns310.3 nsVerbora, 3.56×
French60.6 ns136.2 nsVerbora, 2.25×
Hindi2.96 µs188.9 nswhatlang, 15.68×
Italian43.7 ns134.6 nsVerbora, 3.08×
Japanese371.8 ns284.8 nswhatlang, 1.31×
Dutch29.2 ns131.1 nsVerbora, 4.50×
Portuguese35.3 ns142.4 nsVerbora, 4.04×
Russian1.14 µs140.2 nswhatlang, 8.14×
Swedish73.8 ns130.8 nsVerbora, 1.77×
Vietnamese547.5 ns159.8 nswhatlang, 3.43×
Chinese123.8 ns158.8 nsVerbora, 1.28×

Verbora's per-codepoint Unicode-range classifier loses on Hindi, Japanese, Russian and Vietnamese — the same four languages HashedLinearDetector loses in the language-detection table above, consistent with those scripts needing more per-codepoint classification work in Verbora's 10-script model than in whatlang's wider 25-script one.


Transliteration ​

Japanese kana→romaji, throughput only — against wana_kana 5.0.0 (docs/COMPETITIVE_BENCHMARKS.md§1.11), the only Rust kana↔romaji crate with real current adoption/maintenance found — every alternative investigated is scope-mismatched or effectively abandoned. Never an output-correctness comparison: wana_kana uses a doubled-vowel convention ("スーパー" → "suupaa") while Verbora uses modified Hepburn with macrons ("tōkyō") — a real, executed divergence proven in tests/transliteration_convention_diff.rs, not merely asserted.

RepeatsVerborawana_kanaVerbora advantage
1×758.4 ns1.87 µs2.5×
16×2.75 µs6.71 µs2.4×
256×46.49 µs100.26 µs2.2×
LibraryVersionLanguageTime (median, 1024×)ThroughputRelative
Verbora0.2.0Rust184.66 µs5.4K/s1.00×
wana_kana5.0.0Rust417.82 µs2.4K/s2.26× slower

Verbora wins at every size measured, 2.16×–2.46× depending on length.


POS tagging ​

verbora-tagger (Brill, transformation-based) against postagger 0.0.3 (a pretrained averaged-perceptron model, NLTK weights) and rust-bert 0.23.0's POSModel (a MobileBERT transformer pipeline) — the most widely adopted general Rust NLP crate found in this whole audit (254K downloads, 3,077 stars) (docs/COMPETITIVE_BENCHMARKS.md§1.16). Both are genuinely different algorithm classes — a trained classifier and a transformer forward pass, not rival implementations of the same rule table — so this is reported as a technique comparison, with cold start and steady state kept strictly separate.

This comparison has no Verbora side.verbora-tagger ships no lexicon: the English dictionary and rule set it carried through 0.2 were removed in 0.3 for licensing reasons, and the build-time packed table whose construction a cold-start figure would time went with them. Reinstating the comparison is a measurement-design decision before it is a benchmark run: the harness has to pick a lexicon and hand the same one to every side, or the two are not answering the same question. The competitor measurements below stand.

Cold start — everything needed before tagging one sentence ​

LibraryVersionLanguageTime (median)
postagger (parses a 5.6 MB weights file)0.0.3Rust109.18 ms
rust-bert (loads a ~94 MB MobileBERT checkpoint)0.23.0Rust151.58 ms

Steady state — per-call latency, tagger already constructed ​

LibraryVersionLanguageTime (median, 9 tokens)Time (median, 20 tokens)Time (median, batch of 8×9-tok)
postagger0.0.3Rust58.95 µs75.72 µs451.73 µs
rust-bert0.23.0Rust12.45 ms9.22 ms13.41 ms

What the two tables show on their own is the price of the technique each crate chose: a pretrained model must deserialize its weights before it can answer anything, and then spends a feature-weighted vote or a full transformer forward pass per token. A rule-based tagger pays neither — its construction cost is whatever reading its lexicon and rule set costs, and its per-token work is a dictionary probe plus one pass per rule. That is a structural difference, not a measured one, and this section deliberately puts no number on it until the Verbora configuration is defined again.

This is not an accuracy claim for either technique — see docs/COMPETITIVE_BENCHMARKS.md §1.16 for why both rows are Partial, not Yes; this audit makes no tagging-quality comparison for POS tagging.


Spellcheck ​

Spellcheck::corrections and ::is_correct against three genuinely different algorithms (docs/COMPETITIVE_BENCHMARKS.md§1.17): symspell 0.5.2 (precomputed deletion dictionary), harper-core 2.8.0 (FST + Levenshtein automaton; by far the most widely adopted standalone spellchecking crate found, 14,470 GitHub stars on its parent repo), and spellbook 0.4.2 (Hunspell affix-rule morphology). corrections returns Vec<Correction>, each entry carrying the candidate word alongside the frequency and edit distance behind its ranking; correction_words is the separate call for plain owned Strings where the ranking metadata isn't needed.

Verbora wins construction, membership testing and correction generation against all three competitors, at every size measured.

symspell and harper-core — same corpus as Verbora ​

Both loaded with the identical words.json corpus and per-word frequencies Verbora uses.

GroupCorpusVerborasymspellharper-core
construction (new)1005.25 µs388.13 µs (73.9× slower)74.04 µs (14.1× slower)
construction (new)20,0001.50 ms118.51 ms (78.9× slower)10.63 ms (7.1× slower)
is_correct (hit)20,00020.85 µs296.46 µs (14.2× slower)115.57 µs (5.5× slower)
corrections, distance 1100173.4 ns853.1 ns (4.9× slower)5.01 µs (28.9× slower)
corrections, distance 120,000190.6 ns920.5 ns (4.8× slower)37.40 µs (196.2× slower)
corrections, distance 21,000301.3 ns2.29 µs (7.6× slower)32.25 µs (107.1× slower)
corrections, distance 220,000642.5 ns3.34 µs (5.2× slower)335.64 µs (522.4× slower)

Verbora wins every row in this table, at every size measured. A verbora-borrowed variant (returning views rather than owned corrections) runs alongside verbora in every group above and is marginally faster still (e.g. 164.5 ns vs. 173.4 ns at distance 1, 100-word corpus).

spellbook — matched-workload timing only, not a fair ratio ​

Hunspell's .aff/.dic format has no concept of a flat frequency corpus — spellbook cannot load Verbora's corpus, and Verbora cannot load a Hunspell dictionary. Each side is timed on its own inputs: spellbook's hit and near-miss-typo cases against Verbora's own hit case, and four different spellbook typos against Verbora's own single typo case. This is a timing comparison of two different dictionaries doing conceptually the same job, never presented as a ratio. No Relative column below.
OperationLibraryVersionDictionaryTime (median)
check / is_correct, hitspellbook0.4.2real en_US Hunspell358.4 ns
check / is_correct, near-miss typospellbook0.4.2real en_US Hunspell3.15 µs
is_correct, hitVerbora0.2.0own 20,000-word corpus4.86 µs
suggest, "helo"/"korrect"/"wrold"/"beleive"spellbook0.4.2real en_US Hunspell4.79–8.20 ms
corrections, one typo (typo8)Verbora0.2.0own 20,000-word corpus178.7 ns

spellbook's check is a curated, bundled FST/hash lookup — sub-microsecond as expected for a fixed, pre-built dictionary — while its full affix-aware suggest costs milliseconds, the opposite trade-off from its own check.

fast_symspell — a second deletion-index crate, and Verbora's own answer to it ​

fast_symspell 0.1.10 is a second, independent SymSpell-family implementation. Its published metadata carries no linked repository, but its source is real and readable via crates.io's own tarball — a near-verbatim (confirmed line-for-line) fork of symspell 0.5.2 with three real deltas: ahash hashing, a triple_accel-backed verification pass (which carries its own real, independently-confirmed bug — see Upstream bugs found), and an rkyv zero-copy archived-load path.

GroupCorpusVerborafast_symspell
construction (FuzzyIndex)10023.97 µs361.09 µs (15.1× slower)
construction (FuzzyIndex)20,00025.48 ms110.84 ms (4.4× slower)
corrections, distance 1100173.4 ns817.7 ns (4.7× slower)
corrections, distance 120,000190.6 ns844.6 ns (4.4× slower)
corrections, distance 21,000301.3 ns2.19 µs (7.3× slower)
corrections, distance 220,000642.5 ns3.50 µs (5.4× slower)

Verbora wins both construction and correction generation against fast_symspell at every size measured.

Verbora's own answer: DeletionIndex.verbora_spellcheck::DeletionIndex is a SymSpell-style index built in-house, offered alongside the existing FuzzyIndex BK-tree rather than replacing it. Its own head-to-head figures against FuzzyIndex are not part of this campaign — see the note below.

FuzzyIndex vs. DeletionIndex — not part of this campaign ​

No fresh measurement exists for this comparison.DeletionIndex's internal map was changed to key on a 64-bit hash of each deletion sequence rather than the sequence itself, taking the cost of indexing one word from cubic to quadratic in its length — but this campaign's benchmark run did not include a FuzzyIndex-vs- DeletionIndex group, so no current figures exist to publish here. Neither structure replaces the other — FuzzyIndex stays the default (cheaper, more predictable, no build-time distance ceiling); DeletionIndex is offered for a large dictionary with max_distance known ahead of time and high query volume. A timing comparison between them awaits a future run.

TF-IDF ​

Corpus build/ingestion and query/scoring (docs/COMPETITIVE_BENCHMARKS.md§1.12) against tfidf (afshinm) 0.3.0 (a stateful add()/idf()/tfidf() struct — the architecturally closest Rust match found, but a genuinely different, unsmoothed weighting formula) and rust-tfidf 1.1.1 (query/scoring only — it has no ingestion step, nothing to time as "build"). Both comparisons are explicitly build/query speed only — neither crate's output values are compared against Verbora's, since the weighting formulas differ by design.

LibraryVersionLanguageTime (median, build, 256 docs)ThroughputRelative
afshinm0.3.0Rust59.50 ms16.8/s1.00×
Verbora0.2.0Rust140.05 ms7.1/s2.35× slower
DocsVerboraafshinm
42.11 ms806.41 µs
168.47 ms3.03 ms
6435.55 ms15.03 ms
256140.05 ms59.50 ms
A real, disclosed ingestion loss — with the matching query-time win it buys, shown right below. See PERFORMANCE_GAPS.md entry 13: tfidf's add() is a single space-split pass with zero allocation, no lowercasing, no real tokenizer, no stop-word filtering. Verbora's add_document runs its own full pipeline — lowercasing, real word-boundary tokenization, stop-word filtering, interning — because that is what its own behaviour contract and its own O(1) query-time payoff (below) require.

A second construction shape, build_many_small — many short, per-document ~200-word chunks rather than a few large documents — isolates per-document overhead (interner/document-frequency bookkeeping for Verbora, one push per document for afshinm):

DocsVerboraafshinmRelative
413.69 µs3.26 µs4.20× slower
64214.31 µs57.40 µs3.73× slower
256853.75 µs224.31 µs3.81× slower
10243.42 ms1.19 ms2.87× slower

afshinm wins this shape too, by a narrower and fairly stable margin (2.9×–4.2×) than the few-large-documents shape above.

LibraryVersionLanguageTime (median, tfidf() query, 256 docs)ThroughputRelative
Verbora0.2.0Rust50.5 ns19.82M/s1.00×
rust-tfidf1.1.1Rust1.49 µs672.7K/s29.5× slower
tfidf (afshinm)0.3.0Rust248.97 ms4.0/s~4.9M× slower
DocsVerboraafshinmrust-tfidf
450.4 ns4.95 ms52.0 ns
1650.5 ns17.29 ms101.7 ns
6450.4 ns64.03 ms375.3 ns
25650.5 ns248.97 ms1.49 µs

Verbora's query cost is flat regardless of corpus size (the interned, incrementally-maintained document-frequency table this crate's own build cost pays for); both competitors rescan the whole corpus on every query, so their cost grows linearly with it. idf() shows the same pattern (module "tfidf", group "idf", in results.json). This is the same trade-off in both directions: expensive-but-thorough ingestion buying near-free, corpus-size-independent queries — not a one-sided result either way.


Classifiers ​

BayesClassifier training and prediction against smartcore 0.6.5's MultinomialNB (by far the most downloaded classifier candidate found, 476K downloads, actively maintained) and naivebayes 0.1.2 (ruivieira) — a pre-tokenized, fixed-smoothing-floor Naive Bayes implementation (docs/COMPETITIVE_BENCHMARKS.md§1.13). linfa-bayes no longer appears in the timing rows below: its published fit_with calls an unconditional dbg! once per class on every training call, which turned a training-loop benchmark into gigabytes of stderr output — a defect in the published crate, not something this harness works around. linfa-bayes stays in the accuracy comparison below, where it is called far less often, and in linfa-logistic's own, unrelated rows further down, which are unaffected.

LibraryVersionLanguageTime (median, train, 1024 docs)ThroughputRelative
naivebayes0.1.2Rust935.45 µs1.1K/s1.00×
smartcore0.6.5Rust1.58 ms631.3/s1.69× slower
Verbora0.2.0Rust2.06 ms484.5/s2.21× slower
DocsVerborasmartcorenaivebayes
417.75 µs5.60 µs7.67 µs
1658.79 µs24.12 µs29.19 µs
64197.03 µs92.40 µs89.84 µs
256639.89 µs337.91 µs292.40 µs
10242.06 ms1.58 ms935.45 µs
LibraryVersionLanguageTime (median, predict)ThroughputRelative
Verbora0.2.0Rust1.78 µs560.7K/s1.00×
naivebayes0.1.2Rust3.08 µs324.4K/s1.73× slower
smartcore0.6.5Rust3.93 µs254.5K/s2.20× slower
A genuinely mixed result. Verbora loses training at every size (2.2×–3.2× slower than the faster of the two competitors, narrowing with corpus size) but wins prediction against both — the one row of this section where Verbora comes out fastest. Verbora's per-document training cost includes real, specified tokenization, Porter stemming and stop-word filtering; neither competitor's benchmark adapter does any of that.

Logistic Regression ​

LogisticRegressionClassifier against smartcore 0.6.5's linear::logistic_regression, linfa-logistic 0.8.1, and rustlearn 0.5.0 (SGD-based, unmaintained since 2018, included for historical prominence and flagged as stale).

DocsVerborasmartcorelinfa-logisticrustlearn
424.87 µs284.69 µs71.40 µs4.64 µs
855.43 µs381.38 µs96.84 µs10.53 µs
12103.68 µs489.53 µs144.64 µs14.61 µs
16141.55 µs645.95 µs127.48 µs22.31 µs
Verbora now beats both smartcore and linfa-logistic at almost every size tested. Against smartcore, Verbora wins at every size, by a widening margin as corpus size shrinks (4.6×–11.4×). Against linfa-logistic, Verbora wins at 4, 8 and 12 documents (up to 2.9× faster) and loses narrowly only at 16 (1.11× slower) — a real crossover, but one that now favours Verbora through nearly the whole range measured. rustlearn's single-epoch SGD remains the fastest at every size (5.4×–7.1× faster than Verbora), doing asymptotically less work than every iterate-to-convergence competitor.
LibraryVersionLanguageTime (median, prediction)Relative
rustlearn0.5.0Rust362.1–406.4 ns1.00×
linfa-logistic0.8.1Rust426.8–436.7 ns1.08×–1.18× slower
smartcore0.6.5Rust535.5–593.3 ns1.46×–1.48× slower
Verbora0.2.0Rust730.9–814.6 ns1.80×–2.25× slower

Verbora loses single-document prediction to all three competitors, by 1.8×–2.25×.

Accuracy: is the slower classifier at least more correct? ​

A separate, signal-bearing corpus (four non-overlapping topical vocabularies, generated by tools/bench-data/generate.py) was built specifically because the training corpus above is shape-only random data — useless for accuracy. cargo test --test classifiers_accuracy trains Verbora, smartcore and linfa-bayes at each size and scores them against a fixed, disjoint 128-document test set:

Train sizeVerborasmartcorelinfa-bayes
498.4%93.0%93.0%
16100.0%100.0%100.0%
64100.0%100.0%100.0%
256100.0%100.0%100.0%
1024100.0%100.0%100.0%

All three converge to a perfect score by 16 training documents. Read alongside the speed table above: accuracy is statistically indistinguishable between the three at every size that matters, so the speed numbers stand as measured, not offset by a quality difference that is not actually there on this test set.


Sentiment ​

SentimentAnalyzer's AFINN-based document scoring against sentiment 0.1.1 (mount-research) — the only Rust crate found that scores text against an AFINN-family lexicon, published once in 2017 with no later release.

A narrowed comparison, over a corpus built specifically to make it fair. sentiment embeds the older, smaller AFINN-111 (2,462 entries); Verbora ships AFINN-165 (3,382 entries) and has no negation-free mode. Rather than compare the two lexicons on arbitrary text — which would report a lexicon difference as a speed difference — the benchmarked corpus is drawn from the 2,438-word intersection where the two tables agree exactly (of AFINN-111's 2,462 keys, all but 4 are also in AFINN-165 with the same polarity), excludes the four words in Verbora's negation list (not/no/never/ neither, which sentiment does not implement at all), and uses only lowercase ASCII words joined by single spaces — the one input shape where sentiment's internal, non-swappable tokenizer and Verbora's WordTokenizer produce the same token list. Every exclusion is proved, not just asserted, in tests/sentiment_correctness.rs, which runs the identical corpus through both crates and fails if they ever disagree.
LibraryVersionLanguageTime (median, 1024-word document)ThroughputRelative
Verbora0.2.0Rust28.04 µs35.7K/s1.00×
sentiment0.1.1Rust300.84 µs3.3K/s10.73× slower
Input (words)Verborasentiment
4139.2 ns33.59 µs
16439.9 ns38.63 µs
641.71 µs47.01 µs
2566.84 µs97.71 µs
102428.04 µs300.84 µs

Verbora wins at every size measured, by a wide and widening margin at small input — 241× at 4 words, narrowing to 10.7× at 1024. Two costs are structural to sentiment's published API and stay inside its measured region, because a caller cannot avoid them either: analyze() tokenizes the document twice (once each for its internal positivity/negativity calls) and compiles four Regexes on every call — a fixed per-call cost that is nearly the whole measurement at 4 words and mostly amortized by 1024. Verbora's own scoring loop carries negation state and probes for multi-token phrase keys on every token, a capability this corpus never exercises but still pays for.


WordNet ​

Reading the Princeton WordNet database against wordnet-db 0.1.3 (johanneswd) — a reader for the same index.*/data.* files verbora-wordnet reads, from the same directory, answering the same questions. Both sides read a real Princeton WordNet 3.1 dict/ distribution, which this repository does not vendor (Princeton's own licence, not MIT — see crates/verbora-wordnet/LICENSE-WORDNET).

Verbora can now read this dictionary completely. A defect found while first running this comparison against a real distribution — PointerSymbol::from_symbol rejected the bare ;/ - domain-pointer forms Princeton's index files actually write, affecting 8.8% of WordNet 3.1's index entries, including common words like run, cat and water — is fixed in 0.3.0. Every figure below covers the whole dictionary, not a probe list that had to avoid the unreadable 8.8%.

The two crates are mechanically opposite, which is what makes the comparison worth publishing rather than an objection to it: wordnet-db mmaps or reads all eight files and eagerly parses every index line and data record into HashMaps at open; verbora-wordnet reads bytes (or not, depending on Storage) and binary-searches the index file per query, paying the parse cost only for the one record a query actually touches.

Open and cold start ​

GroupVerbora (fastest strategy)wordnet-db (fastest mode)Verbora advantage
open9.83 µs (Pread)228.17 ms (Mmap)23,210×
cold (open + first lookup, entity)17.40 µs (Pread)232.20 ms (Mmap)13,346×

wordnet-db's two LoadModes (Mmap, Owned) both parse the entire dictionary eagerly at open — mmapping only defers the OS read, not wordnet-db's own parse pass — so both cost roughly 228–237 ms regardless of mode. verbora-wordnet's Pread/LazyResident strategies defer nearly everything, which is why open and cold cost microseconds rather than milliseconds: the real work only happens once a query actually needs it.

Query, once loaded — the headline pair: Resident vs. Owned ​

Both sides read all eight files into owned heap buffers at open, with no unsafe and no OS mapping — the only remaining difference is what each does with the bytes afterwards.

LemmaVerbora (Resident)wordnet-db (Owned)Faster
entity (index entry only)357.5 ns42.1 nswordnet-db, 8.50×
entity (full lookup)714.4 ns90.2 nswordnet-db, 7.92×
dog (full lookup)4.61 µs687.0 nswordnet-db, 6.71×
run (full lookup, 16 senses)10.03 µs1.39 µswordnet-db, 7.22×

wordnet-db wins every query once both sides are loaded, by roughly 7×–8.5× — the payoff of its eager, fully-parsed HashMap representation. Verbora's Synset owns its Strings; wordnet-db's borrows &str out of the mapped buffer, allocating only the Vec<Lemma>/Vec<Pointer> spines — an asymmetry intrinsic to holding the whole file resident, not a shortcut handed to one side.

LazyResident vs. Mmap — both crates' answer to "don't pay for what you don't touch" ​

Not the same mechanism — one defers a read, the other defers a page fault — but the two crates' comparable answers to the same question, and Mmap is wordnet-db's own default.

LemmaVerbora (LazyResident)wordnet-db (Mmap)Faster
entity (full lookup)735.1 ns90.4 nswordnet-db, 8.15×
dog (full lookup)4.60 µs638.5 nswordnet-db, 7.20×

Pread and Indexed (Verbora-internal strategies with no wordnet-db counterpart) are consistently the slowest of Verbora's four once resident — Pread re-reads from disk on every query — and are carried here only for the Verbora-internal ranking, not compared against a competitor row.

The trade-off in one sentence: Verbora opens roughly four orders of magnitude faster and wins any workload dominated by a handful of lookups per process lifetime; wordnet-db wins any workload that stays resident and issues many lookups, by paying its cost once at startup instead of once per query.


No Rust competitor exists: sentence analysis (Analyzers) ​

One of the workspace's 15 benchmarked modules has no fair Rust competitor at all: every candidate found for the composed sentence/text-analysis task Verbora's Analyzers module performs was investigated and rejected on scope grounds — no Rust crate performs the same composed task. Per this project's NO FAIR COMPETITOR FOUND policy, none is forced. A widely-used JavaScript NLP library remains the available baseline for this module. Reproduce or publish any comparison using the method on this site; do not infer a Rust ranking where no equivalent Rust implementation exists.

Full reasoning for every rejected candidate: docs/COMPETITIVE_BENCHMARKS.md § 3.

Phonetic Index / Phonetic Neighbors (PhoneticIndex) has zero competitors of any kind — a Verbora-native extension with no upstream equivalent to compare against. Its own internal build/query benchmark suite lives on the Phonetic neighbors feature page instead of here.

Upstream bugs found ​

Two verification disciplines surfaced real, reproducible defects in third-party dependencies — none in Verbora's own code: re-verifying crates flagged as stale or abandoned before trusting their numbers (this audit's own "do not trust marketing benchmarks — reproduce locally" rule), and the differential fuzzing behind the byte-exact phonetics table (Phonetics). Disclosed here, not filed upstream without separate confirmation.

  • triple_accel 0.4.0 — rdamerau("tac", "tatc") returns 2; the correct restricted-Damerau-Levenshtein (OSA) distance is 1. It over-counts an insertion adjacent to a repeated character. Both of the crate's restricted-Damerau entry points carry it — rdamerau and rdamerau_exp alike — confirmed against a from-scratch three-row OSA implementation, strsim::osa_distance, rapidfuzz's distance::osa and Verbora's osa (all four return 1), and found by a randomized sweep rather than by inspection. Real impact: fast_symspell uses this family as its post-lookup verification pass, so it can silently miss or misrank a correction on an ordinary doubled-letter typo.
  • fst 0.4.7 — its Levenshtein automaton silently returns incomplete results for same-byte-length multi-byte UTF-8 substitutions (e.g. Cyrillic characters one substitution apart). Matches a still-open upstream issue, BurntSushi/fst#38, opened 2017. The ASCII-only corpus this page's own fst comparisons use never exercises it.
  • rphonetic 3.0.6 — several encoders panic on realistic non-ASCII input, all reproduced against 3.0.6 release builds during the differential fuzzing above: Nysiis in strict mode byte-slices its code at offset 6 with no character-boundary check, panicking whenever a longer code's byte 6 splits a multi-byte character (4,233 of the 104,114 fuzzed inputs); Caverphone1/Caverphone2 panic the same way at their fixed 6-/10-byte code cut. The ASCII-only shared corpus never exercises the character- boundary panics above; on every one of those inputs Verbora's own byte-exact encoders return a documented substitute output instead of panicking — the one place they deliberately do not match rphonetic.

Library coverage summary ​

What each library actually covers, not what its scope implies. ✓ = genuine, broadly-equivalent coverage exists; P = coverage exists but only for a narrowed input domain, a reconfigured competitor, or part of the module; — = no fair competitor was found in that ecosystem for this capability at all. This does not claim other libraries are trying to be all-in-one — it shows honestly that none of them are, without implying Verbora's breadth makes it better at any one of these than a specialist crate necessarily is.

CapabilityVerboraJS libraryRust ecosystem
Tokenizers✓✓P
N-grams✓✓P
Stemmers✓✓P
Normalizers✓✓P
Inflectors✓✓P
Phonetics✓✓P
Phonetic index / neighbors✓——
Distances✓✓✓
Language detection✓—P
Script detection✓—P
Transliteration✓✓P
TF-IDF✓✓P
Classifiers (Bayes / logistic / MaxEnt)✓✓P
Sentiment✓✓P
WordNet✓✓P
POS tagging✓✓P
Spellcheck✓✓P
Trie✓✓P
Analyzers✓✓—

19 of 19 — Verbora. 16 of 19 — the JS library (missing language detection, script detection, and phonetic indexing, which it never implemented). 0 of 19 — any single Rust crate at full, unqualified equivalence across a whole module; the Rust ecosystem's real strength shows up inside individual algorithms instead — strsim/rapidfuzz are genuine Yes-equivalence competitors for most of Distances, rust-stemmers for 9 of Verbora's 16 stemmers — not as one library matching Verbora's combined scope. That fragmentation is the whole reason this audit went module-by-module rather than searching for one all-in-one Rust rival: no such rival exists, and claiming one would misrepresent the comparison.

Competitors — attribution ​

Every library compared anywhere on this page, with its official repository, package registry page, and documentation.

LibraryLanguageVersionLicenseRepositoryPackageDocs
JavaScript NLP libraryJavaScript8.1.1MIT——repository README
ngrammaticRust0.7.0MITGitHubcrates.iodocs.rs
strsimRust0.11.1MITGitHubcrates.iodocs.rs
rapidfuzzRust0.5.0MITGitHubcrates.iodocs.rs
triple_accelRust0.4.0MITGitHubcrates.iodocs.rs
editdistancekRust1.0.2MITGitHubcrates.iodocs.rs
stringmetricsRust2.2.2Apache-2.0GitHubcrates.iodocs.rs
tantivyRust0.26.1MITGitHubcrates.iodocs.rs
tokenizers (Hugging Face)Rust0.23.1Apache-2.0GitHubcrates.iodocs.rs
segtokRust0.1.5MITGitHubcrates.iodocs.rs
unicode-segmentationRust1.13.3MIT/Apache-2.0GitHubcrates.iodocs.rs
rust-stemmersRust1.2.0MIT / BSD-3-ClauseGitHubcrates.iodocs.rs
snowball_stemmers_rsRust1.0.1MITGitHubcrates.iodocs.rs
nltk-porterRust0.1.0Apache-2.0GitHubcrates.iodocs.rs
porter-stemmerRust0.1.2MPL-2.0GitHubcrates.iodocs.rs
lindera-analysisRust5.2.0MITGitHubcrates.iodocs.rs
sastrawiRust0.1.1MITGitHubcrates.iodocs.rs
diacriticsRust0.2.2GPL-3.0GitHubcrates.iodocs.rs
kana-converterRust0.1.2MITGitHubcrates.iodocs.rs
ordinalRust0.4.0MPL-2.0GitHubcrates.iodocs.rs
InflectorRust0.11.4BSD-2-ClauseGitHubcrates.iodocs.rs
pluralizerRust0.5.0MIT/Apache-2.0GitHubcrates.iodocs.rs
trie-rsRust0.4.2MIT OR Apache-2.0GitHubcrates.iodocs.rs
qp-trieRust0.8.2MPL-2.0GitHubcrates.iodocs.rs
fast_radix_trieRust1.2.0MITGitHubcrates.iodocs.rs
fstRust0.4.7MIT OR UnlicenseGitHubcrates.iodocs.rs
rphoneticRust3.0.6Apache-2.0GitHubcrates.iodocs.rs
pixelglow/double_metaphoneC++1179dd226 (2014)BSD-2-ClauseGitHubvendored, no registry packageheader comment
whatlangRust0.18.0MITGitHubcrates.iodocs.rs
linguaRust1.8.0Apache-2.0GitHubcrates.iodocs.rs
whichlangRust0.1.1MITGitHubcrates.iodocs.rs
wana_kanaRust5.0.0MITGitHubcrates.iodocs.rs
postaggerRust0.0.3Apache-2.0GitHubcrates.iodocs.rs
rust-bertRust0.23.0Apache-2.0GitHubcrates.iodocs.rs
symspellRust0.5.2MITGitHubcrates.iodocs.rs
harper-coreRust2.8.0Apache-2.0GitHubcrates.iodocs.rs
spellbookRust0.4.2MPL-2.0GitHubcrates.iodocs.rs
fast_symspellRust0.1.10MITno repository URL in its published metadata — re-verified via cargo's registry-cache tarball rather than taken on trust, see the Spellcheck section abovecrates.iodocs.rs
tfidf (afshinm)Rust0.3.0MITGitHubcrates.iodocs.rs
rust-tfidfRust1.1.1MIT OR Apache-2.0GitHubcrates.iodocs.rs
smartcoreRust0.6.5Apache-2.0GitHubcrates.iodocs.rs
linfa-bayesRust0.8.1MIT OR Apache-2.0GitHubcrates.iodocs.rs
naivebayesRust0.1.2Apache-2.0GitLabcrates.iodocs.rs
linfa-logisticRust0.8.1MIT OR Apache-2.0GitHubcrates.iodocs.rs
rustlearnRust0.5.0Apache-2.0GitHubcrates.iodocs.rs
sentimentRust0.1.1MITcrates.iocrates.iodocs.rs
wordnet-dbRust0.1.3MIT OR Apache-2.0GitHubcrates.iodocs.rs

Full research dossier for every candidate considered — including every crate investigated and not selected, and why — lives in docs/COMPETITIVE_BENCHMARKS.md.

Reproducing these numbers ​

Everything on this page regenerates from a clean checkout:

bash
# Shared inputs both sides read (run once)
python3 tools/bench-data/generate.py

cd benchmarks/competitive

# Third-party model/dictionary assets for POS tagging, spellcheck and WordNet
./scripts/fetch-models.sh

# Every module's Criterion benchmarks (this page's numbers)
cargo bench --release

# Machine metadata (results/metadata.json)
./scripts/machine-metadata.sh

# Join Criterion's raw output into results/results.json + results/raw/
python3 scripts/collect-results.py distance levenshtein:verbora,strsim,rapidfuzz ...
# (see benchmarks/competitive/README.md for the full per-module command list)

# Language-detection accuracy report
cargo run --release --example language_accuracy

../../scripts/competitive-benchmarks.sh (repo root) drives all of the above in one command. Full detail, including the exact collect-results.py invocation for every module, is in benchmarks/competitive/README.md and in each module's own dossier in docs/COMPETITIVE_BENCHMARKS.md.

  • String distance results — the JavaScript-library baseline this page's Distance section extends with real Rust competitors.
  • Benchmark method — warmup, sample-count and regression-tracking conventions this page inherits.
  • Parallelism — Verbora's own sequential-vs-parallel numbers, thread counts disclosed, for the APIs this page's competitors have no equivalent to compare against.
  • docs/PERFORMANCE_GAPS.md — every real loss on this page, with its investigated likely cause and — where one exists — a flagged, not-yet-implemented optimization opportunity.
  • docs/COMPETITIVE_BENCHMARKS.md — the full research matrix: every competitor considered, selected or rejected, and why.

Released under the MIT License.