Skip to content

How Verbora uses Rust ​

Verbora is fast for an unglamorous reason: it does not allocate much. The algorithms are the standard ones — Levenshtein is Levenshtein, Jaro–Winkler is Jaro–Winkler. What differs is the data that flows through them, and how often it has to be copied.

This section covers the techniques that produce that, where each one surfaces in the API, and what it means for the code you write.

The techniques, and where to find them ​

TechniqueWhere it shows up in the APIPage
BorrowingEvery tokenizer yields &str slices of your input; every n-gram window borrows your sliceZero-copy
CowAll five normalizers, guaranteed borrowed when nothing changed; Stemmer::stemZero-copy
Lazy iteratorstokens(), ngrams(), char_ngrams(), iter_keys_with_prefix()Iterator vs _into
Caller-owned bufferstokenize_borrowed_into(), pluralize_into(), stem_into()Buffer reuse
Choosing the smallest working setLevenshtein's bit-vector / row / matrix modesCache locality
Struct-of-arraysThe Levenshtein search matrixCache locality
Flat arenasTrie's Vec<Node> addressed by u32Cache locality
Inline small collectionsSmallVec children per trie nodeCache locality
Stack buffers for small inputsJaro–Winkler's match flagsAllocation
Cheaper hash keysDice hashes (char, char) instead of a String per bigramAllocation
Exact fast pathsASCII &[u8] vs promotion to a decoded Vec<char>, in distanceZero-copy
Monomorphised iteratorstokens() returns impl Iterator, not a boxed trait object, so each boundary scan inlinesCache locality

Read these in order ​

How to read the numbers here ​

Timings are measured; allocation counts are not. Published timings come from the benchmark pages, and today they cover verbora-distance. Criterion benchmarks for tokenizers, phonetics, n-grams, normalizers, inflectors and the trie exist in-tree (crates/*/benches/) but their tables are not published yet, so this section says "fewer allocations" rather than quoting a speed figure for those subsystems. Where a page describes allocation behaviour it describes what the code does, read from the source: there is no per-API allocation-count table anywhere in this repository. The one allocation instrument that does exist is verbora-spellcheck's counting_alloc — a #[cfg(test)] global allocator that its own memory-bound tests measure peak bytes with. It is scoped to that crate's test build, it is not compiled into any published library, and it produces no figure this section quotes.
Nothing here has been re-measured against 0.2.0. The benchmark campaign for that release has not been run, so every timing on these pages predates it and is marked pending — several of the kernels underneath them have since been replaced. See Upgrading from 0.1 to 0.2.

Released under the MIT License.