Skip to content

Quick answers ​

Every choice on this site, condensed into one page, for when you know what you need and just want the answer. Each section links to the page that explains it.

Which API shape? ​

You wantShapeExample calls
a result you can hold, index or pass oneagertokenize_borrowed(), process(), keys_with_prefix()
to consume it once, in order — maybe not all of itlazytokens(), ngrams(), iter_keys_with_prefix()
to do this millions of times with the same shape of outputinto-buffertokenize_borrowed_into(), pluralize_into()
generic code over the traitbatchtokenize_batch() (sequential today)

→ The four API shapes

Which tokenization call? ​

Your situationCall
I look at each token once, or stop earlytokens()
I want a Vec to index or iterate, and the input outlives ittokenize_borrowed()
I am in a loop over many documents that all outlive the loopbuf.clear(); tokenize_borrowed_into(doc, &mut buf)
My tokens must outlive the text they came fromTokenizer::tokenize → Vec<String>
My function takes "some tokenizer" and only needs slicesBorrowingTokenizer — all three implement it
I have a slice of documents and want one callTokenizer::tokenize_batch — a sequential map; par_tokenize_batch (feature parallel) if the CPU cost justifies threads

→ Choosing a tokenization API

Which tokenizer? ​

What you are splittingTokenizer
Words, in any language that uses spacesWordTokenizer
Words and the punctuation and whitespace between themSegmentTokenizer — concatenation reproduces the input
SentencesSentenceTokenizer — with_abbreviations if you have a list
Text that needs re-assembly, highlighting or offsetsSegmentTokenizer
Thai, Khmer, Chinese or Japanese word segmentationnothing here — UAX #29 does not segment languages without spaces

→ Tokenizers

Which distance metric? ​

What you are comparingMetricWorking set
Same length by construction (codes, hashes, fixed fields)hamming() → Option<usize>; None when the character counts differ—
Typos, and a swap is honestly two editslevenshtein()bit-vector / 1 row weighted
Typos, adjacent swaps cost 1, never edited againosa()bit-vector / 3 rows weighted
Typos, swaps may be arbitrarily far apartdamerau_levenshtein()Zhao–Sahni rows / full matrix weighted
Position of the best approximate occurrence in a longer stringlevenshtein_search() / damerau_levenshtein_search() / osa_search()bit-vector columns (plain, unit cost) / full matrix
Names or short records, raw scorejaro()—
Names or short records, prefix-boosted (the usual choice)jaro_winkler()—
Shared content, not order or positiondice_coefficient() — bigram set overlap; case and whitespace significant—
Any of those edits priced differentlythe matching *_weighted(), plus a cost set from LevenshteinCosts / OsaCosts / DamerauCostsscalar dynamic program

→ Choosing a distance API

Which phonetic encoder? ​

What you needEncoderKey
English surnames, cheapest possible blocking keySoundExa letter and 3 digits, very coarse
General English words, one key, better precisionMetaphoneletters, unbounded
English text with names of many origins, two indexable keysDoubleMetaphonetwo keys of up to 4 characters; match on either
Slavic / Germanic / Ashkenazi-Jewish surnamesDaitchMokotoff6 digits; codes() returns every branch
American surnames, US-census rulesNysiisletters
German-language names and wordsColognedigits
Names whose language of origin is itself uncertainBeiderMorsea candidate list, per language set
"Which encoder should I even use for this text?"verbora_language::recommenda PhoneticStrategy, whose primary is None rather than a guess when nothing fits

Twelve encoders ship in all — the eight rows above are the common answers. → Phonetics · Language

Which n-gram call? ​

If you…Call
have a slice of elements and want its windows — the defaultngrams(seq, n)
stop early, or fold windows into a counterngrams(seq, n) — it is already lazy
want indexable windowsngrams(seq, n).collect::<Vec<_>>()
want the ends of the sequence to appear in as many windows as the middlePadded::new(seq, n, Some(&start), Some(&end)).ngrams()
need countsfold ngrams(seq, n) into a HashMap, keyed on the window itself
need the windows to outlive the sequencecopy them out with .map(<[_]>::to_vec)
have text and want character windowschar_ngrams(text, n)
have text and want word windowstokenize first, then ngrams over the token slice

→ Choosing an n-gram API

Which trie query? ​

Your questionCall
"Is this exact string stored?"contains()
"How many words are stored?"len() — words, not nodes; O(1)
"How big is the structure?"node_count() — arena nodes, one per scalar; O(1)
"Which stored words start with my string?" — all of them, indexablekeys_with_prefix() → Vec<String>
The same, but only the first N, or I stop on a conditioniter_keys_with_prefix().take(N)
The same, but I only read them — no String per keyfor_each_key_with_prefix()
"How many words start with this?"iter_keys_with_prefix(p).count() — one descent, no traversal
"Does anything start with this?"iter_keys_with_prefix().next().is_some()
"Give me every word in the trie"keys() — lazy; same as iter_keys_with_prefix("")
"Which stored words are prefixes of my string?" — all, shortest firstprefix_matches() → Vec<Cow<str>>
The same, but only the shortest or the first fewiter_prefix_matches().next() / .take(n)
"Where does the longest stored prefix end?" — as textlongest_prefix() → PrefixSplit { word, rest }
The same, but as scalar counts, exact and allocation-freelongest_prefix_lengths() → PrefixSplitLengths
Built once, then queried foreverfreeze() → FrozenTrie; keys_slice() borrows instead of allocating

→ Trie

Which normalizer? ​

What you are normalizingCall
Text you will store and show a human againnfc()
A lookup key that must ignore width, ligation and circlingnfkc()
A lookup key that must ignore accents (Latin script)remove_diacritics()
Both of those at onceremove_diacritics(&nfkc(text))
Text whose combining marks you will inspect yourselfnfd() / nfkd()
Japanese halfwidth katakana and fullwidth alphanumericsnfkc()
Kana into Latin lettersnothing here — that is Transliterators

→ Normalizers

Should I optimise this? ​

Ask in this order:

  1. Has a profiler told me this line is hot? No → use the high-level API and stop reading.
  2. Is the cost the container allocation, or the work inside it? The work → a different API will not help; look at the algorithm, the input size, or how many candidates you are comparing.
  3. Do I consume the result once, in order? Yes → tokens(), no container at all. No → tokenize_borrowed_into(), one container reused.

Still not fast enough, and measured in seconds of CPU? Check whether the operation already has a par_*_batch (fourteen crates ship one, opt-in behind a parallel feature); if not, chunk the input and parallelise at your own call site.

→ Ergonomics vs throughput · Parallelism

Which workload am I in? ​

How work arrivesWorkloadWhat you optimise
One input, answer nowInteractivelatency, ergonomics
More input than memory, or output needed before input endsStreamingbounded memory, laziness, early output
Many documents, offlineBatchmemory reuse, shared setup
Many documents, and seconds of CPU to spendParallel corpuschunking, per-worker state

→ Recipes by workload

Released under the MIT License.