Why there is more than one API
verbora-tokenizers offers three ways to split a string:
tokenizer.tokenize(text) // a Vec you own
tokenizer.tokens(text) // an iterator
tokenizer.tokenize_into(text, &mut buffer) // appended into memory you keepThey are not competing implementations. Each tokenizer has exactly one implementation of its behaviour; these are three ways of moving its output to you, and they differ only in who owns the memory. Same result, every time.
This section tells you which one to pick — here and everywhere else Verbora offers a choice. Every group of similar-looking functions on this site comes with the same information: what each one does, when to use it, when not to, whether it allocates, whether it is lazy, and — when the difference is a performance difference — what measurement supports the recommendation.
The one thing to internalise
tokenize() is the right call for the overwhelming majority of programs. The other shapes are not "the fast version" — they are answers to questions most code never asks. Reaching for tokenize_into() in a web handler that runs once per request buys you nothing and costs you a mutable buffer to manage. The variants exist because these workloads have genuinely different bottlenecks:
| Workload | Bottleneck | Shape that helps |
|---|---|---|
| One string, once | Nothing. Readability wins. | tokenize() |
| Feed tokens into a filter/map chain | Building an intermediate Vec you immediately consume | tokens() |
| Find the first token matching a predicate | Splitting the whole string when you needed a prefix of it | tokens() |
| 40M documents in a loop | One allocation per document | tokenize_into() |
| Documents don't fit in memory | Peak memory | tokens() |
| One query against thousands of candidates | Rebuilding the same per-query state on every comparison | A build-once, query-many type — PreparedPattern, FuzzyIndex, … |
| 16 idle cores | Wall clock | A crate's own par_*_batch, or rayon at your call site |
Build once, query many
The four API shapes is about how a call moves its result to you. There is a fifth arrangement that is a type choice rather than a call shape, and it is the one to reach for whenever one operand is fixed and the other varies: build a value from the fixed side, then query it as many times as you like.
| Type | Built from | Queried with |
|---|---|---|
PreparedPattern | one pattern string | levenshtein, osa against each candidate |
FuzzyIndex | a word list, via FuzzyIndexBuilder | neighbors(query, max_distance) |
DeletionIndex | a word list, via DeletionIndexBuilder | neighbors(query, max_distance) — Err(DistanceBeyondIndex) past the ceiling it was built for |
PhoneticIndex | a word list plus an encoder, via PhoneticIndexBuilder | neighbors(query) |
FrozenTrie | a Trie, once insertion is done | prefix and membership queries |
They all share one contract: construction does work proportional to the fixed side, queries do not repeat it, and the value is immutable afterwards — so a single instance can be shared across threads by reference. What varies is how much they save. An index changes the complexity of a search by ruling candidates out; PreparedPattern changes the constant factor of a comparison that still visits every candidate. Cutting the candidate set is the bigger win where both apply — see Getting the order right.
Where to start
Cow, buffer reuse, batching, parallelism.Subsystems with only one sensible API — phonetics, normalizers, inflectors, tries — carry their "Choosing the right API" section on their own feature page, because the choice there is usually which type rather than which call shape.
What Verbora does not have
Knowing the absences saves you the search:
par_*_batch function behind a parallel Cargo feature — never on by default, never a second implementation, each one added because a benchmark showed a real win. Everything else has no par_* function and no internal thread pool; this site shows you how to write it at your own call site with your own rayon dependency and explains when it actually pays. See Parallelism for the full table. - Batch APIs are minimal.
verbora_core::Tokenizer::tokenize_batchandverbora_core::Stemmer::stem_batchare provided trait methods with sequential default bodies. No other crate has a batch entry point. _intovariants are rare, and they are not where you would guess. The whole list:Tokenizer::tokenize_intoandBorrowingTokenizer::tokenize_borrowed_into;Stemmer::stem_into; every inflector'spluralize_into/singularize_into, plusOrdinalInflector::nth_intoandCaseMode::apply_into;SoundEx::process_intoandMetaphone::process_into— the two phonetic encoders whose keys are most often accumulated in bulk;BrillTagger::tag_into/annotate_into;transliterate_ja_into; andverbora_tfidf::Tokenize::tokenize_intoon that crate's tokenizer trait. Distance, normalizers and n-grams have none — the first has nothing hoistable to write into, and the other two allocate nothing to begin with.- No scratch-buffer API. No function anywhere in Verbora takes mutable working memory you lend it for the duration of a call — there is no
levenshtein_with_scratch, and the Levenshtein family builds its own dynamic-programming working set per call. That is a different thing from prepared state derived from one fixed operand, which does exist: see Build once, query many above, andPreparedPatternfor the distance case specifically.
Where an absence is inconvenient, the relevant page shows the call-site workaround rather than pretending an API exists.
Getting the order right
Suppose you are writing a spell-check suggestion endpoint: one misspelled word per request, a dictionary of 100,000 candidates, and you want the ten closest by edit distance. The instinct is to look for levenshtein_batch. It does not exist — and it would not be the biggest win available anyway.
- Cut the candidate set first. 100,000 Levenshtein calls to return ten results is the wrong shape regardless of how fast each call is. A
Trieprefix query or a phonetic key bucket reduces the candidates by orders of magnitude, and that is the optimisation that matters. - Then pick the metric. For typos,
levenshtein; for names,jaro_winkler, which weights a common prefix. See Choosing a distance API. - Then pick the call shape. Gate on the character-count difference before paying for a comparison, keep inputs ASCII where you can so the byte fast path applies, build the misspelled word into a
PreparedPatternonce since it is fixed and the candidates are what vary, and only then reach forverbora-distance's ownpar_levenshtein_batch(behind itsparallelfeature) orrayonat your call site.
This section is organised to make step 3 easy, so you can spend your attention on steps 1 and 2 — see Recipes by workload for that half of the problem.