The four API shapes
Verbora's naming is regular. Once you know the four shapes, you can predict what a function does — and what it costs — from its name alone.
| Shape | Name pattern | Returns | Allocates | Lazy |
|---|---|---|---|---|
| Eager | verb(input) | owned collection or String | yes — the container | ❌ |
| Lazy | nouns(input), iter_* | an Iterator | no | ✅ |
| Into-buffer | verb_into(input, &mut out) | () | no, after warm-up | ❌ |
| Batch | verb_batch(&[input]) | Vec<…> per input | yes | ❌ |
SoundEx::process_into and Metaphone::process_into, the two encoders whose keys are most often accumulated in bulk. The feature pages say exactly which exist. 1. Eager — tokenize_borrowed(), process(), pluralize(), keys_with_prefix()
Does the work now, hands back a complete result you own.
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};
let tokens = WordTokenizer.tokenize_borrowed("the quick brown fox");
assert_eq!(tokens.len(), 4);
assert_eq!(tokens[2], "brown");Use it when the result is small, you need random access or a length, you are passing it somewhere that wants a slice, or you simply want the code to read well. That is most code.
Do not use it when you are about to consume the result once, in order, and throw it away — that is what the lazy shape is for — or when you are calling it in a loop millions of times and the container allocation shows up in a profile.
Cost. One container allocation, plus growth reallocations as it fills. Note what is not allocated: every tokenizer's tokens are &str slices of your input, so tokenize_borrowed costs no per-token String. (Tokenizer::tokenize is the owned sibling, and that one does cost a String per token — take it only when the tokens must outlive the text.)
2. Lazy — tokens(), ngrams(), char_ngrams(), iter_keys_with_prefix()
Returns an iterator. Nothing happens until you pull.
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};
let shouty: Vec<String> = WordTokenizer
.tokens("the quick brown fox")
.filter(|w| w.len() > 3)
.map(|w| w.to_uppercase())
.collect();
assert_eq!(shouty, ["QUICK", "BROWN"]);Use it when you are building a pipeline, when you might stop early, when the input is large enough that materialising every token at once matters, or when you want to hand a stream to another API that takes IntoIterator — such as phoneticize_tokens, which composes directly with a tokenizer's iterator and never builds the intermediate Vec.
Do not use it when you need the result more than once, need its length up front, or need indexing. Re-running an iterator means re-doing the work; a Vec you can read twice.
Cost. Nothing per token — the iterator is a small struct on the stack, scanning as it goes. Every tokenizer in verbora-tokenizers is lazy end to end, and ngrams/char_ngrams allocate nothing at all.
3. Into-buffer — tokenize_into(), pluralize_into(), stem_into()
Writes into storage you own and keep between calls.
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};
let corpus = ["the quick brown fox", "jumps over the lazy dog"];
let mut buf: Vec<&str> = Vec::new();
let mut total = 0;
for document in corpus {
buf.clear(); // capacity survives; contents do not
WordTokenizer.tokenize_borrowed_into(document, &mut buf);
total += buf.len();
}
assert_eq!(total, 9);Use it when the same operation runs in a tight loop and the container allocation is a real cost — batch jobs, corpus indexing, offline processing.
Do not use it when you call the operation once. You have added a mutable binding and a manual clear() to your code in exchange for one saved allocation.
Cost. After the buffer reaches its high-water mark, zero allocations. Before that, the same growth pattern as the eager shape.
BorrowingTokenizer::tokenize_borrowed_into and Tokenizer::tokenize_into append — they do not clear, so you can accumulate across inputs on purpose. Stemmer::stem_intoclears first. They differ because appending tokens across documents is useful and appending stem fragments is not, but you must check per API. Each one says so in its rustdoc. 4. Batch — tokenize_batch(), stem_batch()
Takes a slice of inputs, returns a result per input.
use verbora_tokenizers::{Tokenizer, WordTokenizer};
let out = WordTokenizer.tokenize_batch(&["one two", "three four five"]);
assert_eq!(out.len(), 2);
assert_eq!(out[1].len(), 3);verbora_core::Tokenizer and verbora_core::Stemmer are provided methods whose default bodies are a sequential map: one fresh Vec<String> per input, no shared buffer, no parallelism. tokenize_batch allocates more than tokenize_borrowed_into does, and produces owned Strings rather than the borrowed &str that tokenize_borrowed gives you. Use it when you are writing generic code over the Tokenizer trait and want the batch operation to improve automatically if an implementation overrides it.
Do not use it when you want throughput today. Reach for verbora-tokenizers's own par_tokenize_batch (behind its parallel feature) or tokenize_borrowed_into with a reused buffer, or rayon over your own slice for anything without a built-in par_* variant — see Parallelism.
Two more shapes you will meet
These are not "levels" — they are type choices that change what you can do with a result.
Cow-returning functions
All five normalizers, and Stemmer::stem, return Cow<'_, str>: borrowed when nothing changed, owned when something did. For the normalizers that is a guarantee rather than a fast-path description — Cow::Borrowed if and only if the result is byte-identical to the input — so branching on it is correct code. See Normalizers.
use std::borrow::Cow;
use verbora_normalizers::remove_diacritics;
// Nothing to fold — no allocation at all.
assert!(matches!(remove_diacritics("plain ascii"), Cow::Borrowed(_)));
// A fold happened — one String.
assert!(matches!(remove_diacritics("café"), Cow::Owned(_)));
assert_eq!(remove_diacritics("café"), "cafe");This matters because these functions are usually called on text that needs no change — an already-composed string handed to nfc, an ASCII token handed to the diacritic fold. See Zero-copy and Cow.
Result-returning constructors
FALLIBLE
Where a configuration value could put a type into a state with no sensible behaviour, the constructor returns Result and the state simply cannot be built. SentenceTokenizer::with_abbreviations is the example: an empty abbreviation would suppress every sentence boundary in the document, so it is rejected rather than documented.
use verbora_tokenizers::{AbbreviationError, SentenceTokenizer};
assert!(SentenceTokenizer::with_abbreviations(["Dr."]).is_ok());
assert_eq!(
SentenceTokenizer::with_abbreviations([""]),
Err(AbbreviationError::Empty { index: 0 })
);The same idea shows up as a type rather than a Result where it can: verbora_ngrams::ngrams takes a NonZeroUsize, so a zero window size is unrepresentable and no call site needs a guard.
Summary
| You want | Shape | Example calls |
|---|---|---|
| a result you can hold, index or pass on | eager | tokenize_borrowed(), process(), keys_with_prefix() |
| to consume it once, in order — maybe not all of it | lazy | tokens(), ngrams(), iter_keys_with_prefix() |
| to do this millions of times with the same shape of output | into-buffer | tokenize_borrowed_into(), pluralize_into() |
| generic code over the trait | batch | tokenize_batch() (sequential today) |
Next: the same reasoning applied concretely, in Tokenization.