Stemmers
verbora-stemmers reduces inflected words to stable stems for indexing, matching and feature extraction. Sixteen stemmers ship: twelve Porter/Snowball implementations, one per language, plus the French Carry stemmer, Lancaster, a Japanese katakana stemmer and an Indonesian dictionary stemmer.
When to use it
- You are building a search index, a TF-IDF model or a classifier and want
running,runsandran-style variants to collide on one key. - You want stop words dropped and tokens stemmed in a single pass over text.
- You need stemming across languages behind one trait.
When not to use it
- You want a lemma. A stemmer is reductive and its output is often not a word (
play→plaiunder Porter). If you want real dictionary forms, see Inflectors for the generative direction. - You want tokenization control.
TokenizeAndStemuses each language's own word-character class. Tokenize with Tokenizers and callstemper token when you need a different split.
Quick example
use verbora_stemmers::{LancasterStemmer, PorterStemmer, TokenizeAndStem};
fn main() {
let porter = PorterStemmer::new();
assert_eq!(porter.stem("running"), "run");
assert_eq!(LancasterStemmer::new().stem("maximum"), "maxim");
assert_eq!(
porter.tokenize_and_stem("My dog is very fun to play with", false),
["dog", "fun", "plai"]
);
}Choosing the right API
| Need | API | Notes |
|---|---|---|
| Stem one token | the stemmer's stem(&str) | returns Cow<'_, str>; borrows when the token is already its own stem |
| Stream stems from text | TokenizeAndStem::stems | the lazy primitive: tokenizes, filters stop words, stems |
| Collect stems from text | TokenizeAndStem::tokenize_and_stem | stems(..).collect() |
| Same tokens, many documents | TokenizeAndStem::tokenize_and_stem_cached | you own a token → stem map, so each distinct token is stemmed once |
| Many documents at once | TokenizeAndStem::par_tokenize_and_stem_batch | feature parallel; one rayon task per document |
| Generic over any stemmer | verbora_core::Stemmer | implemented by every stemmer here |
stems is the primitive; the collecting, caching and parallel entry points all preserve its tokenization, casing and stop-word behaviour exactly. Each takes a keep_stops: bool — pass false to drop stop words.
use std::collections::HashMap;
use verbora_stemmers::{PorterStemmer, TokenizeAndStem};
fn main() {
let porter = PorterStemmer::new();
// Lazy: stop as soon as you have what you need, materialising nothing.
let first = porter.stems("My dog is very fun to play with", false).next();
assert_eq!(first.as_deref(), Some("dog"));
// Cached: one entry per distinct token across the whole corpus.
let mut cache: HashMap<String, String> = HashMap::new();
let corpus = ["dogs playing", "dogs running"];
for doc in corpus {
let stems = porter.tokenize_and_stem_cached(doc, false, &mut cache);
assert!(!stems.is_empty());
}
assert_eq!(cache["dogs"], "dog");
}The cache is caller-owned so eviction, lifetime and hasher stay your decision, and every stemmer but PorterStemmerNl stays zero-sized and Sync — that one carries a one-byte sticky flag and is neither, as the callout below explains. It is consulted only for the string that would have been handed to the stemmer — after the stop-word test, and only for tokens that pass the language's gate — so results are identical to tokenize_and_stem. Entries are trusted verbatim: do not share one map between two different stemmers.
Implementations
| Stemmer | Language |
|---|---|
PorterStemmer | English |
PorterStemmerDe / PorterStemmerEs / PorterStemmerFa / PorterStemmerFr | German, Spanish, Persian, French |
PorterStemmerIt / PorterStemmerNl / PorterStemmerNo / PorterStemmerPt | Italian, Dutch, Norwegian, Portuguese |
PorterStemmerRu / PorterStemmerSv / PorterStemmerUk | Russian, Swedish, Ukrainian |
CarryStemmerFr | French, the Carry algorithm |
LancasterStemmer | English, more aggressive than Porter |
StemmerJa | Japanese katakana |
StemmerId | Indonesian, dictionary-driven |
PorterStemmerDe takes options (PorterStemmerDeOptions), PorterStemmerFr exposes its Regions, and StemmerId exposes its Removal, RemovalKind and RuleResult types. Exact per-language behaviour is in the Rust API reference.
Important behaviour
The text unit is one Unicode scalar value, in every stemmer here and in every other Verbora crate. Region boundaries (R1, R2, RV), the short-word gates and each rule's removal size are all counts of scalar values — which is what the published algorithms ask for, since Snowball and Porter are specified over letters and no such definition names a unit of storage. Below U+10000 a scalar and a UTF-16 code unit coincide exactly, so this changes an answer only for astral text: PorterStemmer::stem("😀s") returns "😀s" unchanged, because "😀s" is two scalar values and Porter's three-letter gate declines to run. A cut at a region boundary is therefore a cut at a character boundary by construction.
Stop-word lists are process-global, per language. English is verbora-core's own list, reached through verbora_core::is_global_stopword and mutated through add_global_stopword/remove_global_stopword (re-exported by verbora-util), so a word added through PorterStemmer is also seen by LancasterStemmer. Not every stemmer exposes mutators:
| Stemmer | add_stop_word(s) | remove_stop_word(s) |
|---|---|---|
PorterStemmer, LancasterStemmer, StemmerId | ✅ | ✅ |
PorterStemmerNo, PorterStemmerSv | ✅ | ❌ |
PorterStemmerPt | ✅ (add_stop_words only) | ❌ |
| all others | ❌ | ❌ |
A stemmer that exposes no mutator does not make its list immutable. Every non-English list is reachable through verbora_stemmers::Language, whose thirteen variants each carry defaults, words, contains, add, add_all, remove, remove_all and reset. Those lists are process-global too, so Language::De.add("foo") changes the answer every PorterStemmerDe in the process gives for the rest of the process, and reset is the only way back.
PorterStemmerNl is stateful. Its suffix_e_removed flag is set by one step, read by a later one, and never reset, so one instance reused across a corpus gives different answers from a fresh instance per word — stem("onaantastbar") is "onaantastbar" on a fresh stemmer and "onaantast" once an earlier word has tripped the flag. It is therefore not Sync, has no par_tokenize_and_stem_batch, and ignores the stem cache. Construct one per document. Parallelism is per document, not per word. Per-word stemming costs tens of nanoseconds to a few microseconds — close enough to rayon's own per-task scheduling cost that word-level parallelism would mostly measure the scheduler. A whole document clears that floor comfortably, so par_tokenize_and_stem_batch fans out one task per document and preserves input order (out[i] is always docs[i]'s result).
Related
- Tokenizers — what to run first when you need a different split
- TF-IDF and Classifiers — the usual consumers
- Sentiment
- Core traits —
verbora_core::Stemmer - Cargo features — enabling
parallel
API reference
cargo doc -p verbora-stemmers --no-deps --openSource: crates/verbora-stemmers/src/.