Features overview
Every subsystem below is implemented, tested, and documented. Start from what you want to do, or from the crate map further down.
Start from the problem
| I want to… | Use |
|---|---|
| Split text into words | Tokenizers — WordTokenizer |
| Split text into sentences | SentenceTokenizer |
| Keep the whitespace and punctuation too, so the pieces re-assemble | SegmentTokenizer |
| Measure how similar two strings are | String distance |
| Correct a typo | levenshtein, usually after narrowing candidates |
| Correct spelling against a corpus | Spellcheck |
| Match names that sound alike | Phonetics |
| Search a whole dictionary for sound-alikes | Phonetic neighbors |
| Match names across language boundaries (genealogy) | Beider-Morse |
| Match names with a shared prefix | jaro_winkler |
| Work out which encoder or stemmer a text even needs | Language |
| Build an autocomplete index | Trie |
| Find repeated phrases | N-grams |
| Strip accents before comparing | remove_diacritics |
| Put text in one canonical spelling before storing or hashing it | nfc |
| Fold halfwidth katakana and fullwidth Latin onto one form | nfkc |
| Romanise Japanese kana | Transliterators |
| Pluralise a noun, or write "23rd" | Inflectors |
| Reduce words to a stem | Stemmers |
| Rank documents by relevance | TF-IDF |
| Classify documents into categories | Classifiers |
| Score sentiment over a token list | Sentiment |
| Look up synonyms and word relations | WordNet |
| Tag parts of speech | POS tagger |
| Split a tagged sentence into subject and predicate | Sentence analyzers |
| Write generic code over any tokenizer | Core vocabulary |
The crates
| Subsystem | Crate | Public surface |
|---|---|---|
| Tokenizers | verbora-tokenizers | WordTokenizer, SegmentTokenizer, SentenceTokenizer — UAX #29 boundaries, borrowed tokens |
| String distance | verbora-distance | 7 metrics; the three edit distances also in weighted and substring-search forms, plus PreparedPattern |
| Phonetics | verbora-phonetics | 12 encoders: SoundEx, Metaphone, Double Metaphone, Daitch–Mokotoff, Beider–Morse, Cologne, Nysiis, Caverphone 1/2, Phonex, Refined Soundex, Match Rating |
| Phonetic neighbors | verbora-phonetics | PhoneticIndex — dictionary-wide candidate generation over SoundEx, Metaphone or DoubleMetaphone |
| Beider-Morse | verbora-phonetics | Cross-language surname matching over up to 18 languages at once, with auto-detection |
| Language | verbora-language | Script detection, optional statistical language detection, and recommend() — language → phonetic encoder |
| N-grams | verbora-ngrams | ngrams, Padded boundary symbols, char_ngrams |
| Normalizers | verbora-normalizers | the four Unicode normalization forms plus remove_diacritics |
| Inflectors | verbora-inflectors | 6 inflectors, runtime rules |
| Trie | verbora-trie | prefix tree, prefix and path queries |
| Transliterators | verbora-transliterators | Japanese kana → romaji, one left-to-right pass over a generated mora index |
| WordNet | verbora-wordnet | lexical database, synsets, relation traversal, 4 storage strategies |
| TF-IDF | verbora-tfidf | TfIdf, TermScore, list_terms/tfidf/tfidfs, JSON persistence |
| Sentiment | verbora-sentiment | 14 lexicons across 10 languages, sticky negation |
| Classifiers | verbora-classifiers | Bayes, logistic regression, MaxEnt + GIS |
| Stemmers | verbora-stemmers | 16 stemmers: Porter/Snowball across 12 languages, plus Carry (French), Lancaster, Japanese, Indonesian |
| Spellcheck | verbora-spellcheck | frequency-ranked correction, BK-tree and deletion indexes |
| POS tagger | verbora-tagger | Brill tagging, training and evaluation over a lexicon you supply |
| Sentence analyzers | verbora-analyzers | phrase annotation, subject/predicate splitting, sentence type |
| Utilities | verbora-util | stop words, abbreviations, weighted graphs, topological ordering and path trees |
| Core vocabulary | verbora-core | 5 traits, StopWordLanguage, StopWords, the process-global stop-word list |
Language support
Tokenization and normalization are not per-language axes any more: both follow the Unicode standard and therefore cover the whole character repertoire at once. The columns that do vary by language are the ones backed by per-language data.
| Language | Inflector | Phonetics |
|---|---|---|
| English | ✅ nouns, verbs, ordinals | ✅ all four core encoders |
| French | ✅ nouns, ordinals | |
| Japanese | ✅ nouns |
Tokenization is UAX #29 word and sentence boundaries, so every language that separates words with spaces is covered by the same three tokenizers — with one stated limitation: the standard's default rules do not segment Thai, Lao, Khmer, Myanmar, Chinese or Japanese, and Verbora ships no dictionary segmenter for them.
Normalization is the four Unicode normalization forms plus a combining-mark fold defined over Canonical_Combining_Class, so remove_diacritics handles Latin, Greek, Cyrillic, Hebrew and Arabic script without a per-language rule set — read Normalizers before applying it to Thai or Devanagari.
Stemmers, sentiment lexicons and Beider-Morse each have their own language axis; see Stemmers, Sentiment and Beider-Morse. The POS tagger has none of its own: it ships no dictionary, so its language is whichever one the lexicon and rule set you give it were written for.
Where else to look
- Choosing the right API — cross-subsystem decisions.
- Performance — allocation, zero-copy, parallelism.
- Benchmarks — measured numbers and how to reproduce.
- Recipes — end-to-end pipelines.
- Status and scope — what "available" means, and the boundaries.
- Exact signatures live in rustdoc; these pages cover selection, composition, costs and common mistakes.