Skip to content

Features overview ​

Every subsystem below is implemented, tested, and documented. Start from what you want to do, or from the crate map further down.

Start from the problem ​

I want to…Use
Split text into wordsTokenizers — WordTokenizer
Split text into sentencesSentenceTokenizer
Keep the whitespace and punctuation too, so the pieces re-assembleSegmentTokenizer
Measure how similar two strings areString distance
Correct a typolevenshtein, usually after narrowing candidates
Correct spelling against a corpusSpellcheck
Match names that sound alikePhonetics
Search a whole dictionary for sound-alikesPhonetic neighbors
Match names across language boundaries (genealogy)Beider-Morse
Match names with a shared prefixjaro_winkler
Work out which encoder or stemmer a text even needsLanguage
Build an autocomplete indexTrie
Find repeated phrasesN-grams
Strip accents before comparingremove_diacritics
Put text in one canonical spelling before storing or hashing itnfc
Fold halfwidth katakana and fullwidth Latin onto one formnfkc
Romanise Japanese kanaTransliterators
Pluralise a noun, or write "23rd"Inflectors
Reduce words to a stemStemmers
Rank documents by relevanceTF-IDF
Classify documents into categoriesClassifiers
Score sentiment over a token listSentiment
Look up synonyms and word relationsWordNet
Tag parts of speechPOS tagger
Split a tagged sentence into subject and predicateSentence analyzers
Write generic code over any tokenizerCore vocabulary

The crates ​

SubsystemCratePublic surface
Tokenizersverbora-tokenizersWordTokenizer, SegmentTokenizer, SentenceTokenizer — UAX #29 boundaries, borrowed tokens
String distanceverbora-distance7 metrics; the three edit distances also in weighted and substring-search forms, plus PreparedPattern
Phoneticsverbora-phonetics12 encoders: SoundEx, Metaphone, Double Metaphone, Daitch–Mokotoff, Beider–Morse, Cologne, Nysiis, Caverphone 1/2, Phonex, Refined Soundex, Match Rating
Phonetic neighborsverbora-phoneticsPhoneticIndex — dictionary-wide candidate generation over SoundEx, Metaphone or DoubleMetaphone
Beider-Morseverbora-phoneticsCross-language surname matching over up to 18 languages at once, with auto-detection
Languageverbora-languageScript detection, optional statistical language detection, and recommend() — language → phonetic encoder
N-gramsverbora-ngramsngrams, Padded boundary symbols, char_ngrams
Normalizersverbora-normalizersthe four Unicode normalization forms plus remove_diacritics
Inflectorsverbora-inflectors6 inflectors, runtime rules
Trieverbora-trieprefix tree, prefix and path queries
Transliteratorsverbora-transliteratorsJapanese kana → romaji, one left-to-right pass over a generated mora index
WordNetverbora-wordnetlexical database, synsets, relation traversal, 4 storage strategies
TF-IDFverbora-tfidfTfIdf, TermScore, list_terms/tfidf/tfidfs, JSON persistence
Sentimentverbora-sentiment14 lexicons across 10 languages, sticky negation
Classifiersverbora-classifiersBayes, logistic regression, MaxEnt + GIS
Stemmersverbora-stemmers16 stemmers: Porter/Snowball across 12 languages, plus Carry (French), Lancaster, Japanese, Indonesian
Spellcheckverbora-spellcheckfrequency-ranked correction, BK-tree and deletion indexes
POS taggerverbora-taggerBrill tagging, training and evaluation over a lexicon you supply
Sentence analyzersverbora-analyzersphrase annotation, subject/predicate splitting, sentence type
Utilitiesverbora-utilstop words, abbreviations, weighted graphs, topological ordering and path trees
Core vocabularyverbora-core5 traits, StopWordLanguage, StopWords, the process-global stop-word list

Language support ​

Tokenization and normalization are not per-language axes any more: both follow the Unicode standard and therefore cover the whole character repertoire at once. The columns that do vary by language are the ones backed by per-language data.

LanguageInflectorPhonetics
English✅ nouns, verbs, ordinals✅ all four core encoders
French✅ nouns, ordinals
Japanese✅ nouns

Tokenization is UAX #29 word and sentence boundaries, so every language that separates words with spaces is covered by the same three tokenizers — with one stated limitation: the standard's default rules do not segment Thai, Lao, Khmer, Myanmar, Chinese or Japanese, and Verbora ships no dictionary segmenter for them.

Normalization is the four Unicode normalization forms plus a combining-mark fold defined over Canonical_Combining_Class, so remove_diacritics handles Latin, Greek, Cyrillic, Hebrew and Arabic script without a per-language rule set — read Normalizers before applying it to Thai or Devanagari.

Stemmers, sentiment lexicons and Beider-Morse each have their own language axis; see Stemmers, Sentiment and Beider-Morse. The POS tagger has none of its own: it ships no dictionary, so its language is whichever one the lexicon and rule set you give it were written for.

Where else to look ​

Released under the MIT License.