Skip to content

Installation ​

Verbora is a Cargo workspace of focused crates. Depend on the ones you need, not on a monolith: verbora-distance pulls in no data assets and no Unicode character database, and verbora-ngrams has no dependencies at all.

Requirements ​

Rust1.85 or newer (rust-version = "1.85")
Edition2024
PlatformAny Rust target the crate supports; no C runtime dependency
unsafeDenied workspace-wide (unsafe_code = "deny"). One narrow exception, test-only: verbora-spellcheck's counting_alloc module. No unsafe is compiled into any published library.

Add a crate ​

toml
[dependencies]
verbora-tokenizers = "0.2"
Upgrading from 0.1? 0.2.0 is a breaking release: a caret requirement of "0.1" will not resolve to it, every module path is gone in favour of the crate root, and several results changed without changing type. Upgrading from 0.1 to 0.2 is the old-signature-to-new-signature guide, including the changes that do not produce a compile error.
Pre-1.0, and versioned per crate. The examples on this site are written against 0.2, which is where seventeen of the nineteen crates are. verbora-wordnet and verbora-tagger are at 0.3: a crate moves only when its own API does, so the numbers are deliberately not in lockstep. If the version you want is not on crates.io yet, use the git or path form below — the code is identical either way.

From git:

toml
[dependencies]
verbora-tokenizers = { git = "https://github.com/addlayerio/verbora" }

From a local checkout:

toml
[dependencies]
verbora-tokenizers = { path = "../verbora/crates/verbora-tokenizers" }

Which crate do I need? ​

Each crate stands alone. verbora-core is pulled in transitively when a crate needs the shared traits or the stop-word tables; you only name it yourself if you are writing generic code or implementing the traits.

CrateDepend on it forPulls in
verbora-tokenizersword, segment and sentence tokenizerscore, unicode-segmentation
verbora-distanceLevenshtein, Damerau, OSA, Jaro–Winkler, Dice, Hammingrustc-hash
verbora-phoneticstwelve encoders plus PhoneticIndexcore, tokenizers, regex
verbora-ngramsn-gram windows, padding, character windowsnothing
verbora-normalizersthe four Unicode normalization forms, diacritic foldingunicode-normalization
verbora-inflectorspluralise/singularise, ordinals (en/fr/ja)regex
verbora-trieprefix tree, prefix search, path matching, FrozenTriesmallvec
verbora-corethe shared traits and StopWordsrustc-hash
verbora-stemmerssixteen stemmers across fourteen languagescore, tokenizers
verbora-spellcheckcorrection and fuzzy/deletion indexesdistance, rustc-hash
verbora-taggerBrill POS tagging, training and testing over a caller-supplied lexiconrustc-hash
verbora-analyzerssentence structure analysis over POS-tagged inputnothing
verbora-languagescript/language detection and phonetic recommendationsphonetics, transliterators, optional whatlang
verbora-transliteratorsJapanese kana-to-romaji transliterationnormalizers
verbora-wordnetWordNet lookup and relation traversalmemchr, rustc-hash
verbora-tfidfsparse TF-IDF indexing, querying and JSON persistencecore, tokenizers, rustc-hash, serde
verbora-sentimentmultilingual lexicon sentimentstemmers, tokenizers, rustc-hash
verbora-classifiersBayes, logistic regression and MaxEntstemmers, tokenizers, rustc-hash
verbora-utilstop words, abbreviations, graphs, path treescore, rustc-hash

rayon is absent from every row above because it is optional everywhere it is used — see Cargo features.

A typical text pipeline:

toml
[dependencies]
verbora-tokenizers = "0.2"
verbora-normalizers = "0.2"
verbora-ngrams = "0.2"

A fuzzy-matching service:

toml
[dependencies]
verbora-distance = "0.2"
verbora-phonetics = "0.2"
verbora-trie = "0.2"

Cargo features ​

Every optional feature is off by default, so a plain dependency stays sequential and pulls in nothing extra. Features add explicit parallel batch APIs and two language-detection implementations — see Cargo features.

Removed in 0.2.0. verbora-core declared a serde feature that nothing in the crate ever used — no item derived or implemented a serde trait. It is gone, and no crate in 0.2.0 has one. A dependency written as verbora-core = { version = "0.1", features = ["serde"] } fails to resolve once the pin moves to "0.2"; drop the features key. Serialization lives in the crates that actually do it — verbora-tfidf and verbora-classifiers — where serde is a plain dependency, not a feature.

Building from source ​

bash
git clone https://github.com/addlayerio/verbora
cd verbora

cargo build --workspace          # everything
cargo test  --workspace          # unit + integration + doctests
cargo bench -p verbora-distance  # Criterion benchmarks

The test suites replay large recorded corpora, so the workspace sets opt-level = 2 for the test profile and for every dependency in dev — an unoptimised debug build makes them unusably slow.

Next ​

Released under the MIT License.