Installation
Verbora is a Cargo workspace of focused crates. Depend on the ones you need, not on a monolith: verbora-distance pulls in no data assets and no Unicode character database, and verbora-ngrams has no dependencies at all.
Requirements
| Rust | 1.85 or newer (rust-version = "1.85") |
| Edition | 2024 |
| Platform | Any Rust target the crate supports; no C runtime dependency |
unsafe | Denied workspace-wide (unsafe_code = "deny"). One narrow exception, test-only: verbora-spellcheck's counting_alloc module. No unsafe is compiled into any published library. |
Add a crate
[dependencies]
verbora-tokenizers = "0.2"0.2.0 is a breaking release: a caret requirement of "0.1" will not resolve to it, every module path is gone in favour of the crate root, and several results changed without changing type. Upgrading from 0.1 to 0.2 is the old-signature-to-new-signature guide, including the changes that do not produce a compile error. 0.2, which is where seventeen of the nineteen crates are. verbora-wordnet and verbora-tagger are at 0.3: a crate moves only when its own API does, so the numbers are deliberately not in lockstep. If the version you want is not on crates.io yet, use the git or path form below — the code is identical either way. From git:
[dependencies]
verbora-tokenizers = { git = "https://github.com/addlayerio/verbora" }From a local checkout:
[dependencies]
verbora-tokenizers = { path = "../verbora/crates/verbora-tokenizers" }Which crate do I need?
Each crate stands alone. verbora-core is pulled in transitively when a crate needs the shared traits or the stop-word tables; you only name it yourself if you are writing generic code or implementing the traits.
| Crate | Depend on it for | Pulls in |
|---|---|---|
verbora-tokenizers | word, segment and sentence tokenizers | core, unicode-segmentation |
verbora-distance | Levenshtein, Damerau, OSA, Jaro–Winkler, Dice, Hamming | rustc-hash |
verbora-phonetics | twelve encoders plus PhoneticIndex | core, tokenizers, regex |
verbora-ngrams | n-gram windows, padding, character windows | nothing |
verbora-normalizers | the four Unicode normalization forms, diacritic folding | unicode-normalization |
verbora-inflectors | pluralise/singularise, ordinals (en/fr/ja) | regex |
verbora-trie | prefix tree, prefix search, path matching, FrozenTrie | smallvec |
verbora-core | the shared traits and StopWords | rustc-hash |
verbora-stemmers | sixteen stemmers across fourteen languages | core, tokenizers |
verbora-spellcheck | correction and fuzzy/deletion indexes | distance, rustc-hash |
verbora-tagger | Brill POS tagging, training and testing over a caller-supplied lexicon | rustc-hash |
verbora-analyzers | sentence structure analysis over POS-tagged input | nothing |
verbora-language | script/language detection and phonetic recommendations | phonetics, transliterators, optional whatlang |
verbora-transliterators | Japanese kana-to-romaji transliteration | normalizers |
verbora-wordnet | WordNet lookup and relation traversal | memchr, rustc-hash |
verbora-tfidf | sparse TF-IDF indexing, querying and JSON persistence | core, tokenizers, rustc-hash, serde |
verbora-sentiment | multilingual lexicon sentiment | stemmers, tokenizers, rustc-hash |
verbora-classifiers | Bayes, logistic regression and MaxEnt | stemmers, tokenizers, rustc-hash |
verbora-util | stop words, abbreviations, graphs, path trees | core, rustc-hash |
rayon is absent from every row above because it is optional everywhere it is used — see Cargo features.
A typical text pipeline:
[dependencies]
verbora-tokenizers = "0.2"
verbora-normalizers = "0.2"
verbora-ngrams = "0.2"A fuzzy-matching service:
[dependencies]
verbora-distance = "0.2"
verbora-phonetics = "0.2"
verbora-trie = "0.2"Cargo features
Every optional feature is off by default, so a plain dependency stays sequential and pulls in nothing extra. Features add explicit parallel batch APIs and two language-detection implementations — see Cargo features.
verbora-core declared a serde feature that nothing in the crate ever used — no item derived or implemented a serde trait. It is gone, and no crate in 0.2.0 has one. A dependency written as verbora-core = { version = "0.1", features = ["serde"] } fails to resolve once the pin moves to "0.2"; drop the features key. Serialization lives in the crates that actually do it — verbora-tfidf and verbora-classifiers — where serde is a plain dependency, not a feature. Building from source
git clone https://github.com/addlayerio/verbora
cd verbora
cargo build --workspace # everything
cargo test --workspace # unit + integration + doctests
cargo bench -p verbora-distance # Criterion benchmarksThe test suites replay large recorded corpora, so the workspace sets opt-level = 2 for the test profile and for every dependency in dev — an unoptimised debug build makes them unusably slow.
Next
- Your first program — the same task written four ways, one per API shape.
- The workspace map — what lives where, and why the crates are split the way they are.
- Cargo features — parallel batch APIs and the other opt-ins.
- Upgrading from 0.1 to 0.2 — what broke, what silently changed, and what to write instead.