The workspace map
Verbora is one Cargo workspace of 19 production crates. Knowing its shape makes the rest of this site easier to navigate.
What each crate is for
| Crate | Public surface |
|---|---|
verbora-core | 5 traits, StopWords, StopWordLanguage, the process-global stop-word list |
verbora-tokenizers | WordTokenizer, SegmentTokenizer, SentenceTokenizer |
verbora-distance | 7 metrics, weighted cost sets, substring search, PreparedPattern |
verbora-phonetics | 12 encoders, PhoneticIndex, BeiderMorse, phoneticize_tokens |
verbora-ngrams | ngrams, Padded, char_ngrams, CharNGrams |
verbora-normalizers | nfc, nfd, nfkc, nfkd, remove_diacritics |
verbora-inflectors | 6 inflectors, Rule, CaseMode, Gender |
verbora-trie | Trie, FrozenTrie, KeysWithPrefix, PrefixMatches |
verbora-transliterators | transliterate_ja, transliterate_ja_into, transliterate_ja_normalized, Rewrite |
verbora-wordnet | WordNet, Storage, Sense, Pointer, PrebuiltIndex, relation traversal |
verbora-tfidf | TfIdf, Document, Analyzer, DocumentScore, TermScore |
verbora-sentiment | SentimentAnalyzer, VocabularyKind, Vocabulary, Contributions |
verbora-classifiers | BayesClassifier, LogisticRegressionClassifier, MaxEntClassifier |
verbora-stemmers | Porter/Snowball language stemmers, Lancaster, Japanese and Indonesian |
verbora-spellcheck | Spellcheck, FuzzyIndex, DeletionIndex, Correction, Neighbor |
verbora-tagger | BrillTagger, Lexicon, RuleSet, Trainer, Evaluation |
verbora-analyzers | analyze, SentenceAnalysis, SentenceType, TaggedWord, Role |
verbora-language | script detection, optional language detectors, phonetic recommendations |
verbora-util | abbreviations, stop-word re-exports, graphs, path trees, topological ordering |
The dependency graph
Deliberately shallow and acyclic. Three crates reach sideways on purpose: verbora-phonetics tokenizes internally for tokenize_and_phoneticize, verbora-transliterators shares the normalizer's kana tables rather than duplicating them, and verbora-language composes the phonetic encoders with the transliterators to turn a detected language into a recommendation.
verbora-core ──┬── verbora-tokenizers ─── unicode-segmentation
├── verbora-phonetics ──── verbora-tokenizers, regex
├── verbora-stemmers ───── verbora-tokenizers
├── verbora-tfidf ──────── verbora-tokenizers, rustc-hash, serde
└── verbora-util ───────── rustc-hash
verbora-ngrams (no dependencies at all)
verbora-inflectors ────── regex
verbora-normalizers ───── unicode-normalization
verbora-distance ──────── rustc-hash
verbora-trie ──────────── smallvec
verbora-tagger ────────── rustc-hash
verbora-analyzers (no dependencies at all)
verbora-wordnet ───────── memchr, rustc-hash
verbora-transliterators ─ verbora-normalizers
verbora-spellcheck ────── verbora-distance, rustc-hash
verbora-sentiment ─────── verbora-stemmers, verbora-tokenizers, rustc-hash
verbora-classifiers ───── verbora-stemmers, verbora-tokenizers, rustc-hash
verbora-language ──────── verbora-phonetics, verbora-transliteratorsrayon is not drawn because it is optional in all fourteen crates that use it, and whatlang is optional in verbora-language — none of them is pulled in by a default dependency.
verbora-core reaches for one crate, rustc-hash, because StopWords::contains runs on every token of every document a stemmer filters. A leaf crate can be used in isolation without dragging in data assets or a regex engine it does not need: verbora-ngrams and verbora-analyzers ship an empty [dependencies] section, and verbora-distance reaches for nothing but rustc-hash, with no Unicode character database of any kind behind it.
The rest of the repository
site/ this site (VitePress)
docs/ internal engineering and design notes
tools/bench-data/ generates the inputs every benchmark harness reads
benches/data/ those generated inputs
benchmarks/competitive/ head-to-head suite against pinned third-party crates
crates/verbora-examples/ compiled documentation snippets (dev-only)verbora-examples is not a crate you depend on: it exists so that the code on this site is real. Every non-trivial snippet here is extracted into a generated example in that package and compiled and run against the actual crates — see Documentation is part of the code.
Next
- Cargo features — the four opt-ins and what they cost.
- Features overview — what each subsystem does and when to reach for it.