Upgrading from 0.1 to 0.2
Verbora 0.2.0 is not source-compatible with 0.1.0. Bumping a dependency from "0.1" to "0.2" will produce compile errors in most programs, and in a few places it will produce different answers without producing a compile error at all. This page is for fixing both.
verbora-wordnet 0.3.0 and verbora-tagger 0.3.0. The tagger one is the larger break: it no longer ships a dictionary. What happened
0.2.0 settles what defines Verbora's behaviour. Every behaviour is now derived from a published standard — UAX #29 for segmentation, UAX #15 for normalisation, Porter (1980), Brill (1992) — or from an explicit Verbora contract, and each is pinned by a test that asserts that contract directly.
Where 0.1 had a fixture recording an output, 0.2 has a rule stating what the output must be and why. That is a stricter standard, and applying it turned up defects a recorded fixture cannot detect: a recorded value agrees with itself forever, including when it is wrong. Most of the churn below is the consequence.
Six rules were applied across all nineteen crates:
- No sentinels. Absence is
Option::None, never a magic value carved out of a numeric range. - No
NaNescapes. It poisons comparison and sorting silently. - No panics outside preconditions the type system cannot express. Invalid states are made unrepresentable by
Result-returning constructors instead. - No function silently rewrites its input. Case folding, trimming and normalisation are the caller's explicit choice.
- The crate root is the entire public surface. Modules are private and everything is re-exported, so each item has exactly one path.
- The text unit is stated and justified per crate, never inherited.
Rule 5 alone breaks every use verbora_distance::levenshtein::…-style import in existing code. Rules 1–3 are why so many return types moved.
Fix the build first
Every module path is gone
This is the single most common error, and the fix is mechanical.
// 0.1
use verbora_distance::levenshtein::{Options, levenshtein};
use verbora_spellcheck::edits::edits;
use verbora_ngrams::text::ngrams_str;
use verbora_core::stopwords::StopWords;Import from the crate root instead. Every public item has exactly one path now, and it is verbora_<crate>::Item. If the root does not export it, it is no longer public — see the per-crate tables below for what replaced it.
The rename table for everything that survived
These items still exist, under a different name or a different signature. This is the list to scan first, because each row is a small edit rather than a redesign.
| Crate | 0.1 | 0.2 |
|---|---|---|
verbora-distance | damerau_levenshtein(a, b, &Options) → f64 | damerau_levenshtein(a, b) → usize |
levenshtein(a, b, &Options) → f64 | levenshtein(a, b) → usize | |
Options { restricted: true, .. } | osa(a, b) → usize — a separate function | |
Options { insertion_cost, .. } | LevenshteinCosts::new(…) / OsaCosts::new(…) / DamerauCosts::new(…), each returning Result<_, CostError>, passed to *_weighted | |
jaro_winkler(a, b, &Options) | jaro_winkler(a, b) | |
hamming(a, b, ignore_case) → i64, -1 for incomparable | hamming(a, b) → Option<usize> | |
hamming_checked, INCOMPARABLE | gone — hamming is the checked form | |
SearchResult { substring: String, distance: f64, offset: isize } | SearchResult<'t, D> with substring() -> &'t str, distance() -> D, range() -> Range<usize> (bytes) | |
StringMetric, Levenshtein, DamerauLevenshtein, JaroWinkler, Dice, Hamming | gone — call the free functions | |
verbora-core | StopWords::english() | StopWords::for_language(StopWordLanguage::En) |
StopWords::remove(&mut self, w) → () | → bool | |
StopWords::remove_all(…) → () | → usize | |
Token, collapse_whitespace, is_whitespace, trim_edge_empties | gone | |
DoubleKeyPhonetic::process_double → (String, String) | → (String, Option<String>) | |
verbora-trie | add_string, add_strings | insert, insert_all |
get_size() | node_count() — and len() is new, and counts words | |
is_case_sensitive() → bool | case_handling() → CaseHandling | |
with_case_sensitivity(bool) | with_case_handling(CaseHandling) | |
find_matches_on_path, iter_matches_on_path, MatchesOnPath | prefix_matches, iter_prefix_matches, PrefixMatches | |
find_prefix → (Option<Cow>, Cow) | longest_prefix → PrefixSplit { word, rest } | |
find_prefix_lengths → (Option<usize>, usize) | longest_prefix_lengths → PrefixSplitLengths { word, rest } | |
verbora-transliterators | transliterate, transliterate_into | transliterate_ja, transliterate_ja_into, transliterate_ja_normalized |
Phase | gone — the pipeline is internal | |
verbora-normalizers | normalize, normalize_token (English contraction expansion: "I'D" → ["I", "would"]) | gone — this crate is Unicode normalization and diacritic folding only |
normalize_ja, normalize_no, normalize_sv | gone — nfkc covers most of normalize_ja, but not all of it: width folding and composite symbols (㍼54年㋃㏪ → 昭和54年4月11日) match, while iteration marks (時々刻々 → 時時刻刻) and the small-tsu rewrite before the n-row (まっなか → まんなか) are not applied — nfkc returns those unchanged, so treat it as a replacement only if you did not rely on those two; normalize_no and normalize_sv existed for the Norwegian and Swedish stemmers, which is exactly where the umlaut-folding defect lived, and they are no longer public | |
verbora-inflectors | pluralize → Result<String, EmptyToken> | → String (total; the empty token has no inflected form and comes back unchanged) |
CountInflector, CountInflectorFr | OrdinalInflector, OrdinalInflectorFr (the French one takes a Gender) | |
CaseMode::{Lower, Capitalize, Upper} | CaseMode::{Preserve, Title, Upper} | |
restore_case(token) | CaseMode::of(token) | |
PatternError | RuleError (#[non_exhaustive]) | |
Rule::apply → Option<String> | → Option<Cow<'t, str>> | |
verbora-spellcheck | Spellcheck::get_corrections → Vec<String> | corrections → Vec<Correction<'_>>, or correction_words → Vec<String> |
par_get_corrections_batch | par_corrections_batch → Vec<Vec<Correction<'_>>> | |
frequency → Option<f64> | → Option<u32> | |
frequencies → (&str, f64) | → (&str, u32) | |
edits, edits_utf16, edits_with_max_distance*, Edits, EditUnit, ALPHABET, sort_by_frequency | gone — candidate generation is internal | |
Spellcheck::trie() | gone | |
DeletionIndex::neighbors → DeletionNeighbors | → Result<DeletionNeighbors, DistanceBeyondIndex> | |
verbora-phonetics | DoubleMetaphone::process → (String, String) | → DoubleMetaphoneCode, with primary(), alternate() -> Option<&str>, into_parts() |
SoundEx::process_with, Metaphone::process_with, *_utf16, try_process | gone — one key, one call, no max_length, no PhoneticError | |
SoundExDM | gone. 0.1 shipped two Daitch–Mokotoff implementations; DaitchMokotoff is the one that remains. They were separate code, so re-check your keys rather than assuming a rename | |
DaitchMokotoffCode | DaitchMokotoff::codes → Vec<String> | |
phoneticize_tokens_with, tokenize_and_phoneticize_with | folded into phoneticize_tokens / tokenize_and_phoneticize | |
verbora-sentiment | SentimentAnalyzer::new(lang: &str, stemmer, kind: &str) | SentimentAnalyzer::with_stemmer(Language, VocabularyKind, stemmer) |
without_stemmer(&str, &str) | without_stemmer(Language, VocabularyKind) | |
get_sentiment → f64 | → Option<f64> (None when no word scored) | |
Score::value → f64 (sum / count, so NaN when nothing scored) | Score::mean → Option<f64> — None is the empty case; Score::over is Option<f64> for the same reason | |
Polarity (an enum of Number/Text) | a struct: value() -> f64, as_written() -> Option<&'static str> | |
Error | UnsupportedPair / UnknownName | |
Vocabulary::shared_for | Vocabulary::shared(kind, Language) | |
Language::from_pattern, as_str | Language::from_code, code (plus FromStr) | |
verbora-language | f32 confidences | Confidence, built with Confidence::new(f32) -> Option<Self> |
PhoneticRecommendation::SoundExDaitchMokotoff | PhoneticRecommendation::DaitchMokotoff (plus new Cologne and BeiderMorse variants) | |
LanguageDetection { candidates } (public field) | candidates(), len(), is_empty(), single(…) | |
verbora-util | CyclicDependency | Cycle |
VertexKey, Vertex | VertexId | |
Bag, StorageBackend, FileBackend, StoragePlugin, StorageType | gone | |
Language (stop words) | still verbora_util::Language, now a re-export of verbora_core::StopWordLanguage | |
verbora-wordnet | Pos | PartOfSpeech |
Find, Source, Probe, Probes, FilePair, IndexHit, IndexRecord, DataRecord, PointerRef | gone or replaced by Synset, SynsetRef, IndexEntry, WordRef, Gloss | |
verbora-stemmers | verbora_stemmers::Token re-export | gone with verbora_core::Token |
stopwords::{add_all, contains, remove, …}(Language, …) | inherent methods on Language: Language::De.contains(w), .add_all(…), .reset() | |
TokenizeAndStem::is_word_char | gone (HYPHEN_JOINS_LETTERS is the remaining knob) |
The crates that were rewritten rather than renamed
For these, there is no row-by-row mapping worth printing, because almost nothing survived under a recognisable shape. Read the feature page and the rustdoc, and plan on rewriting the call site.
| Crate | What is there now |
|---|---|
verbora-tokenizers | Three tokenizers grounded in UAX #29 — WordTokenizer, SegmentTokenizer, SentenceTokenizer. The sixteen AggressiveTokenizer* language variants, TreebankWordTokenizer, RegexpTokenizer, WordPunctTokenizer, CaseTokenizer, OrthographyTokenizer, TokenizerJa and Utf16Token are all gone: their rules could not be traced to any standard. |
verbora-ngrams | ngrams(seq, n: NonZeroUsize) over a slice, Padded for boundary symbols, char_ngrams(text, n) for scalar windows. The *_str, *_with_stats, *_zh, bigrams/trigrams/multrigrams families, NGramStats, ngram_key and the process-global set_tokenizer are all gone. Word n-grams are now the composition you write: tokenize, then ngrams over the token slice. |
verbora-tagger | BrillTagger, Lexicon, RuleSet, Corpus, Trainer, Evaluation, Tag, TaggedToken. BrillPosTagger, BrillPosTrainer, BrillPosTester, TransformationRule, RuleTemplate, Predicate, TaggerError and the rest of the 0.1 surface are gone. |
verbora-tfidf | TfIdf with add_document(&str) -> usize, add_terms, tfidf(query, index) -> Option<f64>, rank(query) -> Vec<DocumentScore>, and an owned Analyzer holding the tokenizer, case folding and stop-word list. The dynamic DocKey/DynValue/JsonValue layer, Interner, TermId, Terms, Encoding, TfIdfError and the process-global tokenizer and stop-word setters are gone. to_json now returns Result<String, ExportError>; from_json returns Result<Self, RestoreError>. |
verbora-classifiers | BayesClassifier, LogisticRegressionClassifier, MaxEntClassifier over a reworked training and persistence surface (TrainingReport, TrainingStep, StopReason, ModelDefect, stamped model files). Context, Feature, FeatureSet, GenerateFeatures, GISScaler, MECorpus, MESentence and the rest are gone. Maximum entropy is now an implementation of Generalised Iterative Scaling rather than a reproduction; in 0.1 Sample::new rejected every non-empty argument, so the feature did not work at all. |
verbora-analyzers | analyze(&[TaggedWord]) -> SentenceAnalysis, with Role, TagClass, Terminator, SentenceType and ImpliedSubject. The mutable SentenceAnalyzer with its part()/type_of() staging, TaggedSentence, Punct, Field, SenType and TypeError are gone. |
verbora-wordnet | WordNet over Synset, SynsetRef, Sense, Pointer, PointerSymbol, Gloss, IndexEntry, PrebuiltIndex and four Storage modes. |
SenType did not simply get renamed. 0.1's SenType had five variants including Unknown and Command. 0.2's SentenceType has four — Declarative, Interrogative, Imperative, Exclamative — and absence is Option::None rather than an Unknown variant. A match carried across mechanically will compile and mean something different. The one break that is not in your code
Every fix above is something the compiler points at inside a .rs file. This one fails earlier, in Cargo.toml, before anything is compiled:
# 0.1
verbora-core = { version = "0.1", features = ["serde"] }
# 0.2 — drop the features key entirely
verbora-core = "0.2"verbora-core declared a serde feature that nothing in the crate ever used: no item derived or implemented a serde trait, and no dependent enabled it. It is gone, and no crate in 0.2.0 has a serde feature. Cargo reports this as the package 'verbora-core' does not have the feature 'serde', which does not name a version and so does not obviously point here.
The crates that do serialize — verbora-tfidf and verbora-classifiers — always used serde as a plain dependency rather than a feature, so their JSON persistence is unchanged.
Every other Cargo feature carried over: parallel on the fourteen crates that had it, and language-detection / fast-language-detection on verbora-language.
#[non_exhaustive], once and properly
34 public enums are #[non_exhaustive], up from 6 in 0.1.0. A downstream match over any of them now needs a wildcard arm, or it fails to compile with E0004: non-exhaustive patterns.
// 0.1: compiled, because the enum was closed.
match err {
CostError::NotFinite { .. } => …,
CostError::Negative { .. } => …,
CostError::TranspositionBelowThreshold { .. } => …,
}// 0.2: add the arm.
match err {
CostError::NotFinite { .. } => …,
CostError::Negative { .. } => …,
CostError::TranspositionBelowThreshold { .. } => …,
_ => …,
}The 21 error enums are the ones you are most likely to be matching on:
verbora_classifiers::{ClassifierError, LoadError, MaxEntError, ModelDefect, RestoreError, StampError} · verbora_distance::CostError · verbora_inflectors::RuleError · verbora_tagger::{CorpusParseError, LexiconError, LiteralError, RuleParseError} · verbora_tfidf::{ExportError, RestoreError, StampError} · verbora_tokenizers::AbbreviationError · verbora_util::{GraphError, PathError} · verbora_wordnet::{Error, ParseSenseError, RecordError}
The other 13 are the data enums those errors carry, or that a caller matches alongside them: verbora_core::StopWordLanguage · verbora_classifiers::TrainingEvent · verbora_distance::Operation · verbora_language::{Language, PhoneticRecommendation, Script, StrategyBasis, TransliterationAdvice} · verbora_sentiment::{Language, VocabularyKind} · verbora_tagger::{Condition, Template} · verbora_util::AbbreviationLanguage.
Why now, and why the payloads too
Pre-1.0 is the window where this costs a wildcard arm rather than a major version. An error enum that is closed cannot gain a variant without a breaking release, which in practice means either shipping a breaking release to describe a newly-distinguished failure or folding it into an existing variant and losing the distinction. Sealing them here buys the freedom to name failures precisely later, at the price of one arm today.
The 13 data enums are marked for the same reason, one level down. Sealing an error type while leaving the enum it carries closed gives the freedom straight back: a newly-distinguished failure almost always needs a new payload value to describe it — a new Operation, a new Language, a new Template — and adding one to a closed enum is the breaking change the outer seal was meant to avoid. They are marked together deliberately.
Three things this does not change:
if letandmatches!are unaffected unless they were exhaustive.- You can still construct these enums' variants from your own crate. The attribute is on the enum, not on its variants, so it constrains matching only:
CostError::Negative { operation, value }still compiles downstream. - A
matchthat already had a_arm, or that binds the whole value, compiles unchanged.
Changes that do not break the build
These compile after the mechanical fixes above and then return something different. Each one is a defect fix — 0.1's answer was wrong — but "wrong" is not the same as "not what your snapshot test asserts".
Ordering
Spellcheck::get_corrections returned Vec<String> ranked by a comparator whose order over f64 frequencies was never specified, and which disagreed with itself on ties. In 0.2 the ranking is distance ascending, then frequency descending, then word ascending, and it is written out as Correction's own Ord — hand-written rather than derived, because a derived Ord compares fields in declaration order and would put word first. Neighbor's order is distance ascending, then word ascending, on the same reasoning. Code that collected 0.1's results and re-sorted them gets a different order. use verbora_spellcheck::Spellcheck;
// Repeats are frequencies: "the" occurs three times, "he" twice, "she" once.
let sc = Spellcheck::new(["the", "the", "the", "he", "he", "she", "th"]);
// Distance first, then frequency descending: the exact match leads, and
// "the" outranks "she" at the same distance because it is more common.
assert_eq!(sc.correction_words("he", 1), ["he", "the", "she"]);
let best = sc.best_correction("he", 1).expect("a correction exists");
assert_eq!((best.word, best.distance, best.frequency), ("he", 0, 2));Frequencies are u32 rather than f64 for the same reason the ranking moved. A count is an integer, and 0.1's f64 frequencies were not merely imprecise: for twelve specific words the stored value was NaN, which a comparator silently reorders everything around. Spellcheck::frequency now returns Option<u32>, and frequencies() yields (&str, u32); there is no NaN left in this crate to guard against.
LogisticRegressionClassifier had a related defect: fit built its target columns in the feature map's enumeration order while classifications() reported weights in insertion order. Those orders differ exactly when a label looks like an integer, so a document trained as "2" classified as "1", and vice versa, confidently. Retrain any model with integer-like labels.
The unit of measurement
verbora-distance and verbora-stemmers counted UTF-16 code units in 0.1 and count Unicode scalar values in 0.2. verbora-trie counts scalars per node. Below U+10000 the two readings coincide; above it they do not.
use verbora_trie::Trie;
let mut t = Trie::new();
t.insert("a👍");
// 0.1's `get_size()` reported 4 here: root, 'a', and the two halves of the
// surrogate pair. One scalar is now one node, whatever plane it lives in.
assert_eq!(t.node_count(), 3);
assert_eq!(t.len(), 1); // and `len()` counts stored words, which is newStemmer output changes for input containing astral characters, for the same reason.
SearchResult carried the same unit problem and a sentinel besides: 0.1's offset was a UTF-16 code unit index into the target, signed, and genuinely -1 when the backtrace exited through column 0. 0.2 returns range() -> Range<usize> in bytes, derived from the borrowed substring, so &target[found.range()] == found.substring() for every input and there is no negative case to handle:
use verbora_distance::levenshtein_search;
let target = "Zürich, Berlin, Wien";
let found = levenshtein_search("Berlin", target);
// "Zürich, " is eight characters but nine bytes, because "ü" takes two.
assert_eq!(found.range(), 9..15);
assert_eq!(&target[found.range()], found.substring());
assert_eq!(found.distance(), 0);Behavioural fixes that change results
Each is now pinned by a test that fails without the fix, and each will move an output your program may be asserting on.
| Crate | What changed | What it means for you |
|---|---|---|
verbora-sentiment | The tokenizer's UAX #29 boundaries split hyphenated lexicon keys, making 2,313 entries unreachable, several of them sign-inverted — "non-approved" scored +1 against a stored polarity of -2. Multi-token span matching now reaches them, and the phrase keys single-token lookup never could. | Scores move, sometimes by a sign. Re-baseline any threshold tuned on 0.1. |
verbora-stemmers | Swedish and Norwegian folded a/o umlauts before consulting stop-word lists spelled with them, so 116 of 428 Swedish entries could never match. In those languages those are distinct letters. | Swedish and Norwegian stop-word filtering removes more than it used to. |
verbora-stemmers | The German stemmer's character gate was a byte-for-byte copy of the Spanish one: it admitted a/e/i/n/o/u accents and omitted a/o/u umlauts and eszett. | German stems change. |
verbora-stemmers | The Dutch stop-word list spelled an entry with a trailing space, so the pronoun je was never filtered. | One more Dutch stop word is filtered. |
verbora-tagger | Tag::new("*") is refused (Err(LiteralError::Wildcard)), because * printed as the wildcard and reparsed as TagPattern::Any — a rule that rewrote one tag became one that rewrote every tag, across the documented persistence path. Corpus::parse_brown inherits this as CorpusParseError::WildcardTag. | A corpus or rule file containing * as a literal tag is now an error instead of silently corrupting the rule set. |
verbora-tagger | Template::instantiate pushed one Condition per position inspected rather than one per distinct condition, so a window template double- (or triple-) counted a corrected token, defeating the trainer's min_score guard. | Trained rule sets differ. Retrain. |
verbora-transliterators | Vowel lengthening consumed the next scalar without asking whether it began a longer key, so ハロウィン came out harōin and スウェーデン came out sūēden. Six keys collide this way after any of seventy morae. | Romanisations change for the affected syllables. |
verbora-language | fold_cyrillic folded two blocks while the script router feeds it a third, so 100 uppercase letters went unfolded — including Ґ, one of four Ukrainian-versus-Russian discriminators. The same text in different case gave a different answer. | Cyrillic detection is now case-invariant, and uppercase Ukrainian text carries its signal. The Cyrillic model was retrained against the corrected extractor. |
verbora-phonetics | Four Beider–Morse Italian rules carried U+FFFD where accented vowels belonged, so real Italian input had the vowel deleted from its encoding. | Beider–Morse encodings change for the affected Italian inputs. |
verbora-distance | jaro(x, x) returned 0.0 for single-unit inputs while jaro_winkler returned 1.0 for the same pair; dice_coefficient("", "") returned NaN. Search could return text absent from the target, because a UTF-16 slice could split a surrogate pair and from_utf16_lossy substituted U+FFFD. | Those cases now return the documented answers, and SearchResult::substring() is a borrow of the target, so it cannot be text the target does not contain. |
verbora-tfidf | The between-documents invariant was restored in finish, which unwinding skips, so a panicking caller iterator left the counter dirty and the next document reported a present term as Some(0) — documented to mean absent. | Corpora built through a panicking tokenizer were wrong; they are not now. |
Resource behaviour
Not a correctness change, but it will change how your program behaves under load:
verbora-spellcheck's deletion generation was cubic in word length. An 800-scalar token cost 4.0 GB of peak RSS; it now costs 32.4 MB.k = 0, documented as a membership test, no longer builds an index at all.DeletionIndexwas re-keyed onto au64hash for the same reason — cubic to quadratic in word length. The index itself stays quadratic in the longest word: that is the symmetric-delete structure's own price, now documented.verbora-trie'sinsert_allfedsize_hint().0straight toVec::reserve, so an iterator that lies — whichsize_hintexplicitly permits — could abort with a capacity overflow or reserve tens of gigabytes. The hint is now clamped.verbora-wordnet'sResidentandLazyResidentallocated by file metadata with no ceiling. The dictionary file is caller-supplied, so its size was always an input; the three preloading storage modes now share the ceilingIndexedalready had. (Storage::Preadpreloads nothing and never had the problem.)MaxEntClassifier::restore's JSON parser descended one stack frame per nesting level with no bound, so a model file of 20,000 nested brackets overflowed the stack — which aborts rather than unwinding, so no caller could catch it. It is bounded at 128 levels.
What this page does not tell you
It also does not cover the places where only the justification moved. Several claims in 0.1's documentation were wrong about code that is unchanged, and several fixtures were re-derived from the publications they should always have come from without any value moving:
- An empty phonetic key does not imply the input had no recognised letter (
Metaphoneon"y",Cologneon"h",MatchRatingApproachon any single vowel), andDaitchMokotoffdoes advance by byte length past a matched pattern, which consumes one character too many for four two-byte keys. - A shortest-path tie-break follows relaxation order, not insertion order, and 149 French inflector entries described as load-bearing are reached by no rule.
- Cologne's 170 expected values were re-derived from Postel (1969) and Match Rating's from NBS Special Publication 500-2 (1977). Where that paper reads its two comparison passes as consecutive, Verbora interleaves them — a difference of 31,044 ratings in 1,129,284 pairs, now stated as a Verbora decision rather than inherited. The implementation did not change; the claim about it did.
If you were relying on one of those claims rather than on the code, your program was already doing something else.
verbora-wordnet 0.3.0
Crates version independently: one moves when its own public API does. Two have moved to 0.3; the other seventeen are at 0.2.
verbora-wordnet 0.2.0 could not read 8.8% of a real WordNet dictionary — 13,606 of 155,467 index entries, including run, cat, light, water, computer and node. PointerSymbol::from_symbol accepted the domain pointers only in their qualified spellings (;c, -c, ;r, -r, ;u, -u) and rejected the bare ; and - that Princeton's index.* files actually write — the class letter belongs to the sense, so it appears only in data.*. index_entry, senses, sense and lookup all failed on any affected lemma.
If you are on 0.2.0 and WordNet lookups were failing for common words, that is why. Move to "0.3".
The fix is breaking in one way:
// 0.2 — compiles
match symbol {
PointerSymbol::Hypernym => "…",
// … every other variant …
}
// 0.3 — `PointerSymbol` gained `Domain` (`;`) and `Member` (`-`) and is
// `#[non_exhaustive]`, so a wildcard arm is required.
match symbol {
PointerSymbol::Hypernym => "…",
// … every other variant …
_ => "…",
}Domain and Member are relations in their own right rather than being folded into one of the three classes: an index entry states that a domain relation exists and deliberately does not state which kind, and choosing one would be a claim the file does not make. They appear only from index entries — a data record that omits the class letter is malformed and is rejected.
verbora-tagger 0.3.0
verbora-tagger 0.2.0 shipped four data files it had no right to redistribute: an English lexicon and rule set that are LGPL-3.0, which is incompatible with this project's MIT licence, and a Dutch lexicon and context-rule set whose terms could not be located at all. All four were removed in 0.3.0. Attribution does not fix a licence mismatch, and the crate is not going to ship data it cannot account for.
What that means for your code: the crate no longer knows any language. It is a tagging engine, and the lexicon is now an input you supply.
// 0.2 — a tagger that arrived knowing English.
use verbora_tagger::{BrillTagger, Language, Lexicon, RuleSet};
let lexicon = Lexicon::bundled(Language::English);
let rules = RuleSet::bundled(Language::English);
let tagger = BrillTagger::new(&lexicon, &rules);// 0.3 — you bring the dictionary. `Corpus` counts the tag frequencies for you.
use verbora_tagger::{BrillTagger, Corpus, RuleSet, Tag};
let corpus = Corpus::parse_brown(&std::fs::read_to_string("brown.txt")?)?;
let lexicon = corpus.build_lexicon(Tag::new("NN")?)?;
let rules: RuleSet = "NN VB PREV-TAG MD".parse()?;
let tagger = BrillTagger::new(&lexicon, &rules);Lexicon::new(default_tag) plus Lexicon::insert is the other way in, when the entries are yours to write down rather than to count, and Trainer learns a rule set from the same corpus. Both paths are shown, compiled, on the POS tagger page.
The removals:
| 0.2 | 0.3 |
|---|---|
Lexicon::bundled(Language) | gone — build one with Corpus::build_lexicon, or Lexicon::new plus Lexicon::insert |
RuleSet::bundled(Language::English) | RuleSet::brill_1992() — the same ten published rules of Brill (1992), Table 1, under a name that carries their citation and without the Language parameter |
RuleSet::bundled(Language::Dutch) | gone — no replacement; train one with Trainer against your own corpus |
brill_paper_rule_strings() | gone — RuleSet::brill_1992() is the same table, and RuleSet::to_string() recovers the rule strings verbatim |
Language | gone — every method on it named data that no longer ships, and its defaults (NN, NNP) were Penn Treebank tags while the one surviving rule set is Brown |
RuleSet::brill_1992 is written in Brown corpus tags. It names AT, PPS, PPO, HVD and NP — not the Penn Treebank DT, PRP, VBD, NNP. Verbora attaches no meaning to a tag beyond string identity, so pairing it with a Penn-tagged lexicon is not an error: the rules simply never fire, the tagger costs a pass per rule, and it returns the initial-state annotation unchanged. If you build a Penn-tagged lexicon, train your own rules against it rather than reaching for this set. If you were relying on the bundled English tagger and have no corpus of your own, the shape of the work is: obtain an annotated corpus under terms you are happy with, run it through Corpus::parse_brown and Corpus::build_lexicon, and train rules with Trainer. That is a deliberate trade — the crate stopped making a licensing decision on your behalf.
If you are stuck
- Check the crate's rustdoc first. The crate root is the whole public surface now, so
cargo doc --open -p verbora-<crate>shows you everything in one list. - Check the crate's
README.md. Every crate has one as of 0.2.0, it is that crate's crates.io landing page, and its examples are compiled as doctests. - Check the feature page for the subsystem — linked from Features — for what the replacement is for, not just what it is called.
- Report a documentation error as a bug. If this page sent you somewhere that does not exist, that is worth an issue: the repository.
Next
- Installation — the version pins and the crate table.
- Your first program — the current API, four ways.
- Choosing the right API — what to reach for now.