Skip to content

Upgrading from 0.1 to 0.2 ​

Verbora 0.2.0 is not source-compatible with 0.1.0. Bumping a dependency from "0.1" to "0.2" will produce compile errors in most programs, and in a few places it will produce different answers without producing a compile error at all. This page is for fixing both.

Read the second half even if the first half fixes your build. The changes that do not break compilation are the dangerous ones — a different ranking, a different tag, a different sentiment score. They are listed under Changes that do not break the build.
Already on 0.2? Two crates have moved past it on their own, and each has its own section at the end of this page: verbora-wordnet 0.3.0 and verbora-tagger 0.3.0. The tagger one is the larger break: it no longer ships a dictionary.

What happened ​

0.2.0 settles what defines Verbora's behaviour. Every behaviour is now derived from a published standard — UAX #29 for segmentation, UAX #15 for normalisation, Porter (1980), Brill (1992) — or from an explicit Verbora contract, and each is pinned by a test that asserts that contract directly.

Where 0.1 had a fixture recording an output, 0.2 has a rule stating what the output must be and why. That is a stricter standard, and applying it turned up defects a recorded fixture cannot detect: a recorded value agrees with itself forever, including when it is wrong. Most of the churn below is the consequence.

Six rules were applied across all nineteen crates:

  1. No sentinels. Absence is Option::None, never a magic value carved out of a numeric range.
  2. No NaN escapes. It poisons comparison and sorting silently.
  3. No panics outside preconditions the type system cannot express. Invalid states are made unrepresentable by Result-returning constructors instead.
  4. No function silently rewrites its input. Case folding, trimming and normalisation are the caller's explicit choice.
  5. The crate root is the entire public surface. Modules are private and everything is re-exported, so each item has exactly one path.
  6. The text unit is stated and justified per crate, never inherited.

Rule 5 alone breaks every use verbora_distance::levenshtein::…-style import in existing code. Rules 1–3 are why so many return types moved.

Fix the build first ​

Every module path is gone ​

This is the single most common error, and the fix is mechanical.

rust
// 0.1
use verbora_distance::levenshtein::{Options, levenshtein};
use verbora_spellcheck::edits::edits;
use verbora_ngrams::text::ngrams_str;
use verbora_core::stopwords::StopWords;

Import from the crate root instead. Every public item has exactly one path now, and it is verbora_<crate>::Item. If the root does not export it, it is no longer public — see the per-crate tables below for what replaced it.

The rename table for everything that survived ​

These items still exist, under a different name or a different signature. This is the list to scan first, because each row is a small edit rather than a redesign.

Crate0.10.2
verbora-distancedamerau_levenshtein(a, b, &Options) → f64damerau_levenshtein(a, b) → usize
levenshtein(a, b, &Options) → f64levenshtein(a, b) → usize
Options { restricted: true, .. }osa(a, b) → usize — a separate function
Options { insertion_cost, .. }LevenshteinCosts::new(…) / OsaCosts::new(…) / DamerauCosts::new(…), each returning Result<_, CostError>, passed to *_weighted
jaro_winkler(a, b, &Options)jaro_winkler(a, b)
hamming(a, b, ignore_case) → i64, -1 for incomparablehamming(a, b) → Option<usize>
hamming_checked, INCOMPARABLEgone — hamming is the checked form
SearchResult { substring: String, distance: f64, offset: isize }SearchResult<'t, D> with substring() -> &'t str, distance() -> D, range() -> Range<usize> (bytes)
StringMetric, Levenshtein, DamerauLevenshtein, JaroWinkler, Dice, Hamminggone — call the free functions
verbora-coreStopWords::english()StopWords::for_language(StopWordLanguage::En)
StopWords::remove(&mut self, w) → ()→ bool
StopWords::remove_all(…) → ()→ usize
Token, collapse_whitespace, is_whitespace, trim_edge_emptiesgone
DoubleKeyPhonetic::process_double → (String, String)→ (String, Option<String>)
verbora-trieadd_string, add_stringsinsert, insert_all
get_size()node_count() — and len() is new, and counts words
is_case_sensitive() → boolcase_handling() → CaseHandling
with_case_sensitivity(bool)with_case_handling(CaseHandling)
find_matches_on_path, iter_matches_on_path, MatchesOnPathprefix_matches, iter_prefix_matches, PrefixMatches
find_prefix → (Option<Cow>, Cow)longest_prefix → PrefixSplit { word, rest }
find_prefix_lengths → (Option<usize>, usize)longest_prefix_lengths → PrefixSplitLengths { word, rest }
verbora-transliteratorstransliterate, transliterate_intotransliterate_ja, transliterate_ja_into, transliterate_ja_normalized
Phasegone — the pipeline is internal
verbora-normalizersnormalize, normalize_token (English contraction expansion: "I'D" → ["I", "would"])gone — this crate is Unicode normalization and diacritic folding only
normalize_ja, normalize_no, normalize_svgone — nfkc covers most of normalize_ja, but not all of it: width folding and composite symbols (㍼54年㋃㏪ → 昭和54年4月11日) match, while iteration marks (時々刻々 → 時時刻刻) and the small-tsu rewrite before the n-row (まっなか → まんなか) are not applied — nfkc returns those unchanged, so treat it as a replacement only if you did not rely on those two; normalize_no and normalize_sv existed for the Norwegian and Swedish stemmers, which is exactly where the umlaut-folding defect lived, and they are no longer public
verbora-inflectorspluralize → Result<String, EmptyToken>→ String (total; the empty token has no inflected form and comes back unchanged)
CountInflector, CountInflectorFrOrdinalInflector, OrdinalInflectorFr (the French one takes a Gender)
CaseMode::{Lower, Capitalize, Upper}CaseMode::{Preserve, Title, Upper}
restore_case(token)CaseMode::of(token)
PatternErrorRuleError (#[non_exhaustive])
Rule::apply → Option<String>→ Option<Cow<'t, str>>
verbora-spellcheckSpellcheck::get_corrections → Vec<String>corrections → Vec<Correction<'_>>, or correction_words → Vec<String>
par_get_corrections_batchpar_corrections_batch → Vec<Vec<Correction<'_>>>
frequency → Option<f64>→ Option<u32>
frequencies → (&str, f64)→ (&str, u32)
edits, edits_utf16, edits_with_max_distance*, Edits, EditUnit, ALPHABET, sort_by_frequencygone — candidate generation is internal
Spellcheck::trie()gone
DeletionIndex::neighbors → DeletionNeighbors→ Result<DeletionNeighbors, DistanceBeyondIndex>
verbora-phoneticsDoubleMetaphone::process → (String, String)→ DoubleMetaphoneCode, with primary(), alternate() -> Option<&str>, into_parts()
SoundEx::process_with, Metaphone::process_with, *_utf16, try_processgone — one key, one call, no max_length, no PhoneticError
SoundExDMgone. 0.1 shipped two Daitch–Mokotoff implementations; DaitchMokotoff is the one that remains. They were separate code, so re-check your keys rather than assuming a rename
DaitchMokotoffCodeDaitchMokotoff::codes → Vec<String>
phoneticize_tokens_with, tokenize_and_phoneticize_withfolded into phoneticize_tokens / tokenize_and_phoneticize
verbora-sentimentSentimentAnalyzer::new(lang: &str, stemmer, kind: &str)SentimentAnalyzer::with_stemmer(Language, VocabularyKind, stemmer)
without_stemmer(&str, &str)without_stemmer(Language, VocabularyKind)
get_sentiment → f64→ Option<f64> (None when no word scored)
Score::value → f64 (sum / count, so NaN when nothing scored)Score::mean → Option<f64> — None is the empty case; Score::over is Option<f64> for the same reason
Polarity (an enum of Number/Text)a struct: value() -> f64, as_written() -> Option<&'static str>
ErrorUnsupportedPair / UnknownName
Vocabulary::shared_forVocabulary::shared(kind, Language)
Language::from_pattern, as_strLanguage::from_code, code (plus FromStr)
verbora-languagef32 confidencesConfidence, built with Confidence::new(f32) -> Option<Self>
PhoneticRecommendation::SoundExDaitchMokotoffPhoneticRecommendation::DaitchMokotoff (plus new Cologne and BeiderMorse variants)
LanguageDetection { candidates } (public field)candidates(), len(), is_empty(), single(…)
verbora-utilCyclicDependencyCycle
VertexKey, VertexVertexId
Bag, StorageBackend, FileBackend, StoragePlugin, StorageTypegone
Language (stop words)still verbora_util::Language, now a re-export of verbora_core::StopWordLanguage
verbora-wordnetPosPartOfSpeech
Find, Source, Probe, Probes, FilePair, IndexHit, IndexRecord, DataRecord, PointerRefgone or replaced by Synset, SynsetRef, IndexEntry, WordRef, Gloss
verbora-stemmersverbora_stemmers::Token re-exportgone with verbora_core::Token
stopwords::{add_all, contains, remove, …}(Language, …)inherent methods on Language: Language::De.contains(w), .add_all(…), .reset()
TokenizeAndStem::is_word_chargone (HYPHEN_JOINS_LETTERS is the remaining knob)

The crates that were rewritten rather than renamed ​

For these, there is no row-by-row mapping worth printing, because almost nothing survived under a recognisable shape. Read the feature page and the rustdoc, and plan on rewriting the call site.

CrateWhat is there now
verbora-tokenizersThree tokenizers grounded in UAX #29 — WordTokenizer, SegmentTokenizer, SentenceTokenizer. The sixteen AggressiveTokenizer* language variants, TreebankWordTokenizer, RegexpTokenizer, WordPunctTokenizer, CaseTokenizer, OrthographyTokenizer, TokenizerJa and Utf16Token are all gone: their rules could not be traced to any standard.
verbora-ngramsngrams(seq, n: NonZeroUsize) over a slice, Padded for boundary symbols, char_ngrams(text, n) for scalar windows. The *_str, *_with_stats, *_zh, bigrams/trigrams/multrigrams families, NGramStats, ngram_key and the process-global set_tokenizer are all gone. Word n-grams are now the composition you write: tokenize, then ngrams over the token slice.
verbora-taggerBrillTagger, Lexicon, RuleSet, Corpus, Trainer, Evaluation, Tag, TaggedToken. BrillPosTagger, BrillPosTrainer, BrillPosTester, TransformationRule, RuleTemplate, Predicate, TaggerError and the rest of the 0.1 surface are gone.
verbora-tfidfTfIdf with add_document(&str) -> usize, add_terms, tfidf(query, index) -> Option<f64>, rank(query) -> Vec<DocumentScore>, and an owned Analyzer holding the tokenizer, case folding and stop-word list. The dynamic DocKey/DynValue/JsonValue layer, Interner, TermId, Terms, Encoding, TfIdfError and the process-global tokenizer and stop-word setters are gone. to_json now returns Result<String, ExportError>; from_json returns Result<Self, RestoreError>.
verbora-classifiersBayesClassifier, LogisticRegressionClassifier, MaxEntClassifier over a reworked training and persistence surface (TrainingReport, TrainingStep, StopReason, ModelDefect, stamped model files). Context, Feature, FeatureSet, GenerateFeatures, GISScaler, MECorpus, MESentence and the rest are gone. Maximum entropy is now an implementation of Generalised Iterative Scaling rather than a reproduction; in 0.1 Sample::new rejected every non-empty argument, so the feature did not work at all.
verbora-analyzersanalyze(&[TaggedWord]) -> SentenceAnalysis, with Role, TagClass, Terminator, SentenceType and ImpliedSubject. The mutable SentenceAnalyzer with its part()/type_of() staging, TaggedSentence, Punct, Field, SenType and TypeError are gone.
verbora-wordnetWordNet over Synset, SynsetRef, Sense, Pointer, PointerSymbol, Gloss, IndexEntry, PrebuiltIndex and four Storage modes.
SenType did not simply get renamed. 0.1's SenType had five variants including Unknown and Command. 0.2's SentenceType has four — Declarative, Interrogative, Imperative, Exclamative — and absence is Option::None rather than an Unknown variant. A match carried across mechanically will compile and mean something different.

The one break that is not in your code ​

Every fix above is something the compiler points at inside a .rs file. This one fails earlier, in Cargo.toml, before anything is compiled:

toml
# 0.1
verbora-core = { version = "0.1", features = ["serde"] }
# 0.2 — drop the features key entirely
verbora-core = "0.2"

verbora-core declared a serde feature that nothing in the crate ever used: no item derived or implemented a serde trait, and no dependent enabled it. It is gone, and no crate in 0.2.0 has a serde feature. Cargo reports this as the package 'verbora-core' does not have the feature 'serde', which does not name a version and so does not obviously point here.

The crates that do serialize — verbora-tfidf and verbora-classifiers — always used serde as a plain dependency rather than a feature, so their JSON persistence is unchanged.

Every other Cargo feature carried over: parallel on the fourteen crates that had it, and language-detection / fast-language-detection on verbora-language.

#[non_exhaustive], once and properly ​

34 public enums are #[non_exhaustive], up from 6 in 0.1.0. A downstream match over any of them now needs a wildcard arm, or it fails to compile with E0004: non-exhaustive patterns.

rust
// 0.1: compiled, because the enum was closed.
match err {
    CostError::NotFinite { .. } => …,
    CostError::Negative { .. } => …,
    CostError::TranspositionBelowThreshold { .. } => …,
}
rust
// 0.2: add the arm.
match err {
    CostError::NotFinite { .. } => …,
    CostError::Negative { .. } => …,
    CostError::TranspositionBelowThreshold { .. } => …,
    _ => …,
}

The 21 error enums are the ones you are most likely to be matching on:

verbora_classifiers::{ClassifierError, LoadError, MaxEntError, ModelDefect, RestoreError, StampError} · verbora_distance::CostError · verbora_inflectors::RuleError · verbora_tagger::{CorpusParseError, LexiconError, LiteralError, RuleParseError} · verbora_tfidf::{ExportError, RestoreError, StampError} · verbora_tokenizers::AbbreviationError · verbora_util::{GraphError, PathError} · verbora_wordnet::{Error, ParseSenseError, RecordError}

The other 13 are the data enums those errors carry, or that a caller matches alongside them: verbora_core::StopWordLanguage · verbora_classifiers::TrainingEvent · verbora_distance::Operation · verbora_language::{Language, PhoneticRecommendation, Script, StrategyBasis, TransliterationAdvice} · verbora_sentiment::{Language, VocabularyKind} · verbora_tagger::{Condition, Template} · verbora_util::AbbreviationLanguage.

Why now, and why the payloads too ​

Pre-1.0 is the window where this costs a wildcard arm rather than a major version. An error enum that is closed cannot gain a variant without a breaking release, which in practice means either shipping a breaking release to describe a newly-distinguished failure or folding it into an existing variant and losing the distinction. Sealing them here buys the freedom to name failures precisely later, at the price of one arm today.

The 13 data enums are marked for the same reason, one level down. Sealing an error type while leaving the enum it carries closed gives the freedom straight back: a newly-distinguished failure almost always needs a new payload value to describe it — a new Operation, a new Language, a new Template — and adding one to a closed enum is the breaking change the outer seal was meant to avoid. They are marked together deliberately.

Three things this does not change:

  • if let and matches! are unaffected unless they were exhaustive.
  • You can still construct these enums' variants from your own crate. The attribute is on the enum, not on its variants, so it constrains matching only: CostError::Negative { operation, value } still compiles downstream.
  • A match that already had a _ arm, or that binds the whole value, compiles unchanged.

Changes that do not break the build ​

These compile after the mechanical fixes above and then return something different. Each one is a defect fix — 0.1's answer was wrong — but "wrong" is not the same as "not what your snapshot test asserts".

Ordering ​

Correction ranking is a documented total order now, and re-sorting the results cannot disagree with it. In 0.1, Spellcheck::get_corrections returned Vec<String> ranked by a comparator whose order over f64 frequencies was never specified, and which disagreed with itself on ties. In 0.2 the ranking is distance ascending, then frequency descending, then word ascending, and it is written out as Correction's own Ord — hand-written rather than derived, because a derived Ord compares fields in declaration order and would put word first. Neighbor's order is distance ascending, then word ascending, on the same reasoning. Code that collected 0.1's results and re-sorted them gets a different order.
rust
use verbora_spellcheck::Spellcheck;

// Repeats are frequencies: "the" occurs three times, "he" twice, "she" once.
let sc = Spellcheck::new(["the", "the", "the", "he", "he", "she", "th"]);

// Distance first, then frequency descending: the exact match leads, and
// "the" outranks "she" at the same distance because it is more common.
assert_eq!(sc.correction_words("he", 1), ["he", "the", "she"]);

let best = sc.best_correction("he", 1).expect("a correction exists");
assert_eq!((best.word, best.distance, best.frequency), ("he", 0, 2));

Frequencies are u32 rather than f64 for the same reason the ranking moved. A count is an integer, and 0.1's f64 frequencies were not merely imprecise: for twelve specific words the stored value was NaN, which a comparator silently reorders everything around. Spellcheck::frequency now returns Option<u32>, and frequencies() yields (&str, u32); there is no NaN left in this crate to guard against.

LogisticRegressionClassifier had a related defect: fit built its target columns in the feature map's enumeration order while classifications() reported weights in insertion order. Those orders differ exactly when a label looks like an integer, so a document trained as "2" classified as "1", and vice versa, confidently. Retrain any model with integer-like labels.

The unit of measurement ​

verbora-distance and verbora-stemmers counted UTF-16 code units in 0.1 and count Unicode scalar values in 0.2. verbora-trie counts scalars per node. Below U+10000 the two readings coincide; above it they do not.

rust
use verbora_trie::Trie;

let mut t = Trie::new();
t.insert("a👍");

// 0.1's `get_size()` reported 4 here: root, 'a', and the two halves of the
// surrogate pair. One scalar is now one node, whatever plane it lives in.
assert_eq!(t.node_count(), 3);
assert_eq!(t.len(), 1); // and `len()` counts stored words, which is new

Stemmer output changes for input containing astral characters, for the same reason.

SearchResult carried the same unit problem and a sentinel besides: 0.1's offset was a UTF-16 code unit index into the target, signed, and genuinely -1 when the backtrace exited through column 0. 0.2 returns range() -> Range<usize> in bytes, derived from the borrowed substring, so &target[found.range()] == found.substring() for every input and there is no negative case to handle:

rust
use verbora_distance::levenshtein_search;

let target = "Zürich, Berlin, Wien";
let found = levenshtein_search("Berlin", target);

// "Zürich, " is eight characters but nine bytes, because "ü" takes two.
assert_eq!(found.range(), 9..15);
assert_eq!(&target[found.range()], found.substring());
assert_eq!(found.distance(), 0);

Behavioural fixes that change results ​

Each is now pinned by a test that fails without the fix, and each will move an output your program may be asserting on.

CrateWhat changedWhat it means for you
verbora-sentimentThe tokenizer's UAX #29 boundaries split hyphenated lexicon keys, making 2,313 entries unreachable, several of them sign-inverted — "non-approved" scored +1 against a stored polarity of -2. Multi-token span matching now reaches them, and the phrase keys single-token lookup never could.Scores move, sometimes by a sign. Re-baseline any threshold tuned on 0.1.
verbora-stemmersSwedish and Norwegian folded a/o umlauts before consulting stop-word lists spelled with them, so 116 of 428 Swedish entries could never match. In those languages those are distinct letters.Swedish and Norwegian stop-word filtering removes more than it used to.
verbora-stemmersThe German stemmer's character gate was a byte-for-byte copy of the Spanish one: it admitted a/e/i/n/o/u accents and omitted a/o/u umlauts and eszett.German stems change.
verbora-stemmersThe Dutch stop-word list spelled an entry with a trailing space, so the pronoun je was never filtered.One more Dutch stop word is filtered.
verbora-taggerTag::new("*") is refused (Err(LiteralError::Wildcard)), because * printed as the wildcard and reparsed as TagPattern::Any — a rule that rewrote one tag became one that rewrote every tag, across the documented persistence path. Corpus::parse_brown inherits this as CorpusParseError::WildcardTag.A corpus or rule file containing * as a literal tag is now an error instead of silently corrupting the rule set.
verbora-taggerTemplate::instantiate pushed one Condition per position inspected rather than one per distinct condition, so a window template double- (or triple-) counted a corrected token, defeating the trainer's min_score guard.Trained rule sets differ. Retrain.
verbora-transliteratorsVowel lengthening consumed the next scalar without asking whether it began a longer key, so ハロウィン came out harōin and スウェーデン came out sūēden. Six keys collide this way after any of seventy morae.Romanisations change for the affected syllables.
verbora-languagefold_cyrillic folded two blocks while the script router feeds it a third, so 100 uppercase letters went unfolded — including Ґ, one of four Ukrainian-versus-Russian discriminators. The same text in different case gave a different answer.Cyrillic detection is now case-invariant, and uppercase Ukrainian text carries its signal. The Cyrillic model was retrained against the corrected extractor.
verbora-phoneticsFour Beider–Morse Italian rules carried U+FFFD where accented vowels belonged, so real Italian input had the vowel deleted from its encoding.Beider–Morse encodings change for the affected Italian inputs.
verbora-distancejaro(x, x) returned 0.0 for single-unit inputs while jaro_winkler returned 1.0 for the same pair; dice_coefficient("", "") returned NaN. Search could return text absent from the target, because a UTF-16 slice could split a surrogate pair and from_utf16_lossy substituted U+FFFD.Those cases now return the documented answers, and SearchResult::substring() is a borrow of the target, so it cannot be text the target does not contain.
verbora-tfidfThe between-documents invariant was restored in finish, which unwinding skips, so a panicking caller iterator left the counter dirty and the next document reported a present term as Some(0) — documented to mean absent.Corpora built through a panicking tokenizer were wrong; they are not now.

Resource behaviour ​

Not a correctness change, but it will change how your program behaves under load:

  • verbora-spellcheck's deletion generation was cubic in word length. An 800-scalar token cost 4.0 GB of peak RSS; it now costs 32.4 MB. k = 0, documented as a membership test, no longer builds an index at all. DeletionIndex was re-keyed onto a u64 hash for the same reason — cubic to quadratic in word length. The index itself stays quadratic in the longest word: that is the symmetric-delete structure's own price, now documented.
  • verbora-trie's insert_all fed size_hint().0 straight to Vec::reserve, so an iterator that lies — which size_hint explicitly permits — could abort with a capacity overflow or reserve tens of gigabytes. The hint is now clamped.
  • verbora-wordnet's Resident and LazyResident allocated by file metadata with no ceiling. The dictionary file is caller-supplied, so its size was always an input; the three preloading storage modes now share the ceiling Indexed already had. (Storage::Pread preloads nothing and never had the problem.)
  • MaxEntClassifier::restore's JSON parser descended one stack frame per nesting level with no bound, so a model file of 20,000 nested brackets overflowed the stack — which aborts rather than unwinding, so no caller could catch it. It is bounded at 128 levels.

What this page does not tell you ​

No performance figure on this site has been re-measured against 0.2.0. The benchmark campaign for this release has not been run. Every timing published here and on the benchmark pages was produced against 0.1.0-era code and is marked pending; the kernels underneath several of them have been replaced since. Treat every number as historical, and measure your own workload rather than reasoning from a ratio on this site. The resource figures in the section above are peak-RSS measurements recorded with the fix, not throughput.

It also does not cover the places where only the justification moved. Several claims in 0.1's documentation were wrong about code that is unchanged, and several fixtures were re-derived from the publications they should always have come from without any value moving:

  • An empty phonetic key does not imply the input had no recognised letter (Metaphone on "y", Cologne on "h", MatchRatingApproach on any single vowel), and DaitchMokotoff does advance by byte length past a matched pattern, which consumes one character too many for four two-byte keys.
  • A shortest-path tie-break follows relaxation order, not insertion order, and 149 French inflector entries described as load-bearing are reached by no rule.
  • Cologne's 170 expected values were re-derived from Postel (1969) and Match Rating's from NBS Special Publication 500-2 (1977). Where that paper reads its two comparison passes as consecutive, Verbora interleaves them — a difference of 31,044 ratings in 1,129,284 pairs, now stated as a Verbora decision rather than inherited. The implementation did not change; the claim about it did.

If you were relying on one of those claims rather than on the code, your program was already doing something else.

verbora-wordnet 0.3.0 ​

Crates version independently: one moves when its own public API does. Two have moved to 0.3; the other seventeen are at 0.2.

verbora-wordnet 0.2.0 could not read 8.8% of a real WordNet dictionary — 13,606 of 155,467 index entries, including run, cat, light, water, computer and node. PointerSymbol::from_symbol accepted the domain pointers only in their qualified spellings (;c, -c, ;r, -r, ;u, -u) and rejected the bare ; and - that Princeton's index.* files actually write — the class letter belongs to the sense, so it appears only in data.*. index_entry, senses, sense and lookup all failed on any affected lemma.

If you are on 0.2.0 and WordNet lookups were failing for common words, that is why. Move to "0.3".

The fix is breaking in one way:

rust
// 0.2 — compiles
match symbol {
    PointerSymbol::Hypernym => "…",
    // … every other variant …
}

// 0.3 — `PointerSymbol` gained `Domain` (`;`) and `Member` (`-`) and is
// `#[non_exhaustive]`, so a wildcard arm is required.
match symbol {
    PointerSymbol::Hypernym => "…",
    // … every other variant …
    _ => "…",
}

Domain and Member are relations in their own right rather than being folded into one of the three classes: an index entry states that a domain relation exists and deliberately does not state which kind, and choosing one would be a claim the file does not make. They appear only from index entries — a data record that omits the class letter is malformed and is rejected.

verbora-tagger 0.3.0 ​

verbora-tagger 0.2.0 shipped four data files it had no right to redistribute: an English lexicon and rule set that are LGPL-3.0, which is incompatible with this project's MIT licence, and a Dutch lexicon and context-rule set whose terms could not be located at all. All four were removed in 0.3.0. Attribution does not fix a licence mismatch, and the crate is not going to ship data it cannot account for.

What that means for your code: the crate no longer knows any language. It is a tagging engine, and the lexicon is now an input you supply.

rust
// 0.2 — a tagger that arrived knowing English.
use verbora_tagger::{BrillTagger, Language, Lexicon, RuleSet};

let lexicon = Lexicon::bundled(Language::English);
let rules = RuleSet::bundled(Language::English);
let tagger = BrillTagger::new(&lexicon, &rules);
rust
// 0.3 — you bring the dictionary. `Corpus` counts the tag frequencies for you.
use verbora_tagger::{BrillTagger, Corpus, RuleSet, Tag};

let corpus = Corpus::parse_brown(&std::fs::read_to_string("brown.txt")?)?;
let lexicon = corpus.build_lexicon(Tag::new("NN")?)?;
let rules: RuleSet = "NN VB PREV-TAG MD".parse()?;
let tagger = BrillTagger::new(&lexicon, &rules);

Lexicon::new(default_tag) plus Lexicon::insert is the other way in, when the entries are yours to write down rather than to count, and Trainer learns a rule set from the same corpus. Both paths are shown, compiled, on the POS tagger page.

The removals:

0.20.3
Lexicon::bundled(Language)gone — build one with Corpus::build_lexicon, or Lexicon::new plus Lexicon::insert
RuleSet::bundled(Language::English)RuleSet::brill_1992() — the same ten published rules of Brill (1992), Table 1, under a name that carries their citation and without the Language parameter
RuleSet::bundled(Language::Dutch)gone — no replacement; train one with Trainer against your own corpus
brill_paper_rule_strings()gone — RuleSet::brill_1992() is the same table, and RuleSet::to_string() recovers the rule strings verbatim
Languagegone — every method on it named data that no longer ships, and its defaults (NN, NNP) were Penn Treebank tags while the one surviving rule set is Brown
RuleSet::brill_1992 is written in Brown corpus tags. It names AT, PPS, PPO, HVD and NP — not the Penn Treebank DT, PRP, VBD, NNP. Verbora attaches no meaning to a tag beyond string identity, so pairing it with a Penn-tagged lexicon is not an error: the rules simply never fire, the tagger costs a pass per rule, and it returns the initial-state annotation unchanged. If you build a Penn-tagged lexicon, train your own rules against it rather than reaching for this set.

If you were relying on the bundled English tagger and have no corpus of your own, the shape of the work is: obtain an annotated corpus under terms you are happy with, run it through Corpus::parse_brown and Corpus::build_lexicon, and train rules with Trainer. That is a deliberate trade — the crate stopped making a licensing decision on your behalf.

If you are stuck ​

  1. Check the crate's rustdoc first. The crate root is the whole public surface now, so cargo doc --open -p verbora-<crate> shows you everything in one list.
  2. Check the crate's README.md. Every crate has one as of 0.2.0, it is that crate's crates.io landing page, and its examples are compiled as doctests.
  3. Check the feature page for the subsystem — linked from Features — for what the replacement is for, not just what it is called.
  4. Report a documentation error as a bug. If this page sent you somewhere that does not exist, that is worth an issue: the repository.

Next ​

Released under the MIT License.