Skip to content

Iterator vs reusable buffer ​

These are the two shapes people most often assume are alternatives. They are not. They optimise different things, and each is wrong in the other's situation.

The two problems ​

An iterator removes the container. There is no Vec, so there is nothing to allocate, nothing to fill, and nothing to keep in memory. You get one item at a time and you can stop whenever you like.

_into reuses the container. There is still a Vec, still filled completely, still holding every item at once — but its allocation is paid for once and then borrowed by every subsequent call.

text
tokens()                          tokenize_borrowed_into()

input                             allocate buffer once
  │                                        │
  ├─ token ─▶ consumer            document 1 ─▶ fill ─▶ consume ─▶ clear
  ├─ token ─▶ consumer            document 2 ─▶ fill ─▶ consume ─▶ clear
  ├─ token ─▶ consumer            document 3 ─▶ fill ─▶ consume ─▶ clear
  └─ …                            document 4 ─▶ fill ─▶ consume ─▶ clear

no container exists               one container, reused
peak memory: one token            peak memory: one document's tokens
can stop early                    always produces everything
single pass                       result is re-readable, indexable, sortable

When the iterator wins ​

You consume once, in order.

rust
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};

let letters: usize = WordTokenizer
    .tokens("counting letters in every word")
    .map(str::len)
    .sum();

assert_eq!(letters, 26);

Materialising here would allocate a Vec that exists for one traversal.

You might not need all of it.

rust
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};

// Scanning stops at the first hit; the rest of the string is never split.
let found = WordTokenizer.tokens("alpha beta gamma delta").any(|w| w == "beta");

assert!(found);

With a materialised result you have already done all the work before you can look at any of it. On a 9.7 kB document that is the difference between splitting two tokens and splitting fifteen hundred.

The input is big and you do not want it all in memory at once. Peak memory for a lazy pass is one token. For a materialised pass it is every token in the document, simultaneously.

You are feeding another API that takes IntoIterator.

rust
use verbora_core::{StopWordLanguage, StopWords};
use verbora_phonetics::{SoundEx, phoneticize_tokens};
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};

let soundex = SoundEx::new();
let stops = StopWords::for_language(StopWordLanguage::En);

let keys = phoneticize_tokens(WordTokenizer.tokens("the quick fox"), &stops, |t| {
    soundex.process(t)
});

assert_eq!(keys, ["Q200", "F200"]);

No Vec<String> is built between the tokenizer and the encoder.

When _into wins ​

You need the whole result, and you need it again and again.

rust
use verbora_tokenizers::{BorrowingTokenizer, WordTokenizer};

fn index(corpus: &[&str]) -> usize {
    let mut buf: Vec<&str> = Vec::new();
    let mut hits = 0;

    for document in corpus {
        buf.clear();
        WordTokenizer.tokenize_borrowed_into(document, &mut buf);

        // Two passes over the same tokens. An iterator would have to re-scan.
        let longest = buf.iter().map(|w| w.len()).max().unwrap_or(0);
        hits += buf.iter().filter(|w| w.len() == longest).count();
    }

    hits
}

assert_eq!(index(&["a bb ccc", "dddd ee"]), 2);

Two passes over an iterator means tokenizing twice. Two passes over a buffer is free.

You need random access, a length, or sorting. All of these want a slice.

You are calling the operation an enormous number of times with output of similar size each time. This is the case the API was added for: the buffer reaches its high-water mark within the first few documents and then never calls the allocator again.

When neither wins ​

Calling something once. Use tokenize_borrowed(). Both of the shapes on this page cost you something in exchange for a saving you will not measure.

The decision ​

Do you need every item at once?Then
No — one at a time, and you might stop earlytokens() — the iterator wins twice
No — one at a time, always all of themtokens() — still no container
Yes, oncetokenize_borrowed()
Yes, repeatedly in a looptokenize_borrowed_into() with one buffer

What Verbora actually offers ​

Not every subsystem has both. Reading a name is not enough — check the page.

SubsystemLazy_into
Tokenizerstokens() on every tokenizertokenize_into, tokenize_borrowed_into
N-gramsngrams(), Padded::ngrams(), char_ngrams() — all lazy already— (nothing is allocated to write into)
Trieiter_keys_with_prefix(), keys(), iter_prefix_matches()— (for_each_key_with_prefix is the allocation-free enumeration)
Inflectors—pluralize_into, singularize_into, OrdinalInflector::nth_into, CaseMode::apply_into
Core traits—stem_into (clears first)
Distance——
Phonetics—SoundEx::process_into, Metaphone::process_into — those two encoders only
Transliterators—transliterate_ja_into
TaggerBrillTagger::tag_streamtag_into, annotate_into
Normalizers—— (they return Cow, which is the analogous saving)
_into is not the same as allocation-free. It removes the container allocation. Whether the elements allocate depends on the API: tokenizers put borrowed &str in the buffer, so nothing else allocates — but NounInflector::pluralize_into still allocates inside the matching rule, so it saves one String per call rather than all of them. The feature pages state this per API.

Released under the MIT License.