← Back to case studies

Applied AI Blog: Notes from Building spreezy.ai

Embed or Not to Embed

Any product that takes free-text input eventually has to answer the same question: is this new thing the same as something already there? The instinct is to reach for embeddings and a vector database, since “semantic search” sounds like the sophisticated answer to anything involving text. It usually isn’t: duplicate detection is three problems wearing one name, and each needs a different tool.

Seriesspreezy.ai · 0→1 Notes
RoleFounder
Published2026
CategoryApplied AI / Duplicate Detection

The Problem

Same thing, worded differently

A shopping list app has to decide if “Whole Milk” is a new item or the same as “Milk.” A CRM has to decide if “Jon Smith” and “Jonathan Smith” are one contact. The inputs are messy and inconsistent, and near-duplicates are sometimes intentional.

But duplicate detection is really three problems wearing one name: literal repeats with formatting noise, typos and near-misses, and conceptually identical items worded completely differently. Each needs a different tool; the wrong one either misses obvious duplicates or silently merges things that were never the same.

A concrete case makes this vivid: a list already has “Milk.” Adding “Whole Milk” should probably prompt a check, since it’s a variant of the same product. Adding “Almond Milk” shouldn’t: it’s a genuinely different product that happens to share a word. Getting that right cheaply forces you to understand what each method does.

The Landscape

Seven ways to catch a duplicate

  • Exact / normalized matching (lowercase, strip punctuation, compare directly) is essentially free, has zero false positives, and should be in every pipeline. It catches nothing spelled differently, which is a large share of duplicates but far from all.
  • Wildcard / substring matching (LIKE, regex) extends this to containment, useful for search UIs but not similarity: it can’t tell “banana” is close to “bananas” but not to “bandana.”
  • Trigram similarity breaks strings into overlapping 3-character windows and scores overlap. It’s the workhorse for typos and formatting differences, and it’s cheap: implementations like PostgreSQL’s pg_trgm are indexable (GIN/GiST), so it stays fast at millions of rows. Its blind spot is meaning: it has no idea “soda” and “pop” are the same thing, since there’s no character overlap.
  • Levenshtein distance counts the minimum edits to turn one string into another. It’s more precise than trigrams for pure typo detection and easier to tune (“off by 2 characters”), but there’s no efficient general-purpose index for “everything within edit distance N.” In practice it re-ranks a shortlist a cheaper method already narrowed, rather than scanning a whole dataset.
  • Phonetic matching (Soundex, Metaphone) encodes words by sound rather than spelling, catching things like “Smith” vs. “Smyth.” It’s narrow-purpose, mostly names.
  • Vector embeddings turn text into vectors that capture meaning, so “soda”/“pop” or “OJ”/“orange juice” land close together with zero character overlap. This is the only method here that catches genuinely conceptual duplicates. The cost is real: generating an embedding means a model call, and comparing at scale means either brute force or a dedicated index.
  • Vector databases (Pinecone, Weaviate, pgvector) make embedding comparison fast once you’re searching millions of vectors, via approximate nearest-neighbor search. They matter at scale and are overkill below it: comparing one vector against a few hundred cached ones in memory is instant. A related technique, locality-sensitive hashing (MinHash, SimHash), approximates similarity over enormous corpora in web-scale document dedup without a full embedding pipeline.
Where each method sits, positioned by what it can catch (literal spelling → conceptual meaning) against how much infrastructure and cost it takes to run. Hover or tab through a point for details.

The Tradeoffs

Cost tracks meaning, not quality

MethodCostInfra dependencyComplexityCatches
Exact matchNegligibleNoneVery lowLiteral repeats
WildcardNegligibleDatabase onlyLowContainment only
TrigramNegligible–lowDB extension at scaleLow–mediumTypos, near-spelling
LevenshteinLow (on shortlists)NoneLowPrecise typo distance
PhoneticNegligibleNoneLowSound-alike names
EmbeddingsPer-call costEmbedding modelMediumConceptual duplicates
Vector DBHosting costDedicated serviceMedium–highFast semantic search at scale

The pattern: cost and infrastructure scale with how much meaning a method understands, not with how good it is at its job. String methods are cheap because they do something narrow: compare characters. Embeddings are more capable but add a real dependency and cost. A managed vector database adds a second, larger dependency on top, justified only once brute-force search is actually slow, typically tens of thousands of vectors per request, not hundreds.

The mistake runs both directions: a vector database because “semantic search” sounds correct when an instant, free in-memory loop would do; or string methods alone when the product genuinely needs same-meaning-different-words matches, which no amount of trigram tuning will solve.

The Recommendation

Layer the methods, add meaning last

Layer the methods rather than picking one, and let the dataset’s size and nature decide how many layers you need.

  • Layer one, always: normalization and exact matching: free, and always the first check.
  • Layer two, for most real-world duplicates: trigram similarity with Levenshtein re-ranking, for typos and near-identical spelling; this alone handles the large majority of real-world duplicates at essentially no cost.
  • Layer three, only when meaning genuinely matters: vector embeddings, and even then, embed only the items that survive the first two filters unmatched, rather than embedding everything by default.

Skip the vector database until the numbers force it. Within a small, bounded set (a user’s own list, one account’s records), brute-force cosine similarity over cached vectors in memory beats standing up Pinecone or Weaviate. Reach for pgvector once you need a shared catalog, and a dedicated managed vector database only once query volume makes fast approximate search a real requirement, not a precaution.

Closing Thought

Some judgments aren’t a similarity score

For cases where the same-vs-different judgment follows a known pattern (like fat-content variants of milk being the same product family while plant-based alternatives are a different one), don’t expect any similarity score, string-based or semantic, to encode that judgment reliably. A small, explicit rule set for known patterns will outperform a generic model on the exact cases you can anticipate, leaving the embedding layer to handle the long tail you can’t.

Concepts & methods

Duplicate Detection Exact / Normalized Matching Trigram Similarity (pg_trgm) Levenshtein Distance Phonetic Matching Vector Embeddings Approximate Nearest Neighbor Search Locality-sensitive Hashing pgvector Layered Matching Pipelines Cost-aware Architecture