Applied AI Blog: Notes from Building spreezy.ai
Any product that takes free-text input eventually has to answer the same question: is this new thing the same as something already there? The instinct is to reach for embeddings and a vector database, since “semantic search” sounds like the sophisticated answer to anything involving text. It usually isn’t: duplicate detection is three problems wearing one name, and each needs a different tool.
The Problem
A shopping list app has to decide if “Whole Milk” is a new item or the same as “Milk.” A CRM has to decide if “Jon Smith” and “Jonathan Smith” are one contact. The inputs are messy and inconsistent, and near-duplicates are sometimes intentional.
But duplicate detection is really three problems wearing one name: literal repeats with formatting noise, typos and near-misses, and conceptually identical items worded completely differently. Each needs a different tool; the wrong one either misses obvious duplicates or silently merges things that were never the same.
A concrete case makes this vivid: a list already has “Milk.” Adding “Whole Milk” should probably prompt a check, since it’s a variant of the same product. Adding “Almond Milk” shouldn’t: it’s a genuinely different product that happens to share a word. Getting that right cheaply forces you to understand what each method does.
The Landscape
LIKE, regex) extends this to containment, useful for search UIs but not similarity: it can’t tell “banana” is close to “bananas” but not to “bandana.”pg_trgm are indexable (GIN/GiST), so it stays fast at millions of rows. Its blind spot is meaning: it has no idea “soda” and “pop” are the same thing, since there’s no character overlap.pgvector) make embedding comparison fast once you’re searching millions of vectors, via approximate nearest-neighbor search. They matter at scale and are overkill below it: comparing one vector against a few hundred cached ones in memory is instant. A related technique, locality-sensitive hashing (MinHash, SimHash), approximates similarity over enormous corpora in web-scale document dedup without a full embedding pipeline.The Tradeoffs
| Method | Cost | Infra dependency | Complexity | Catches |
|---|---|---|---|---|
| Exact match | Negligible | None | Very low | Literal repeats |
| Wildcard | Negligible | Database only | Low | Containment only |
| Trigram | Negligible–low | DB extension at scale | Low–medium | Typos, near-spelling |
| Levenshtein | Low (on shortlists) | None | Low | Precise typo distance |
| Phonetic | Negligible | None | Low | Sound-alike names |
| Embeddings | Per-call cost | Embedding model | Medium | Conceptual duplicates |
| Vector DB | Hosting cost | Dedicated service | Medium–high | Fast semantic search at scale |
The pattern: cost and infrastructure scale with how much meaning a method understands, not with how good it is at its job. String methods are cheap because they do something narrow: compare characters. Embeddings are more capable but add a real dependency and cost. A managed vector database adds a second, larger dependency on top, justified only once brute-force search is actually slow, typically tens of thousands of vectors per request, not hundreds.
The mistake runs both directions: a vector database because “semantic search” sounds correct when an instant, free in-memory loop would do; or string methods alone when the product genuinely needs same-meaning-different-words matches, which no amount of trigram tuning will solve.
The Recommendation
Layer the methods rather than picking one, and let the dataset’s size and nature decide how many layers you need.
Skip the vector database until the numbers force it. Within a small, bounded set (a user’s own list, one account’s records), brute-force cosine similarity over cached vectors in memory beats standing up Pinecone or Weaviate. Reach for pgvector once you need a shared catalog, and a dedicated managed vector database only once query volume makes fast approximate search a real requirement, not a precaution.
Closing Thought
For cases where the same-vs-different judgment follows a known pattern (like fat-content variants of milk being the same product family while plant-based alternatives are a different one), don’t expect any similarity score, string-based or semantic, to encode that judgment reliably. A small, explicit rule set for known patterns will outperform a generic model on the exact cases you can anticipate, leaving the embedding layer to handle the long tail you can’t.
Concepts & methods