Applied AI Blog: Notes from Building spreezy.ai
Every ML feature starts with the same problem: you need user data to train a good model, and you need a good model to get users to generate that data. Waiting for the chicken to lay the egg means shipping delays, or shipping a feature with nothing to learn from — and for a B2C app, time-to-value is everything.
The Problem
At spreezy.ai this showed up almost immediately. On day zero there is no user, no history, nothing to learn from. My first instinct was to wait: collect a few months of real behavior, then train. In practice, “it’ll get smarter later” is a hard sell when the first version lacks the smarts.
This isn’t a niche situation. Every new feature inside an existing product, every new market a company enters, every startup before its debut faces the identical problem: a model that needs signal to be useful, sitting in front of zero signal.
It matters for a specific reason. The costs of a bad cold-start experience compound quickly. A recommendation engine that gets it wrong in week one doesn’t just embarrass itself once — it trains itself on its own bad guesses, or worse, drives away the very users whose behavior would have corrected it. The system that most needs real data is the one least likely to get any, because a broken early experience suppresses the engagement that would fix it.
The Landscape
The Tradeoffs
| Method | Cost | Realism | Failure mode |
|---|---|---|---|
| Rule-based heuristics | Negligible | None (not a model) | Never personalizes |
| Transfer learning | Fine-tuning compute | High, if domains match | Negative transfer if domains diverge |
| Synthetic data | Generation cost (LLM calls) | Approximate | Bakes in generator’s blind spots |
| Simulated interactions | Higher when you simulate at scale | Closer to real behavior | Content-behavior gap persists |
| Cross-tenant aggregation | Data-sharing / privacy overhead | High, if base is related | Requires enough related tenants to pool |
The pattern here mirrors any build-vs-buy tradeoff: the more a method tries to approximate real behavior, the more it costs to build and the more carefully you must audit it for the assumptions it’s quietly encoding. Rules cost nothing and encode nothing. Simulated interactions cost the most and encode the most — including, silently, whatever biases went into building the simulator.
The Approach
Don’t choose a single method — sequence them, and let real data retire each stage as it arrives. This is the approach we are taking at spreezy.ai:
Closing Thought
Cold start poses a real challenge when you are bootstrapping a new product under real budget constraints, and time-to-value is crucial for driving real user adoption. There is no single definitive approach for all domains — but sequencing cheap methods first and letting real usage retire them, rather than betting everything on one technique, is what has kept spreezy.ai moving.
Concepts & methods
Sources