← Back to case studies

Applied AI Blog: Notes from Building spreezy.ai

Cold Start: Achilles Heel of AI Activation

Every ML feature starts with the same problem: you need user data to train a good model, and you need a good model to get users to generate that data. Waiting for the chicken to lay the egg means shipping delays, or shipping a feature with nothing to learn from — and for a B2C app, time-to-value is everything.

Seriesspreezy.ai · 0→1 Notes
RoleFounder
Published2026
CategoryApplied AI / Cold-Start Problem

The Problem

No data, but a real product decision to make

At spreezy.ai this showed up almost immediately. On day zero there is no user, no history, nothing to learn from. My first instinct was to wait: collect a few months of real behavior, then train. In practice, “it’ll get smarter later” is a hard sell when the first version lacks the smarts.

This isn’t a niche situation. Every new feature inside an existing product, every new market a company enters, every startup before its debut faces the identical problem: a model that needs signal to be useful, sitting in front of zero signal.

It matters for a specific reason. The costs of a bad cold-start experience compound quickly. A recommendation engine that gets it wrong in week one doesn’t just embarrass itself once — it trains itself on its own bad guesses, or worse, drives away the very users whose behavior would have corrected it. The system that most needs real data is the one least likely to get any, because a broken early experience suppresses the engagement that would fix it.

The Landscape

Five ways we are tackling this at spreezy.ai

  • Rule-based heuristics. Sort by popularity, recency, or category defaults instead of a learned model. It’s not machine learning at all, and that’s by design. It’s honest about what it is, costs nothing, and gives you a baseline to beat later. It can’t personalize, but it’s the correct first move almost everywhere.
  • Transfer learning. Take a model trained in a related, data-rich domain and fine-tune it on whatever small sample you do have. A model pretrained on a mature catalog can be adapted to a new one with a fraction of the data, provided the domains overlap. This is the standard move when a similar dataset exists somewhere, even if yours doesn’t.
  • Synthetic data generation. LLMs are your friends. Use an LLM, a simulator, or hand-written generative rules to manufacture plausible examples that stand in for real interactions. Synthetic data approximates reality, it doesn’t reproduce it, and a model trained purely on synthetic examples will carry its generator’s blind spots. Airbnb’s search team did this explicitly for a natural-language search launch: no real queries and no relevance labels existed yet, so they combined LLM-generated queries with seed examples from user research to train and evaluate ranking models before real traffic arrived, with a deliberate plan to hand off to real data once it showed up.
  • Simulated user interactions. One step further, simulate an agent acting like a user, generating a full behavioral trace instead of isolated labels. This is the approach behind recent work simulating user interactions for “cold” items in recommendation systems that have content but no behavioral history — using an LLM to stand in for the missing interaction data rather than deriving synthetic embeddings from content features alone, which several papers note still leaves a real gap between content-derived and behavior-derived signal.
  • Cross-tenant or aggregated data. If you’re a platform serving many customers, pool data across them rather than treating each as a cold start in isolation. This is standard in multi-tenant SaaS: a model predicting rare events (workplace injuries, fraud, churn) is trained on the aggregate because any single customer’s data is too sparse alone, then specialized per-tenant as their own data accumulates.
Machine learning cold-start pathway: a five-stage data pipeline from synthetic data, internal dog-fooding, closed beta, and agent simulations through to real usage data, paired with a model-maturity progression from heuristics and rules to scaled ML to high maturity and continuous learning
The cold-start pathway in two layers: a data pipeline moving from synthetic data to real usage data, and the model maturity it unlocks along the way, from simple heuristics to continuous learning.

The Tradeoffs

What each approach costs you

MethodCostRealismFailure mode
Rule-based heuristicsNegligibleNone (not a model)Never personalizes
Transfer learningFine-tuning computeHigh, if domains matchNegative transfer if domains diverge
Synthetic dataGeneration cost (LLM calls)ApproximateBakes in generator’s blind spots
Simulated interactionsHigher when you simulate at scaleCloser to real behaviorContent-behavior gap persists
Cross-tenant aggregationData-sharing / privacy overheadHigh, if base is relatedRequires enough related tenants to pool

The pattern here mirrors any build-vs-buy tradeoff: the more a method tries to approximate real behavior, the more it costs to build and the more carefully you must audit it for the assumptions it’s quietly encoding. Rules cost nothing and encode nothing. Simulated interactions cost the most and encode the most — including, silently, whatever biases went into building the simulator.

The Approach

Sequence the methods, and let real data retire each stage

Don’t choose a single method — sequence them, and let real data retire each stage as it arrives. This is the approach we are taking at spreezy.ai:

  • Stage one, always: ship the rule-based baseline. It’s not embarrassing to launch on popularity sort or recency; it’s dishonest to dress it up as personalization, and it’s the fastest way to start collecting the real signal that eventually replaces it.
  • Stage two: while internal testing and pilot users are a good way to get real data early, use one of the techniques above to scale for your training data needs. Synthetic data is imperfect, but it’s still better than starting from nothing.
  • Stage three: generate simulated interactions, but build a cutover into the plan from day one at a fixed point (a data volume, a time window, a shadow metric) at which real interaction data takes over training, rather than letting synthetic or simulated data become a permanent crutch.

Closing Thought

There is no one right answer, only the right sequence

Cold start poses a real challenge when you are bootstrapping a new product under real budget constraints, and time-to-value is crucial for driving real user adoption. There is no single definitive approach for all domains — but sequencing cheap methods first and letting real usage retire them, rather than betting everything on one technique, is what has kept spreezy.ai moving.

Concepts & methods

Cold-Start Problem Rule-Based Heuristics Transfer Learning Synthetic Data Generation LLM-Simulated User Interactions Cross-Tenant / Multi-Tenant Aggregation Recommendation Systems Staged Model Activation

Sources