Two data scientists at their workstations reviewing abstract charts while working on a recommendation system

Recommendation System Design: A Practical Guide from Objectives to Evaluation

A recommendation system does more than predict what a person might click. It decides which products, courses, documents, services or media receive attention—and in what order. That makes recommendation system design an operational and governance problem as much as a modelling problem.

The most reliable starting point is not an algorithm. It is a precisely defined decision, a useful outcome and a measurement plan that prevents one convenient metric from distorting the experience. This practical guide takes a system from objectives through candidate retrieval, scoring, re-ranking, evaluation and production monitoring.

What a recommendation system actually does

A recommender selects a small, ordered set of items for a user or context from a much larger catalogue. Inputs can include explicit preferences, previous interactions, item attributes, session context, availability and organisational rules. Outputs may appear as “next best action”, “similar items”, a personalised home page or a ranked search result.

Google’s official recommendation-systems overview describes a common three-stage production architecture: candidate generation, scoring and re-ranking. Candidate generation rapidly reduces a large corpus to a manageable subset. A more precise model scores that subset. Re-ranking then applies requirements such as freshness, diversity and exclusions.

That separation is useful because each stage has a different job. Retrieval must be fast and inclusive enough not to miss strong candidates. Scoring estimates value more carefully. Re-ranking protects the final experience and business rules. Trying to make one model perform all three jobs can be expensive, opaque and difficult to control.

Start with the decision, not the click

“Increase engagement” is not a sufficient objective. It does not say which behaviour is valuable, for whom, over what period or at what cost. A learning platform might value course completion and skill progression, while a retailer might value suitable purchases with low return rates. A public-service portal may prioritise successful task completion rather than time spent.

A good objective has four parts:

  • User outcome: what becomes easier, safer or more relevant?
  • Organisational outcome: what sustainable value should improve?
  • Guardrails: which harms, exclusions or quality losses are unacceptable?
  • Time horizon: is success measured in the session, after a transaction or over months?

Objective choice changes ranking behaviour. Google’s scoring guidance warns that optimising clicks can favour clickbait, while optimising watch time can favour very long content. The lesson generalises: the model will exploit the proxy it receives, so pair every optimisation metric with outcome and guardrail measures.

The EPW RANKED design framework

RANKED is a six-part framework for moving from a vague personalisation idea to a governable production system.

R — Result and decision

State where recommendations appear, who sees them, which action they support and what success means. Define the eligible catalogue and the moment at which the decision is made. Write the primary metric, secondary outcomes and guardrails before building a model.

A — Audience, catalogue and availability

Profile the users, items and contexts the system must serve. Examine catalogue size, turnover, language, geography, stock or capacity constraints and the proportion of new users and items. Availability rules belong upstream: recommending an unavailable product or unsuitable course is not a ranking error alone; it is a system-design failure.

N — Notice signals and negatives

List the events that might represent interest: views, saves, purchases, completions, ratings or repeat use. Then test what each signal really means. A long dwell time may show interest, confusion or an unattended screen. Non-interaction is not automatically dislike because the user may never have seen the item.

Log exposure as well as response. Without knowing which items were displayed, training data confuses “not chosen” with “not offered”. Define negative signals carefully—returns, hides, complaints and rapid abandonment can be more informative than the absence of a click.

K — Keep candidates, scores and constraints separate

Choose retrieval methods that fit the data. Content-based filtering uses item and user features, making it useful for new items and explainability. Collaborative filtering learns from patterns of user–item interaction and can uncover latent relationships, but it struggles when interaction data is sparse. Hybrid systems combine sources such as popularity, rules, content similarity and learned embeddings.

Google’s candidate-generation guidance notes that content-based and collaborative methods can represent queries and items in an embedding space, then retrieve candidates using measures such as cosine similarity, dot product or Euclidean distance. Those measures are not interchangeable: for example, dot product can favour high-norm, often frequent items.

After retrieval, a ranking model can use richer features. Re-ranking should then remove ineligible items and enforce diversity, freshness, fairness, frequency limits or contractual rules. Keep these policies visible rather than hoping the model learns them indirectly.

E — Evaluate offline and online

Offline evaluation helps compare models safely. Use a time-based split where appropriate so the test set represents future behaviour rather than a random mixture of past and future. Metrics must match the interface: precision@k measures how many displayed items are relevant; recall@k measures how many relevant items are recovered; ranking metrics such as NDCG reward putting useful items near the top. Coverage, novelty and diversity reveal whether relevance is concentrated on a narrow set.

Offline performance is necessary but not sufficient. Historical data was produced by the previous policy, position effects and exposure choices. Online experiments should therefore measure the stated user and organisational outcome, guardrails, subgroup effects and longer-term behaviour. Predefine stopping rules and minimum practical effect, not only statistical significance.

D — Deploy, detect and decide again

Monitor candidate availability, feature freshness, retrieval recall, score distributions, latency, fallback rates and business outcomes. Track exposure concentration and subgroup performance. Define ownership for incidents, rollback and model retirement. Feedback changes the data the next model learns from, so periodically reconsider the objective and exploration policy rather than treating deployment as the end.

EPW RANKED framework for recommendation-system design and continuous improvement
RANKED keeps objectives, data, architecture, evaluation and production feedback connected.

Architecture choices by problem condition

Condition Useful starting method Main strength Main risk
Limited interaction data Rules, popularity and content-based retrieval Works before a dense behaviour history exists Can be generic or over-similar
Rich user–item interactions Collaborative filtering or matrix factorisation Finds latent taste patterns Cold start and popularity bias
Large, fast-changing catalogue Multiple candidate generators plus learned ranking Balances scale and precision Operational complexity and retrieval blind spots
Strict eligibility constraints Rule filter before retrieval and policy re-ranking Prevents unsuitable results Rules can become fragmented or stale
Several competing outcomes Multi-objective ranking with explicit guardrails Makes trade-offs testable Weights can hide policy choices
High-stakes decision support Conservative retrieval, explanations and human review Supports accountability Personalisation may be inappropriate for some decisions

A simple baseline should survive long enough to be beaten convincingly. Popular items within an eligible category, recent successful choices or a rule-based shortlist often create a strong reference. Compare any complex model against that baseline on outcome, robustness, cost and explainability.

Worked example: recommending professional courses

Imagine a catalogue that recommends professional courses. The initial request is to “show people courses they will click”. RANKED exposes the weakness of that objective. A click on an advanced course that the learner cannot use is not success. The team reframes the result as helping a visitor identify a relevant, feasible next course, measured by qualified enquiries and later enrolment, with guardrails for prerequisite fit, catalogue exposure and complaint rate.

The audience and catalogue analysis records role, stated goal, experience, preferred location and schedule only when appropriately collected. Course records include topic, level, prerequisites, delivery locations and dates. Eligibility filters remove unavailable sessions and courses whose mandatory prerequisites are not met.

For retrieval, the system combines three candidate sources: content similarity to the visitor’s stated goal, popular courses within the relevant professional family, and courses related to recent voluntary browsing. A ranker estimates usefulness using those features, while re-ranking limits near-duplicates and adds some diversity across skills or dates.

Offline evaluation uses time-ordered data and reports precision@5, catalogue coverage and results for new versus returning visitors. The online test measures qualified course enquiries, not clicks alone. It also monitors backtracking, unsuitable-course reports and concentration of exposure. New visitors receive a robust content-and-popularity fallback instead of a fabricated personal profile.

This design is modest, but it is testable. It separates eligibility from prediction, includes a cold-start path and measures whether the recommendation helped someone make a better choice.

Worked recommendation-system design from user goal and eligibility to ranking and outcome measurement
A practical recommender separates eligibility, candidate sources, scoring, re-ranking and outcome evaluation.

How to evaluate recommendation quality

Use a scorecard rather than a single headline number:

  • Relevance: precision@k, recall@k, NDCG or task-specific success.
  • Reach: user coverage, item coverage and performance for sparse-history cases.
  • Experience: diversity, novelty, repetition, latency and explicit dissatisfaction.
  • Outcome: completion, retained use, suitable purchase, qualified enquiry or another durable result.
  • Guardrails: returns, complaints, unsafe or ineligible exposure, subgroup disparities and excessive concentration.
  • Operations: feature freshness, empty candidate sets, fallback rate, service errors and cost per recommendation.

Break metrics down by meaningful conditions: new and established users, new and established items, device type, region or catalogue segment. Aggregate gains can hide a system that improves for data-rich users while degrading for everyone else.

Cold start, feedback loops and responsible control

Cold start is not one problem. New users lack interaction history; new items lack exposure; a new platform lacks both. Use declared preferences, item metadata, contextual popularity and brief onboarding where proportionate. Exploration can gather evidence for new items, but it should be bounded by suitability and risk.

Recommendation creates feedback. Items shown near the top gain interactions, which can make them more likely to be shown again. Monitor exposure as well as response, compare popularity-normalised measures and reserve controlled opportunities for suitable new or less-exposed items. Do not force diversity where it conflicts with safety or relevance; treat it as an explicit policy trade-off.

Privacy also shapes design. Collect only signals needed for the stated purpose, explain personalisation where appropriate, respect user controls and set retention limits. Sensitive attributes should not be inferred casually. In higher-impact contexts, a recommendation may require explanations, appeal routes and human oversight—or personalisation may be the wrong approach entirely.

Production-readiness checklist

  1. Write the decision, primary outcome, time horizon and guardrail metrics.
  2. Define catalogue eligibility, availability and exclusion rules.
  3. Document each behavioural signal and plausible alternative interpretation.
  4. Log exposure, position, response and outcome with appropriate privacy controls.
  5. Build and retain a simple baseline.
  6. Measure candidate recall separately from ranking quality.
  7. Create explicit fallbacks for new users, new items and service failure.
  8. Use time-aware offline evaluation and prevent feature leakage.
  9. Test online with pre-agreed success, guardrail and stopping criteria.
  10. Monitor subgroup performance, catalogue concentration, latency and cost.
  11. Version features, models and policies; rehearse rollback.
  12. Assign owners for quality review, incidents and objective changes.

Developing recommendation-system design skills

Recommendation systems require a blend of problem framing, data interpretation, retrieval, ranking, experimentation and production governance. EPW’s Recommendation Systems Design and Optimization course covers collaborative and content-based methods, matrix factorisation, hybrid and context-aware approaches, evaluation, deployment, scalability and feedback-loop optimisation.

The discipline also connects to wider organisational use of AI. EPW’s guide to the benefits of AI and machine learning for organisations explains how to link model outputs to decisions, controls and measurable outcomes.

Conclusion

Good recommendation system design aligns the ranked list with a useful decision. Start with the result and guardrails; understand the audience, catalogue and signals; keep retrieval, scoring and policy controls distinct; evaluate online as well as offline; and monitor the feedback the system creates.

The RANKED framework makes those dependencies visible. It also makes a crucial point: a better prediction metric is not automatically a better recommendation. The system succeeds when it helps people make better choices while remaining relevant, resilient and accountable in production.