essays
2026-09-02· 8 minPull UpEngineeringRecommendersData

How Pull Up Learns What You Like

The data sources behind a small music recommender, what each one is for, and the arithmetic that turns taps into a taste.

Ayush Upneja
AI product engineer

Pull Up turns three catalogs, eight audio measurements, and a trail of user actions into recommendations for roughly 1,300 New York music-event cards. The important design choice is that no source gets to answer a question it was not built to answer.

A genre roster can tell us who matters inside drum and bass. It cannot tell us who is playing Brooklyn next Friday. An event listing can tell us who is booked. It cannot safely define an artist’s genre from one support slot. A Deezer result can verify identity, artwork, and audience size. It cannot establish local relevance by itself.

The algorithm is mostly the machinery for keeping those facts separate, then combining them only after their individual gates pass.

Three sources, one crate: an artist enters the bank only when the genre roster, the event index, and Deezer all agree.

Three sources, one crate: an artist enters the bank only when the genre roster, the event index, and Deezer all agree.

Three sources, three questions#

The artist bank begins with three fact sources.

The first is a scrape of noteverynoise.com, a community reconstruction of Every Noise at Once. Our snapshot contains 777 genre pages and 43,186 artists. Within each page, artists are ordered by Spotify followers.

We use that ordering to answer: who matters within this genre?

That qualification matters. A globally famous artist should not automatically lead every genre lane, and a locally recognized DJ should not disappear because their global follower count is modest. The ranked pages provide genre-relative ordering, not a claim that any artist currently has a New York show.

The second source is Pull Up’s New York event index. Its job is to answer: what is actually playing?

The index currently holds more than 1,300 cards. Six deterministic adapter paths enumerate structured sources first, including venue calendars, WordPress feeds, RSS, ICS, Ticketmaster, and DICE. An LLM-powered search lane fills gaps that structured feeds do not expose.

That LLM lane does not publish prose as fact. Its output enters the same normalization pipeline as every other row:

validate fields and date
→ normalize artist, venue, links, and genres
→ reject hallucinated or unsupported rows
→ create occurrences
→ merge into stable cards

An empty or invalid research batch cannot replace the last good one. If a future event still has no defensible genre after enrichment, it stays out of public recommendation surfaces.

The third source is Deezer. We use it to answer: does this artist identity resolve, and can we attach verified artwork and fan count?

Name verification is exact after normalization. That rule exists because fuzzy matching can quietly transfer another artist’s audience to the wrong name. A candidate without a valid exact-name result does not inherit the nearest search result’s millions of fans. For roster-derived entries, an exact match with artwork is required before the artist enters the generated bank.

Each source therefore contributes a different column:

SourceFact it owns
noteverynoiseGenre membership and follower-ranked order
NYC event indexCurrent, bookable local supply
DeezerVerified identity, artwork, and fan count

An artist crosses into the generated banks only when the applicable admission gates agree. Ranked genre evidence supplies the lane. Event evidence supplies local reality. Deezer supplies verified identity and display metadata. Missing evidence produces a dropped or limited candidate, not an invented fact.

Audio is measured locally#

Text metadata describes where an artist sits culturally. Audio describes what the available recording sounds like.

For artists with a usable 30-second preview, Pull Up downloads the preview and computes eight DSP dimensions locally with librosa:

  • tempo
  • energy
  • brightness
  • spectral rolloff
  • percussiveness
  • spectral contrast
  • dance
  • tonal character

No model listens to the track and writes a genre description. The extractor returns numbers.

Coverage is only about half of eligible artists, because many artists do not have a preview that passes resolution and quality checks. Missing audio remains missing. The vector builder renormalizes around the absent block instead of filling it with averages.

The other constraint is sampling. We currently analyze one preview per artist, so n = 1. A 30-second single preview is evidence about one recording, not a complete account of an artist’s catalog. Audio is consequently an additive signal, not the authority over genre or identity.

Two vector spaces do different jobs#

Pull Up has two related taste representations because asking good questions and ranking future shows are different problems.

The quiz brain and the feed brain, joined by a deliberate handoff.

The quiz brain and the feed brain, joined by a deliberate handoff.

The adaptive quiz uses a 28-dimensional posterior: 24 genre buckets plus four room-preference axes. Each answer updates a Gaussian mean and covariance with assumed-density filtering, or ADF. The mean represents the current estimate. The covariance represents what the quiz still does not know.

The question selector uses BALD, an information-gain criterion. It prefers a comparison expected to reduce uncertainty, subject to coverage and redundancy constraints. Entropy provides the stopping signal.

The quiz asks where it is least sure, and stops when the cloud stops shrinking.

The quiz asks where it is least sure, and stops when the cloud stops shrinking.

The arithmetic is useful here. The quiz can reason over:

24 genre buckets + 4 axes = 28 dimensions.

The persistent recommendation vector is wider:

24 buckets + 102 microgenres + 8 audio dimensions + 8 traits = 142 dimensions.

This vector must keep learning after onboarding. It represents declared taste, artist and venue affinities, microgenre evidence, audio fit, and behavioral history.

Not everything in the quiz posterior crosses into it today. Ranked genre selections seed bucket weights. Resolved named artists can contribute artist vectors. The quiz stores its posterior and axis means, but posterior covariance and every inferred bucket magnitude do not become feed weights wholesale. The persistent model is reconstructed from the committed genre and artist facts, then updated by behavior.

The current v6 deck makes that boundary safer. Every left or right swipe first posts through the production taste endpoint. A right swipe contributes +1.5; a left swipe contributes -1.5. The session posterior advances only after that production write succeeds. A swipe cannot teach the quiz while silently failing to teach the recommender.

Behavior supplies the levers#

A stored signal first decays the existing state, then adds a normalized entity vector multiplied by the signal’s weight:

v' = v * 0.5^(dt / 90 days)  +  w_signal * (x / ||x||)

After 90 days, an untouched contribution retains one half of its strength. After 180 days, it retains one quarter:

0.5^(180 / 90) = 0.5^2 = 0.25

That lets recent behavior matter without erasing older taste overnight.

Not every input is worth the same: lever height is the literal weight constant in the code.

Not every input is worth the same: lever height is the literal weight constant in the code.

Actions are weighted by commitment:

SignalWeight
Good attendance+3.0
Join a plan / Going+2.5
Save or follow+2.0
Keep in the deck+1.5
Reject in the deck-1.5
Plan activity or audition+0.5
Search click+0.4
Tap or link tap+0.3

A tap is weak evidence. Attending and reporting a good experience is ten times stronger:

3.0 / 0.3 = 10.

Negative actions matter too. The system should not interpret repeated rejection as a request for more exploration in the same direction.

How candidate scoring works#

For artist suggestions, Pull Up combines three interpretable terms.

The graph term uses the strongest related-artist edge from the user’s seeds. The bucket term uses Jaccard overlap, with a 0.15 bonus when the seed’s primary bucket appears in the candidate. Popularity is logarithmic, so the difference between 100 and 10,000 fans matters more than the difference between 1,000,000 and 1,009,900.

With a usable graph row:

score = 0.60(graph) + 0.25(bucket similarity) + 0.15(log popularity).

The weights sum on the page:

0.60 + 0.25 + 0.15 = 1.00.

Without a graph edge, the algorithm does not pretend the edge is zero-quality evidence. It changes models:

score = 0.70(bucket similarity) + 0.30(log popularity).

Three faders, one score per candidate.

Three faders, one score per candidate.

For show ranking, artist, venue, time, audio, direct affinities, friend activity, freshness, and availability enter later stages. Search keeps the full eligible inventory and sorts it by taste fit. Recommendation shelves can be narrower because they must also satisfy inventory, confidence, overlap, and repetition constraints.

Genre truth is correctable#

Genre is stored as its own artist fact. It is not merely copied from an event card.

All three genre mappers, runtime TypeScript, build-time JavaScript, and nightly Python, now use token and phrase boundaries. One shared fixture tests the same cases across implementations. dance may match “dance,” but it cannot match the inside of “dancehall.”

Artist-owned evidence wins: the ranked dataset, bucket rosters, and stored artist facts. Event evidence is a fallback only when those sources are empty, and it requires two corroborating headlining bills on distinct dates at distinct venues. A support slot or one unusually tagged festival cannot redefine the artist.

Merges are also correctable. A complete, non-empty refresh from every source already attached to a card may replace its genre set. Partial refreshes still union conservatively, because an absent source is not proof that its fact became false.

What I would not claim yet#

Nineteen of the 102 microgenres have no ranked noteverynoise page, and two top-level buckets are empty in the current ranked coverage. The algorithm can tolerate those holes, but tolerance is not coverage.

There are also 566 upcoming cards that remain untaggable. We hide them from recommendation surfaces rather than guess. That choice reduces recall in exchange for protecting genre truth.

String matching is not musical understanding. Token boundaries prevent known substring failures, but they do not resolve aliases, scenes, language, or artistic evolution.

Finally, one 30-second preview per artist is a thin audio sample, and roughly half coverage means audio cannot be the foundation of ranking. It is supporting evidence.

Those limits point to the system’s actual shape. Pull Up is not one taste model with omniscient inputs. It is a set of deliberately narrow sources, gates, vectors, and behavioral updates, each allowed to say only what its evidence can support.

Ayush Upneja
Written by Ayush Upneja

AI product engineer at Google. Nights and weekends I build AI that brings people together in person, including Pull Up.