essays
2026-09-01· 8 minPull UpEngineeringRecommendersPostmortem

Skrillex, Because You Like Burna Boy

A recommender produced an absurd result with correct arithmetic. The bug was three genre mappers, event-lineup inheritance, and a merge policy that never forgot.

Ayush Upneja
AI product engineer

Pull Up shipped a live recommendation caption that read: “Skrillex, because you like Burna Boy.” The similarity calculation was working as written. The artist data underneath it was wrong.

Before the repair, both Skrillex and Zedd scored 0.65 against a Burna Boy seed. Wizkid, Burna Boy’s actual Afrobeats peer, scored 0.3333. After the repair, Wizkid and Davido scored 0.65, while Skrillex and Zedd scored 0.0000.

That inversion was the receipt. The bug was not one bad edge or an unfortunate ranking weight. Three genre mappers, event-lineup inheritance, and an accumulate-only merge policy had collaborated to manufacture a false musical world.

Three sources, one crate: an artist enters the bank only when the genre roster, the event index, and Deezer all agree.

Three sources, one crate: an artist enters the bank only when the genre roster, the event index, and Deezer all agree.

The recommendation was embarrassing, but internally consistent#

Pull Up recommends artists and upcoming shows from several kinds of evidence: selected genres, named artists, an artist graph, genre buckets, popularity, event metadata, and behavioral taste signals.

The suggestion path behind the bad caption used three terms when graph data was available:

score(c) = 0.60 * G_c  +  0.25 * B_c  +  0.15 * P_c

Where:

  • G_c is the strongest similar-artist graph edge across the user’s seeds.
  • B_c is the strongest genre-bucket similarity across those seeds.
  • P_c is normalized popularity.

Bucket similarity uses Jaccard overlap plus a bonus when the candidate contains the seed’s primary bucket:

J(c, s) = min( 1,  |B_c intersect B_s| / |B_c union B_s|  +  0.15 if s's primary bucket is in B_c )

This is intentionally simple. If two artists share one bucket out of a two-bucket union, their base overlap is 1/2 = 0.50. If the candidate also contains the seed’s primary bucket, the result becomes 0.50 + 0.15 = 0.65.

The suspicious result therefore had a coherent explanation: the system believed Burna Boy, Zedd, and Skrillex shared genre evidence. Given that premise, 0.65 was not a math error. It was the expected answer.

Three faders, one score: the graph does the real work at 0.60; the bucket term is the one that was poisoned.

Three faders, one score: the graph does the real work at 0.60; the bucket term is the one that was poisoned.

Pull Up actually had two taste spaces#

The quiz looked more sophisticated than the bad caption suggested. It maintained a 28-dimensional posterior: 24 genre buckets plus four room-preference axes. Pairwise choices updated that posterior with assumed-density filtering. BALD acquisition used its mean and covariance to choose questions expected to reduce uncertainty.

The persistent recommender used a different, 142-dimensional space:

  • 24 genre buckets
  • 102 microgenres
  • 8 audio features
  • 8 scene traits

The quiz posterior and the recommender vector were not two views of one state. They were separate systems joined by a narrow handoff.

At onboarding commit, ranked genre labels and index-resolved artist names crossed that handoff. The posterior’s pairwise bucket magnitudes, covariance, uncertainty, and microgenre leans did not. Four room-axis means were written, but the code comment at the write site was unusually candid:

Axes have no reader yet.

So the adaptive quiz could learn something, display an archetype from it, and persist it for research while the feed reconstructed taste from much coarser labels. The quiz was not feeding its full posterior into ranking.

The quiz brain and the feed brain are different sizes, joined by a narrow handoff.

The quiz brain and the feed brain are different sizes, joined by a narrow handoff.

That disconnect did not create the Skrillex recommendation. It made the data feeding the recommender even more important. The narrow handoff elevated genre classification from one feature among many into a load-bearing source of truth.

The first bug was three substring matchers#

Genre classification existed in three implementations:

  1. Runtime TypeScript.
  2. JavaScript used to derive rosters and artist banks.
  3. Python used by nightly ingestion and verification.

They were similar, but not identical. All three performed substring matching somewhere in their logic.

That produced deterministic classifications such as:

"dancehall" contains "dance"  -> EDM
"dubstep"   contains "dub"    -> Reggae / Dancehall
"djent"     contains "dj"     -> EDM

These were not fuzzy-model mistakes. They were plain string containment.

The build-time mapper had an additional divergence. It could recognize an exact canonical microgenre and then continue applying substring fallbacks. An exact dancehall page could therefore become both Reggae / Dancehall and EDM during roster generation, even though the runtime mapper stopped after the exact match.

The stamp only reads the first five letters: dancehall becomes dance, dubstep becomes dub, djent becomes dj.

The stamp only reads the first five letters: dancehall becomes dance, dubstep becomes dub, djent becomes dj.

The fix was token and phrase boundaries in all three paths, backed by one shared regression corpus. dance still maps to EDM when it is the token dance. It no longer fires inside dancehall. Canonical microgenres short-circuit keyword fallback consistently.

That closed the obvious classifier bug. It did not explain why Zedd had become Afrobeats.

Zedd played Randall’s Island, so Zedd became Afrobeats#

The precompute pipeline split an event’s artist and lineup text, took up to the first three names, and assigned every event genre to every extracted artist.

That rule treats a bill as a description of each performer. A support act absorbs the headliner’s genres. A festival artist absorbs the festival card’s genres. Playing one mixed bill can rewrite an artist’s musical identity.

Zedd appeared on a Randall’s Island event whose card carried Afrobeats among its genres. The pipeline consequently attached Afrobeats to Zedd. The false tag then flowed into generated artist banks, bucket co-occurrence, and similarity suggestions.

The replacement rule gives artist-owned evidence priority:

  1. Stored artist buckets.
  2. The checked-in ranked microgenre dataset.
  3. Checked-in bucket rosters.
  4. Event evidence only when the first three are empty.

Event evidence now has to corroborate itself. The same canonical bucket must appear on at least two headlining events with distinct dates and distinct venues. Each qualifying event may carry no more than two canonical buckets. Support slots, one-off bills, repeated venue appearances, and broad festival cards do not establish an artist genre.

That is a deliberately high bar. Missing a weakly supported tag is less damaging than teaching the entire recommendation system that one booking defines an artist.

Bad genres could enter, but they could not leave#

The nightly merge policy completed the trap:

sorted(set(existing.genres + fresh.genres))

A fresh source could add a genre. It could never retract one.

That policy is attractive in ingestion systems because union feels safe. If one scrape omits a field, retaining the old value prevents accidental data loss. But union is safe only when facts are monotonic. Genre classification is not monotonic. A better mapper, a corrected source, or an artist-specific record must be able to invalidate an old tag.

We changed the merge contract so a full, non-empty re-report from the sources that previously contributed can replace stale genres. A partial refresh still unions. An empty genre list cannot erase a card, because an empty scrape may be a data gap rather than a retraction.

The distinction is provenance, not freshness. Newer data is not automatically better. A correction is valid only when we know which earlier evidence it supersedes.

The review changed the diagnosis#

A second model, GPT-5.6 Sol, adversarially reviewed the initial diagnosis. That review surfaced the misattributed root cause: fixing substring boundaries alone could not explain or remove Zedd’s Afrobeats tag. Following the value backward exposed event-genre inheritance, then the accumulate-only merge that preserved its output. The useful contribution was not another proposed ranking formula. It was refusing to accept a plausible first explanation when the artifact still could not be derived from it.

Regeneration was part of the fix#

Correcting source code did not correct the checked-in and stored derivatives already built with the old rules.

The dependency chain included quiz precompute output, wide and slim artist banks, bucket rosters, the roster pool, genre-space metadata, and Redis-backed artist facts. We regenerated them in dependency order against the corrected mappers and artist-evidence policy.

The resulting bank became smaller:

2,078 - 1,788 = 290 artists removed
290 / 2,078 = 13.96%

Those 290 entries were not removed to hit a target size. They fell out because the corrected evidence no longer admitted them into the generated bank.

Ten artists had carried the false Afrobeats plus EDM pair. After regeneration, zero did, a reduction of 10.

The direct similarity receipt then flipped:

Burna Boy seed candidateBeforeAfter
Wizkid0.33330.65
Davidon/a0.65
Zedd0.650.0000
Skrillex0.650.0000

Same formula, fixed catalog: the inversion is the receipt.

Same formula, fixed catalog: the inversion is the receipt.

The test suite had been agreeing too easily#

The same review found a separate failure in the Python test infrastructure. The ingestion module imported an optional dependency at module load time. When that dependency was unavailable, the module called:

sys.exit(0)

Test discovery imported the module, received a successful process exit, and stopped. The suite looked green because the interpreter had been instructed to report success before collecting everything.

After moving the dependency failure to the executable path and making it exit loudly with code 78, Python discovery increased from 194 to 200 real tests:

200 - 194 = 6 newly executed tests
6 / 194 = 3.09% more than the apparent old count

The important number is six, not the percentage. Six tests existed and had not been running.

The audit also found a more subtle dead measurement. Receipt-ranking duels compare two artist vectors using:

d = v_winner - v_loser

When both artists have identical bucket vectors, d = 0. The UI can change their displayed order, but the posterior receives a zero direction and learns nothing. Nina Kraviz versus Will Sparks was one concrete identical-vector case in the bank.

The pipe is much narrower than the tank: only ranked genres and artist names reach the recommender.

The pipe is much narrower than the tank: only ranked genres and artist names reach the recommender.

What I regret, and what remains unresolved#

I regret treating generated genre arrays as inert build artifacts. They were executable policy with a longer lifetime than the code that produced them. The repair was not complete until every dependent artifact had been regenerated and the before-and-after recommendation had been scored directly.

I also regret that the three mappers were allowed to evolve independently. A shared fixture now constrains them, but one classifier implementation or one generated matcher would be a stronger design.

Several uncertainties remain:

  • latin pop still has an unresolved dual-mapping decision.
  • 566 upcoming cards remain genuinely untaggable. They stay hidden rather than receive fabricated genres.
  • Token-boundary matching is still string matching. It avoids djent becoming EDM because it contains dj, but it does not understand music.
  • The 28-dimensional posterior still does not reach the recommender as a posterior. The axes write still has no reader.
  • Identically tagged artist duels still need contrast gating or richer artist vectors.

The lesson I am keeping is narrower than “data quality matters.” When a recommender produces an absurd explanation, inspect the exact features that made the score rational before tuning the score. Our 0.65 was correct arithmetic over a fictional catalog.

Ayush Upneja
Written by Ayush Upneja

AI product engineer at Google. Nights and weekends I build AI that brings people together in person, including Pull Up.