Back to work
AICreator SupportMobile App

Where the the model stops, and the creator decides

AI-assisted topic tagging for an advertising platform

A shot of three wireframed screens for a mobile app, in dark mode, annotated with green lettering

Skills

User research · Designing AI features · Human-in-the-loop design · Working within a controlled taxonomy · Two-sided marketplace design

It's hard to find the right podcast

This client had a platform that connected small and mid-sized podcasts with local advertisers looking for niche audiences. Discoverability made this market work: advertisers needed to find the right show, listeners needed to find the right episode, and creators needed to be found at all.

I talked to several podcast creators about how their days actually went: what ate the most time in making a show, and where they wished they could get some of it back. The striking pressure all the creators mentioned was how hard it was to find their audience, and generally be heard in a crowded podcast market.

The notion of tagging the content of podcast episodes sat inside the pressure of constant growth. It wasn’t a filing chore, tagging is one way a show gets discovered by new listeners. Automating the taggin wasn’t about saving ten minutes, it was about taking one piece of the never-ending job off someone’s plate without taking away their say in how their show gets represented.

A digital whiteboard with notes from research

Use a shared vocabulary that already exists...

We tagged podcast episodes against the IAB taxonomy, rather than allowing the generation of free-text labels.

Advertisers already bought against IAB categories, so tags that matched their vocabulary were immediately useful on the buy side. Listeners got refined, consistent topics to search rather than a thousand near-identical variations. And constraining the model to an existing, agreed-upon set meant it couldn’t invent a plausible-sounding tag that no one else in the ecosystem used.

...and keep the creators in the loop

Tags were how their show was described to advertisers and to anyone searching. The system generated topic tags and notified the creator that they were ready to be reviewed. From there, creators could remove, edit, or add anything they felt was more accurate. Tags went public only after they approved them.

If a tag was wrong, the creator removed it, and it never reached anyone. There was no bad-tag cleanup problem, no advertiser matched against a category the show didn’t belong to, no listener misled.

Not every feature made the final cut

I wanted model confidence surfaced alongside each proposed tag. These were very busy people. Realistically, a lot of them would glance at a list of eight tags, see nothing obviously wrong, and hit approve. Confidence would have told them where to actually look to avoid potential trouble.

It didn’t make the final cut, because nothing published without human review. A wrong tag was a low-cost risk. Confidence scoring was a real build cost to reduce a failure that was already contained to an acceptable degree.

In the End

Better models would change what the feature could be, just a couple years later, but they shouldn’t change whose call it is.

The feature wasn’t about AI in the end, it was about authority. Who decides when a system makes a judgment about someone’s work? If anything, a system proposing which moments of your show represent you needs that gate more than a system proposing topic labels, not less.