Off-the-shelf topic and sentiment models tag feedback in someone else's categories. The useful tags are yours, and they change every time the product does.
| Pitfall | What it looks like | How to check, before spending anything |
|---|---|---|
| Frequent tags dominate | Overall scores rise while rare tags go unfound | Macro-average per-tag recall; read the rare tags on their own. |
| Agreement ceiling | Scores plateau well short of perfect | Compare against how often your analysts agree. That is roughly the ceiling. |
| "Other" as a crutch | A growing share lands in a catch-all tag | Read a sample of "other". It is usually a missing tag or a vague definition. |
| Mixed taxonomies | The same comment carries old and new tags | Relabel under one version of the taxonomy before scoring. |
| Sentiment bleed | Negative comments get tagged as problems | Keep sentiment a separate field from topic. |
Tools that run this loop: DSPy and its GEPA optimizer, bpto (ours, open source), and Impromptune, the studio built on it.