Building a testing system for UGC ideas before scaling—how do you actually know what's worth big budget?

we ran into a real validation problem at the agency: we’d see one UGC format test at low budget, it’d show decent numbers—maybe 6-7% CTR—and then we’d say “okay, this is the one” and push $20K behind it. half the time it’d work, half the time it’d completely underperform at scale.

i started paying attention to what’s different between the tests that actually scaled vs. the ones that crashed, and i realized we weren’t testing the right things. we were testing creative, but we weren’t testing audience density, creative fatigue curves, or competitive environment.

so i built this informal testing framework:

  • week 1: small budget, 2-3 creator variations of the same core idea
  • week 2: double budget if CTR stays above threshold, rotate creators, add different placements
  • week 3: scale if the cohort retention metrics look good (not just raw CTR)

but honestly, i’m not confident this is the “right” way. i’m mostly doing this because it feels like it reduces waste. the problem is—i don’t have a clean metric for “when to scale.” is it when CTR plateaus? when LTV looks stable? when third-party validation shows viral potential?

how do you actually call the moment to scale? what indicators actually matter? and how do you know if w test failed because the idea was bad vs. because the execution was weak?

this is actually a well-studied problem in performance marketing, so let me share what actually works.

the mistake most people make is they look at week-one performance. that’s noise. what you need is cohort stability. here’s the framework i use:

Week 1-2 (discovery phase): run small tests, 3-4 creative variations. goal here isn’t to find the winner—it’s to eliminate the losers. anything below your baseline CTR gets cut. this is ruthless.

Week 3-4 (validation phase): take your top performers and run them at 2-3x week-one budget, but on new audience segments. if the CTR drops more than 25% on a new segment, that’s a red flag. that usually means the first test caught a niche audience, not a broadly resonant idea.

Week 5+ (scaling phase): only enter this if you’ve proved the idea works across at least two distinct audience segments AND your LTV cohorts are stable (meaning day-7 ROAS isn’t degrading vs. day-1).

about your question on execution vs. idea: i separate them like this. if CTR is good but conversion drops hard, it’s usually execution—landing page, messaging mismatch, etc. if both CTR AND conversion are bad, it’s the idea. if only CTR is bad, the audience might just be wrong.

i’d honestly recommend building a simple scoring system:

  • CTR stability across cohorts = 40% weight
  • ROAS by day (day 1 vs. day 7) = 40% weight
  • creative fatigue curve (does engagement decline linearly or exponentially?) = 20% weight

if your score is above 70 on this system, scale. if it’s below 50, kill it. this removes a lot of the guesswork.

я полностью согласна с Mark, но добавлю чисто аналитический момент: смотрите на вариацию результатов внутри вашего теста.

еслиодин креатор дал 8% CTR, второй дал 5%, третий дал 10%—это не значит, что идея хороша в среднем на 7.6%. это значит, что идея очень сильно зависит от контекста исполнения, и когда ты масштабируешь, ты рискуешь попасть на “низкого” креатора и получить 4% на 20K бюджета.

наоборот, если все креаторы дали 6.8%, 6.2%, 7.1%—это значит, что идея robust. ее можно масштабировать, потому что она работает независимо от исполнителя.

второе: я смотрю на метрику, которую люди часто пропускают—view duration и completion rate. высокий CTR с низким watch time может быть просто кликбейт, который вообще не конвертит. высокий CTR + высокий watch time = идея реально резонирует.

и последнее: не забывайте про seasonal effects и competitive landscape. я тестировала UGC в начале сентября, и она убивала. в октябре тот же контент упал на 30%. это не потому что идея плохая—конкуренция изменилась, сезонные тренды, что-то еще. всегда учитывайте этот контекст при принятии решения о масштабировании.

я бы добавил одно практическое правило: никогда не масштабируй на реальный клиентский бюджет, пока не прошёл две недели в живых условиях.

у меня была ситуация, когда бриф выглядел отлично на тесте, но когда я запустил его полноценно, выяснилось, что конкурентные объявления занимают тот же ad space, и мой контент просто теряется. тестовый бюджет был слишком мал, чтобы это заметить.

когда я начал требовать минимум 10-14 дней в реальных условиях, перед тем как рекомендовать масштабирование, результаты улучшились на 40%.

и еще—всегда тестируй не только контент, но и аудиторию. один и тот же UGC может убивать на холодной аудитории, но не резонировать с теплой. или наоборот. понимание, где этот контент работает лучше, это часто более ценно, чем сам контент.

с моей стороны—я замечаю, что много людей тестируют контент, но не тестируют timing и frequency. я могла создать идеальный UGC, но если его запустить в 3 часа ночи по московскому времени, это не покажет истинный потенциал.

также—people fatigue. если один творческий кусок работает хорошо, это не значит, что он будет работать, когда его покажут одному и тому же человеку 5 раз подряд. я тестирую не только raw performance, но и как вычитается это в реальном юзер-джорни.

и еще один момент—я всегда советую тестировать с разными ценовыми точками. контент, который хорошо конвертит товар за 50 баксов, может вообще не работать на товар за 500. это разные психологические барьеры.