Case study · growth engagement · Nov 2025 – Apr 2026

ManyChat: affiliate growth on an AI agent fleet

For five months, as a contractor, I ran ManyChat’s affiliate program end to end — recruiting, activation, reactivation, partner quality, partner support — and built the first agentic growth engine in the company’s history: five production AI agents on top of PartnerStack and ManyChat’s own ops systems, from scratch. Channel revenue grew about 1.2× in four months, roughly $2M in ARR added, #1 Trending on PartnerStack — and the company still runs on what I built. The machine is on this page.

The problem

ManyChat pays affiliates to bring it customers. The channel had tens of thousands of partners on the books, applications arriving faster than anyone could research them, most of the book dormant, and a revenue line that had just fallen for two straight months. The brief was growth: more good partners in, dormant ones back, more output from the active ones — built on revenue the program can trust.

Results

Exact internal figures stay inside the program — the agreement is explicit and I keep it. What I publish are multiples, ranges and shapes.

~1.2×partner-channel revenue over the four months I ran itNov 2025 → Feb 2026 · month-by-month chart below
~$2MARR added across the engagementannualised; the order of magnitude reported at delivery
25×partner-chat messages per month, Nov → Janthe program’s own slides say 29×; I print the stricter count
#1Trending on PartnerStack Discover Partnershipsthe handoff document’s own words
~300×faster application decisions — about 10 minutes, from 2–3 days manualmeasured on the first two hundred applicants, Dec 2025
1,900+applications researched by the agents, Nov – Febevery one scored, answered, and decided by a person

What I owned

The whole affiliate function, not a component of it. Tens of thousands of partners on the books, a thousand-plus active in any given month — recruiting, the application gate, partner communications, reactivation, quality review and the reporting all ran through the machine described below, with two ops people and me above it. Built from scratch inside the contract window, migrated onto ManyChat’s own infrastructure in February, and still in use after I left.

The strongest artifact of scope is the successor JD. You only get to write that document when the function is yours to hand over.

Its “What You’ll Own” section lists five areas: AI agent quality & performance · operational metrics & analytics · partner journey optimization · system & workflow management · strategic program growth. Its “Key Metrics” row names scoring accuracy, approval-to-activation conversion, decision time, AI response approval rate, partner activation rate, revenue per partner and reactivation success rates. JOB_DESCRIPTION — Head of Affiliate, written 17 Dec 2025 for the hire replacing me

The growth machine

The first agentic growth engine in ManyChat’s history — five stable production AI agents, built from scratch on top of PartnerStack and the company’s own ops systems and processes. The five working parts below are what that engine ran during the engagement — software, not a plan. The decision pipeline, the dossiers and the review workspace further down this page are lifted from the repository and pointed at anonymised fixtures.

1 · Inbound activation — the application funnel

Applications were the raw ore: by January the program drew close to a thousand inbound applications a month. Every one was researched by the agent in about ten minutes, scored 0–100 against the rubric below, and answered — approvals got an onboarding message and follow-ups on the same rail, holds got clarifying questions instead of silence. And the gate became real selection. Before the machine, approval was a formality — effectively every applicant got access, ~99%. With scored dossiers and authenticity checks it settled near ~62%: the program now chooses the two applicants in three who will actually ship customers, and the book stopped diluting. Speed at the front door had its public side effect too: ManyChat became #1 Trending on PartnerStack Discover Partnerships, which itself feeds the top of the funnel.

2 · Reactivation — waking the dormant book

Most of a mature affiliate book is asleep, and the expensive part is knowing whom to wake. The CRM classified every partner’s trajectory from transaction history — dormant, declining, at risk — and the agents wrote personal outreach against each partner’s own numbers. Thousands of partners got that outreach over the engagement, and the priority queue started at the top of the revenue book, where most of the largest historical earners had churned at some point. In a channel this concentrated a single recovered whale pays for the whole rail — several came back.

3 · Chat operations with active partners

The channel that talks to partners went from dormant to daily: monthly partner-chat volume grew 25× between November and January, and active partner rooms went from a couple hundred to over two thousand. None of it was spray. The book was segmented by what a partner actually brings — how many users, of what size, on what trajectory — and activation and account attention were prioritised down that ranking, so the rooms that opened were the rooms that move revenue. That prioritisation, not volume for its own sake, is where the 25× and the active-room growth came from. By January roughly half the messages were agent-written — each one drafted against the partner’s dossier and the program’s pricing truth, with ops able to approve, edit or regenerate. V3, live 18 February, added plan-vs-fact per partner: it reads a partner’s trajectory and reaches out before the decline shows up in a quarterly review.

4 · Revenue quality — growth the program keeps

A channel that pays for customers has to know where its customers come from — growth is only worth reporting on revenue you can trust. The agents gave the program that certainty at both ends. At the gate, every applicant carried an authenticity check next to the score — growth patterns, engagement in proportion to audience, identity consistent across platforms; the dossier cards below show the signals. Across the live book, the system read acquisition sources against the partnership terms for the thousand-plus active partners on a schedule — coverage that had previously depended on someone finding a spare hour.

The most consequential finding was quiet: partnerships acquiring through channels the program reserves for the brand itself — its own search demand — and reselling that demand back with a commission on top. Their users were real, paying and staying, which is exactly why the pattern survives casual review. The program parted ways with the partnerships behind more than 10% of channel revenue, kept every one of their users — they had come for the product — and re-based its growth on revenue it fully owns. The chart below is measured after that reset. And every decision stayed with a named person; the agents brought the evidence.

5 · The dashboards management steered by

A Go + React dashboard over the program database carried the CRM, the application queue and eight analytics reports. The view that mattered most was concentration: out of tens of thousands of partners in the database, a few hundred produced 80% of all customers — and the raw PartnerStack exports on my disk show the same shape independently of my own dashboards. That concentration is why per-partner chat operations pay for themselves: moving one large partner moves the month.

The turn

The four months I ran the channel, indexed to the month the contract started (November 2025 = 100). Absolute dollars stay inside the program — the shape is the argument: a line that had just fallen two months in a row turned in the exact month V1 went live, and compounded every month after.

indexed, November 2025 = 100 · the four contract months only — the months before I started are not mine to show · absolute channel revenue is deliberately not published

Underneath the total, the composition moved the right way: first-transaction (new-customer) revenue jumped about 60% in the first full month after V1, and returning-customer revenue turned up behind it. New money moved first — which is what you want to see if the funnel work is real. Where raw PartnerStack exports exist on my disk, they agree with the dashboard series month for month.

Attribution, kept honest

From application to decision

The same five steps for every applicant, with a person holding the last one.

01 IntakePartnerStack applications land in Mongo / CSV exports
02 Researchn8n “3. Research [final]” agent + Firecrawl / Bright Data tools
03 Score0–100 rubric: Reach 40 · Fit 35 · Engagement 15 · Identity 10
04 DecideAPPROVED_EXCELLENT / MONITOR / MANUAL_REVIEW_*
05 Ops gateops.html — search, filter, open dossier, human decides

What a dossier looks like synthetic fixture

Three example applicants as the review workspace renders them. The people are invented; the shape is the production one — channels found, audience per channel, the four factor scores, an authenticity check, and a decision the human gate can still override. The dossier excerpts are written in the agent's actual register: claims checked, sources counted, doubts stated.

Marina Voss
Course creator · education (1–10) · Netherlands
86A · Platinum
approved_excellent
Reach32 / 40
Fit34 / 35
Engagement13 / 15
Identity7 / 10
YouTube 96,400 newsletter 14,200 site 31,000 / mo Instagram 12,800

Teaches automation to course creators. Every audience claim checked out against two independent sources. Last posted 4 days ago; the comment threads are real conversations, not engagement pods.

authenticity 0.985.8 min · 13 tool calls
Diego Ferrer
Marketing agency (11–50) · Brazil
58B- · Silver
monitor
Reach24 / 40
Fit23 / 35
Engagement6 / 15
Identity5 / 10
Instagram 41,300 YouTube 8,900 site 6,400 / mo

A real agency with real clients, but the Instagram audience grew in two step-changes and engagement is thin for the follower count. Approved with monitoring; the first payouts get a routine second look.

authenticity 0.826.9 min · 15 tool calls
Anya Petrov
SaaS review blog (1–10) · Germany
44C · Bronze
manual_review_clarify
Reach12 / 40
Fit22 / 35
Engagement6 / 15
Identity4 / 10
blog 9,800 / mo LinkedIn 1,140

The niche fit is genuine — three years of automation tutorials — but the claimed email list could not be verified from any public source, and the application email is a free mailbox with no domain behind it. Held with two clarifying questions.

authenticity 0.697.4 min · 16 tool calls

Research time per applicant, illustrative fixture values: median 6.2 min · p90 11.4 min · 14 median tool calls · a tool fallback fired in about 2 of 10 dossiers. The real timings lived in n8n execution logs and were not exported.

Score distribution and decision mix

Both charts are computed from the anonymised fixture the demo workspace runs on. The fixture kept the production distribution shapes and dropped the people, so the bars below are the real shape of the pipeline's output — how rarely a clean application arrives, and how much of the queue lands on a human.

Scores, 79 scored dossiers

computed from the fixture · shape preserved from production, people invented

Decision mix, 84 dossiers

Approve · 38 Manual review · 35 Decline · 6 Unscored · 5

ops-level enum; the agent's four outcomes collapse into these three chips

Two readings worth holding together. Nearly half the queue is a clean approve, which is the point of the machine — the human's day goes to the 42% in the middle, where a dossier earns its keep. And the decline chip exists at all only because a person issues it; the agent can only recommend.

The line between approve and monitor

Same pipeline, same week, two invented applicants. Diego's reachable audience is larger than Marina's newsletter. The 28-point gap comes from engagement and identity, the two factors that answer "is this audience real" rather than "how big is it".

Marina Voss · 86 approved_excellent

Every claim survived contact with a second source.
  • Audience claims verified against two independent sources each
  • Engagement proportional to follower count; posted 4 days ago
  • Same person across three platforms, same face, same name history
  • Authenticity 0.98 — every check clean

Diego Ferrer · 58 monitor

A real business — with two signals worth a second look.
  • Follower count grew in two step-changes, months apart
  • Engagement 0.4% where the niche usually shows ~2%
  • Identity matched on one of three platforms
  • Authenticity 0.82 — fine to proceed, worth keeping an eye on

Both cards are synthetic fixtures. The signals they illustrate — step-change growth, engagement-to-follower ratio, cross-platform identity match — are the real ones the rubric scores.

Scoring rubric — actual weights

The production research agent scored each applicant out of 100 on four weighted factors. Reach and fit are scored independently on purpose: a specialist with a small audience of practitioners can take full marks on fit while scoring low on reach, so the total says something about the partner rather than restating their follower count.

FactorWeightWhat it measures
Audience Reach40 / 100Verified total audience across every channel found
Profile Fit35 / 100Expertise and use case — deliberately independent of audience size
Engagement Quality15 / 100Recency and genuineness of interaction, not raw counts
Identity Verification10 / 100Whether the person is who the application says

The score and a separately-accumulated authenticity check then map onto four operational outcomes rather than a verdict — approve, approve with monitoring, hold with clarifying questions, or hold as a recommended decline. Two of the four are holds, and the agent has no authority to issue a decline itself. The exact thresholds were the client's and are not published here; the anatomy page explains why each rule exists.

Ops UI decision chips collapse research outcomes into APPROVE / MANUAL_REVIEW / DECLINE for the human gate — that mapping is in the unified dataset’s summary.decision, not invented here.

The working software (lifted)

Everything above describes the machine; this is the machine. The production review workspace, running here against anonymised fixtures with a visual polish pass on top — the dossier cards earlier on the page are three of its rows, drawn out for reading. The demo population is the demo’s only: 91 applicants, 84 dossiers — fixture counts, no relation to the client’s volumes above.

app/ops.html · Partner approvals workspace Open full screen

The approvals workspace — where every AI dossier waits for its human decision. Search, filters and scoring logic are the production ones.

Two more pages of the same app open full screen rather than embedded here: the applications intake queue · the program home — score mix, pipeline and a partner leaderboard the page itself labels synthetic.

n8n research workflow (artifact)

Rendered from the exported workflow JSON — nodes, types, and connections only. No screenshot, no invented studio UI, no execution payloads. In production this was one of ten workflows in the ai-partner-program project, organised as two pipelines — applicant processing and the chat response agent.

When Executed by Another WorexecuteWorkflowTrigger
Structured Output Parser1outputParserStructured
OpenRouter Chat ModellmChatOpenRouter
#agent
BrightData - Custom API callbrightDataTool
OpenRouter Chat Model4lmChatOpenRouter
Loop Over Items1splitInBatches
Json parser1agent
Get applicants from Mongo1mongoDb
Set new status1set
Update status in Mongo1mongoDb
Set result data1set
Update result in Mongo1mongoDb
MCP Firecrawl1mcpClientTool
Access and extract data frombrightDataTool
Extract structured data frombrightDataTool
List available datasets in BbrightDataTool
Call next workflow1executeWorkflow
Aggregate1aggregate
Webhook1webhook

Rendered from the exported workflow definition — node types and connections only.

Denominators, kept apart

More populations run through this page than any single number admits: the full partner database in the tens of thousands; the thousand-plus active in a given month; the applicants inside a single export window; a regional recruiting book of hundreds inside the global program; and the 91-applicant anonymised fixture the demo workspace runs on. Each figure above names its population and window, and none of them is ever added to another.

A rate quoted against the wrong denominator answers a question the reader never asked. Most reporting mistakes in a channel like this are that mistake rather than a modelling failure.

What was hard

What is real

The review workspace, the applications list and the program home are the production pages, running here against anonymised fixtures with only a visual polish pass on top. The scoring weights are the ones the system used. The workflow graph is rendered from the real exported definition. The growth numbers come from named artifacts — the program’s reporting dashboards, the February 2026 technical handoff, the delivery summary to marketing leadership — cross-checked against raw PartnerStack exports on this disk. Absolute channel revenue is deliberately not published; dynamics are shown as an index.

What is reconstructed / transformed

All names, emails, employers, URLs, handles, and dossier prose are anonymised. The three dossier cards above and the research-time figures are invented illustrations, labelled where they appear; the partner leaderboard on the app's Program home page is synthetic and says so on the page. The February reporting dashboards contain real partner identities, so they are quoted as aggregates here and never embedded or screenshotted. Applications rows were reshaped to the UI schema the table already expected, and 331 rows whose group column had captured CSV answer fragments were repaired to a real tier. Prompt anatomy page describes the real production prompt section by section without reproducing the client's file.

Also read: Anatomy of a production prompt

More from my work