Strategy

Keep It or Kill It: An Operator Case from an Interview

01 Oct 2026StrategyAI & ProductPrevious Work

Written as an interview assignment for an Operator role at a community startup. The role didn't go ahead, so I'm publishing the work, with the company and its data kept anonymous.

The company runs a membership community. Its biggest bet is connection and discovery, and that side of the business is moving towards running itself: self-run meetups, member search, small groups that meet without a host.

One thing doesn't fit. Its live learning sessions, with one facilitator, one topic and one session in the room, are the best-performing thing it runs, and they can't run without a person.

The case had two parts. First, should the company keep or kill learning over the next eighteen months, and how would you prove it? Second, a stretch: automate the LinkedIn outreach a senior team member was doing by hand.

Part 1: rebuild every claim from the rows

The brief came with a summary and the raw registration log. I rebuilt every claim from the log before accepting any of it, and labelled each finding as observed, calculated, assumed or a proposed threshold, so nobody could mistake a guess for a fact.

The summary and the log disagreed in places. Approvals had been read as paid seats, though some were free. Approval was also being read as attendance, when the log had no way to show who actually turned up. None of that is unusual. Summaries drift from the data they summarise, which is why the decision should come from the rows.

What held up

Repeat demand was real. A large share of registrations, and of revenue, came from people who had already been to an event. I computed a stricter version of each metric alongside, and the conclusion didn't move.

The more interesting finding was where people went next. Most repeat customers crossed topic categories. They were following practical progress, or following the community itself.

Price mattered less than the brief implied. At the same ticket price, demand varied widely between the strongest and weakest topics, so strong topics could carry a higher price. The data couldn't say how sensitive anyone was to price, because topic mix, timing and marketing all changed at the same time.

One category looked like the weakest. But a specific, practical session inside it did well, so the lesson was about the proposition.

The recommendation: keep it, but make it earn its calendar

Keep learning for eighteen months as a measured portfolio inside connection and discovery, in four layers:

As an operating rule: of every twelve sessions, nine go to proven formats, two to adjacent tests and one to something genuinely new.

How to prove it right or wrong

The data could show demand. It couldn't show cost, attendance, learning or whether any of this kept members. So the plan starts by measuring what's missing.

Months one to three add costs, attendance, capacity, acquisition source and facilitator to every event, and run comparable price tests. Months four to nine raise the share of proven themes and test one or two pathways. Months ten to eighteen keep what's profitable and stop what misses its thresholds.

The thresholds are set in advance, from the current median, so nobody can move them after seeing the results. Learning is judged on contribution margin, paid attendance, repeat purchase, membership renewal and connections created. Gross revenue alone doesn't decide it.

Part 2: an outreach machine that knows when to stop

The manual process was: find someone who fits the community, decide if they really do, write a personal invite, send it from a senior team member's account. The brief asked for all of it automated, end to end.

I built it as a local pipeline. A Chrome extension captures a batch of search results. A second capture records evidence from a profile a person has opened. Claude classifies each profile, drafts a note in the senior sender's voice only when every check passes, and puts it in a review queue. The last state is ready to send. There is no send button, by design.

What sourcing taught me

I tested four search approaches on a real account before trusting any of them. Broad queries returned mostly the wrong profession. Narrower queries returned mostly the wrong audience, because LinkedIn can't filter on what the community actually screens for. Natural-language queries were ignored, and strict ones returned nothing.

That shaped three rules. Search cards can't establish fit, so nothing gets classified from a card alone. LinkedIn's own filters aren't reliable, so location is checked again downstream. And audience fit is never inferred from a name or a photo. It counts only when the profile says it, or when the source is a community that already screens for it.

How the AI decides

Three separate gates: career fit, location and audience evidence. Career fit is yes only when a quoted line from the profile shows the signal the community looks for, such as a recent business, a job plus a side build, or a pivot with current activity. Thin or conflicting evidence comes back uncertain, and uncertain is a valid answer.

On a live run of ten real profiles, seven came back as a fit and three as uncertain. The uncertain ones went to a person to decide. Six got drafts. Four were held back, each with a recorded reason.

The best catch was a strong founder who scored near the top on career fit and still got no draft. Her profile put her in a different country. LinkedIn's location filter had let her through, and the pipeline caught it.

Where it went wrong

The judgement held. I didn't override a single classification on that batch. The miss was in the voice: an early batch of drafts used punctuation the sender would never use. I fixed the generator, so the rule lives in code.

The bigger risk was sourcing noise. LinkedIn puts your mutual connections inside each result card as profile links, and early batches picked up people who were never in the results. Cleaning that up mattered more than tuning any prompt.

Before it runs on its own

A senior person's name on a note is a promise that they'll follow up warmly, so they have to approve the voice and sample drafts first. A restricted personal account is an inconvenience. A restricted founder account is a business problem.

Before it runs unsupervised at volume, I'd want measured precision on the yes decisions, weak evidence reliably routed to uncertain, close to zero invented details in drafts, a falling edit rate, hard rate limits with a kill switch, and proof that accepted connections turn into conversations worth having. Until then, a person stays in the loop.

What I'd keep from it

The brief asked keep or kill. The honest answer was keep it, with a date, set now, by which you'll know if that was wrong.

Both parts came down to the same habit: deciding in advance what would prove you wrong. The most useful thing the pipeline did was refuse: the uncertain profiles it wouldn't force, and the strong founder it wouldn't invite.

Building something where execution feels heavier than it should?

Tell me what's slipping. I'll tell you what I see.

letsbuild@yashasvishailly.com Or start with The One Fix