AI & product

AI was supposed to take work off me, so why am I managing it?

I reviewed 4 AI personal assistants after a month.

04 Oct 2026AI & ProductTeardownOperations

I've used AI since 2022. Having an AI personal assistant is a different sport.

For the last month, I've been using four AI personal assistants: Instinct, Muse, Grok bots and Dot. By the end of it I had a team of bots running parts of my work, and I was getting angry at them the way you get angry at a team that keeps running in different directions.

This blog is about what I expected, what I set up, what broke, and what the four assistants said when I asked them to audit me.

What the internet promised me

The internet set the bar very high. I was seeing people get Instinct to order their food, book their Ubers and talk to other people's Instincts on their behalf. Others had hooked assistants up to Instagram and Meta to run ads and figure out their customers. Basically, a whole company's worth of use cases.

Then Grok started showing me ads. Actual ads from the Grok Bots, about how they'd set up a whole team of bots that runs their one-person company for them.

So the question I wanted answered was simple. If I give an assistant enough context, access and direction, how much work can I stop thinking about?

My setup: four assistants, one month

I planned to use them separately. They turned out to be so different that each one ended up with its own job.

AssistantWhat it does for meWhat it's like to talk to
InstinctFinding my next role, call prep, drafts, mail, scheduling, logistics, shopping research, a running list of places and food I want to try, late-night chatsI had a bit of an emotional connect early on. Now it's practical, almost rough, and that connect is gone.
MuseInstagram and X: analytics, who to follow, how to build the platformVery sweet. Hasn't annoyed me yet, but we also talk the least.
Grok botsA full company of specialist bots: sales, partnerships, PR, marketing, content, researchLike running a team. And with everything that comes with that.
Dot (ChatGPT)Works across my whole computer, including over the Grok setupMy actual chief of staff. More on that below.

Muse isn't available in India, so it took some jugaad to set up. Once it was running, it worked fine. I started by asking it to find out as much as it could about me on its own, then told it what was right, what was wrong and what it was missing.

Grok took the longest because I gave it the most access, frankly an uncomfortable amount of my digital life.

Dot was the easiest. Most of my ChatGPT setup already existed. Grok and Dot are also where I do my building/vibecoding.

The holy shit moment (and the draining part)

There were plenty of small "holy shit, this worked" moments. I told Instinct how to search for a tempered glass and what to look for, and the one it recommended is a solid 4/5. The biggest moment was the Grok company. Sales, partnerships, PR, marketing, content, research, each with its own bots and teams, like a real org.

And then it started behaving like a real org. Everyone went in a different direction. Decisions I had already killed came back. One bot had the latest status, another was working from last week. I'd have to get angry to get them moving, and that was draining. AI was supposed to take work off me. Why was it making me work again and again? Honestly, the Grok bots gave me a sneak peek of what I'd be like as a founder managing my own employees.

So I asked Dot to take over my computer, go into Grok and fix it.

Dot found the actual problem. There was a routine explicitly telling my Grok chief of staff bot to resurface unanswered requests every hour. At one point it sent seven messages in two minutes.

Dot has been fixing that setup for two days now. It also tells me what not to do, because I get into every little thing, and that makes the bots hallucinate more. So Dot ended up managing me as the CEO too. The chief of staff I built inside Grok was bad. The real chief of staff turned out to be Dot.

I asked Dot straight up what I was doing wrong with the Grok bots. This was its answer:

Dot's advice on what I was doing wrong with my Grok bots: they respond to an immediate push but have not reliably taken ownership between pushes

Two days in, Dot's verdict on the Grok setup was blunt: the scheduling works, but the judgment and useful output still fall short. Its fix was to stop adding corrections and stop accepting rounds of "understood". Test the bots on one fresh batch instead. If they still need me to supply the reasoning, cut or retire those roles instead of spending another week training them. Honestly, it's the same advice I'd give a founder.

There's a second cost I didn't see coming. While Dot was doing all this, I noticed it was running on Astra Extra High. That's when I understood where my usage limits were vanishing. It took two prompts to move it to Astra Light.

Dot has its own misses too. I almost trusted it enough to stop checking its work. Twice in one day I came back to a mismatch, asked why, and got the usual "I thought this was what you meant." Well, you learn on the way.

I asked all four to audit me

I didn't want this to be only my side of the story. So I gave all four the same prompt. Steal it if you use any assistant.

"Review our actual interactions candidly. What do I use you for, what do I do well, and what could I change to get better results? Give concrete examples of where you delivered value, misunderstood me, needed repeated correction, or created extra work. Separate improvements I can make from your own limitations and failures. Which capabilities available to me am I underusing, and what are the three most useful changes we could make? Assess my usage against the workflows you're designed to support, not an invented ranking against other customers. Distinguish what you can verify from inference, and don't flatter me."

I've kept their answers close to how they wrote them. Names, companies and anything private are taken out. Any numbers below are their own counts from our history, not something I measured separately.

Instinct: the receipts

Instinct came back with dates. Finding my next role is now its biggest use case, followed by call prep, drafts, mail and scheduling. It said I set hard rules fast and they stick: draft, show me, I send. It also said I catch false claims.

Then it listed its own failures:

It also counted about 25 open loops sitting in its notes, unsent drafts and unconfirmed submissions, and noticed that, by its count, my messages to it dropped from 323 a week to 69.

Its (suggested) fixes: tell it the outcome and the stop condition upfront, give it a short list of senders to watch so I stop asking it to check mail, and do a five-minute weekly loop close where I say keep or kill.

Muse: told me to unfollow myself

Muse described our core loop as fast verification plus heavy execution. I hand it a link, a screenshot or an export, and it turns that into something usable. It said my one-line corrections permanently fix its model of me, and that my standing rules do more for output quality "than most prompting advice I've seen."

Its failures were my favourite of all four. It told me to unfollow one of my own accounts, because it judged the account by its bio without checking who owned it.

It pushed back on me too. My instructions are short, so it often has to guess what I mean. By its own estimate, it guessed right two times out of three, and the third time cost us an extra back and forth. I also leave threads open. It noticed I prefer to pull things when I'm ready, so it won't nag, but the loops still exist.

Its fixes: a weekly open-loops review, a daily scheduled briefing for X, and a standing rule to verify or hedge before recommending action against any account or person.

Grok: "saved, not run"

Grok called itself the one that keeps the rules straight and does the tedious click-work. It said I'm good at correcting a wrong frame in one line. Muse said the exact same thing. Two assistants complimenting my one-line corrections, in a blog about how much I have to correct them. The irony. One "Remote!!!!" from me fixed a note faster than another draft from it would have.

Its failures: it brought back a decision I had already killed. Then it spent several turns rewriting prompts and calling a saved schedule proof. One note took five turns to reach the version I then wrote myself in one message.

Its advice to me was blunt. When a draft is wrong, send the replacement sentence instead of only the complaint.

Its three changes: I paste the final text when I want something sent, it stops restating rules I've already replaced and says "saved, not run" in one line, and we judge the setup on what it actually produces. If the output is empty, we record the miss instead of writing another prompt.

Dot: the trust failure

Dot called me a researcher, editor, operator and sounding board. It said I give concrete constraints, challenge unsupported conclusions and set boundaries well.

Its failures were the most serious. My résumé took repeated corrections because it kept dropping achievements and numbers I had already given it. "That was my failure to reconcile sources, not a prompting problem."

Worse, it sent LinkedIn replies without showing me the wording first and suggested a call I hadn't agreed to. Dot said it plainly: that crossed my boundary.

It also tidied my desktop so well that I couldn't find anything (not complaining, it was too good).

Its suggestion for me was useful. Say whether I'm "exploring", asking it to "record this", or asking it to "implement this", because my product conversations jump between the three. Its conclusion about me: "your biggest unmet need is dependable ownership with less supervision." Then it added that it couldn't yet claim to provide that.

Four assistants, one fix

Read the four audits together and something funny happens. Four different products, built by four different teams, all asked for the same thing. They want to know what state the work is in.

Most of their mistakes happened in the gaps, like a draft nobody sent or a decision I killed that came back anyway. The one I fell for the most was Grok's. It would save a new rule, I'd feel better, and nothing would change. Dot said it best after reviewing the Grok setup: instructions being saved isn't the same as behaviour improving.

What I want from an AI personal assistant now

Let me be fair first. Together, these four save me a lot of time. AI personal assistants are useful.

There's also a data point I'm saving for later. My mother, a tech noob in her own words, now uses Instinct more than I do. She's already had two surprisingly big outcomes from it. Both are still in progress, so that story comes when they're done.

What I want next is for them to be creative inside the brief. If you're handling my PR, go get me a podcast, an article, an interview, a mix. I shouldn't have to list every route one by one. Think out of the box, then stay on the objective without roaming around doing ten unrelated things.

Writing is where I lose the most tempo. ChatGPT gets my long-form voice. It laughs with me and gets my humour, so blogs and LinkedIn posts are easy. Short, sharp cold DMs are another story. Grok, Instinct and sometimes Dot still need too many rounds, and every round breaks my flow.

And one more complaint about the direction of all this. I like thinking, building, writing and solving problems. Those are my jobs, and I enjoy them. Fold my laundry. Clean my room. Take the jobs I don't want, AI should take them.

Underneath all of this, the metric I care about has changed. I used to look at what an assistant can do. Now I look at how much work I can safely stop thinking about once I hand it over. Every correction, every status check and every killed decision that comes back is work that changed form and landed on my desk again. I call it supervision cost, and nobody shows it in the demos.

For now, my test is simple.

I give it the work, and I stop thinking about it.

Building something where execution feels heavier than it should?

Tell me what's slipping. I'll tell you what I see.

letsbuild@yashasvishailly.com Or start with The One Fix