I used up a week of Muse on my first day.
I never checked the meter. I just kept going until the app told me to stop, and that evening I paid for the $20 plan so I could keep going. That was Saturday, September 19. By Wednesday it was running twelve jobs for me on a schedule, and I’d figured out that my role wasn’t asking it questions. It was managing it.
Muse is the personal AI agent Meta launched on September 8. You talk to it like a chatbot, but it keeps one running conversation, remembers what you’ve taught it, and does work in the background on a schedule, then reports back. What I told it yesterday, it acts on today, whether I’m around or not.
TL;DR
- In July I spent four days vibe coding a deal-finding tool that never worked the way I needed. In September, Muse did the same job after half a day of direction, and I never built an app.
- A week in, it runs twelve scheduled jobs for me. The work isn’t prompting. It’s managing.
- It knows a lot, and it assumes it knows the rest. Most of my effort goes into catching the details.
- Every rule I run it by came from something going wrong once. Every run reports. It drafts, I decide. “Done” isn’t proof.
- It never touches my day job. Keeping that line is part of doing this well.
- Personal agents are the next wave, and the skill they reward is management.
I stopped building the tool
My lab for all of this is my baseball card hobby. I buy, price, and sell vintage and ’90s cards through Facebook groups, Marketplace, and eBay. There’s real money on the line, the feedback is fast, and the market is full of mislabeled listings and sharp resellers. It punishes sloppy work quickly, which makes it a good place to learn.
In July I tried to build my way to an edge. I wrote a tool to find underpriced cards: four days, dozens of revisions, a six-phase plan, a database, a dashboard. I built it with AI help, and it still never worked quite the way I needed. The half that was supposed to help me sell never got built. I shelved it.
On my first day with Muse, I described what I’d been trying to do. By that evening it was searching Facebook Marketplace for deals. Clunky at first, and it took a few rounds of refining. Then I added the private Facebook groups I buy from. Then the sell side: I update my inventory spreadsheet, tell Muse a card is ready, and it drafts the listing for Facebook, X, or eBay. The first one went live on eBay on September 23. I took the photos and Muse did the rest, right up to the button. I pressed that myself.
Same job, both times. What changed is that I stopped building and started delegating. Those are different skills, and only one of them is something I do every day at work.
What it does while I’m not looking
Twelve scheduled jobs, as of this week. Most of them hunt for cards. Some sweep hourly, and every one of them pulls recent sold prices before it calls anything a deal. One sorts my personal inbox and hands me a morning briefing. Family messages break through right away, and everything else waits its turn. One is health-related. Three more were meant to help me kick a bedtime snack habit. Those are paused, and we’ll leave it there.
What surprised me most was how little I had to manage the tool itself. I don’t start a new chat for each request. It’s one conversation that keeps building, and it remembers the rules I’ve given it. That’s the part that makes it feel less like software and more like a colleague.
It always assumes it knows
It knows a surprising amount about cards. It still gets the details wrong, and it’s confident about them.
Its first eBay description was written in the third person, like a stranger selling my card. It valued a large vintage lot around the wrong Mickey Mantle, and I had to point out which card was actually in the photo. Small things, mostly. But in this hobby, small things are the price. Get one card number wrong and the whole valuation falls apart. Business works the same way. The details an AI gets confidently wrong are the ones a customer, an auditor, or a board member notices first.
The one that got me was the inventory sheet. I asked it to add some new rows, and I assumed it would know to put them at the bottom, above the totals line. It didn’t. It told me the sheet was updated and correct. I said it wasn’t. It told me again. After a few rounds of that I decided I must be looking at an old copy of the file, and I went looking for the problem on my end.
It was so sure it was done that I started doubting my own screen.
The rows had landed on top of the totals line, and some below it. The new cells were bolded like the totals, aligned wrong, and missing the grid borders the rest of the sheet had. I had to spell out every one of those before it got it right.
In We’re All Leaders Now I wrote that AI turned everyone into the leader of a capable, confidently wrong intern. I didn’t expect to prove it on my own spreadsheet.
The rules I didn’t plan to write
I didn’t sit down and design a governance model for a baseball card hobby. Every rule below came from something going wrong once. That’s how most management rules get written, if we’re honest.
Teach once, correct fast
Early on, one of the sweeps flagged three “good flips.” I killed all three in one sentence. None of the cards were what the titles claimed. Within the hour, that became a standing rule: read the photo against the title before pricing anything. It also became a new job. If sellers mislabel cards as better than they are, some of them mislabel cards as worse, and those are the real bargains.
It runs the other direction too. I told it eBay requires signature confirmation on orders over $800. It checked and came back with $750, with a link to eBay’s own policy. No arguing, no deferring to the boss. Just the right number. I want that from anyone who works with me.
Every run reports
At first, some scheduled jobs finished quietly and I couldn’t tell whether they’d run at all. So I told it: “I don’t like silent runs. Can’t tell if it finished.” Now every job reports, even when there’s nothing to report. “Nothing qualified” is an answer. Silence isn’t.
It drafts, I decide
It researches, drafts, and prepares. It doesn’t message anyone, buy anything, or publish anything unless I’ve said yes. The eBay listing went up as a private draft first, and I was the one who listed it. That boundary is built into how I use it, not a promise it makes.
“Done” isn’t proof
The totals line taught me this, and so did a worse moment. I asked for a change one tab at a time. It had already started on every tab at once, tried to undo that, and broke rows on one of them. I only caught it because I’d asked to go one tab at a time.
Now any change to my data follows the same steps: propose it, back it up, make it, check it. And when I ask whether it actually updated something, it doesn’t say yes. It opens the file and shows me the lines that changed.
My eye is final
When I pass on a card, it stays passed unless the price drops or it gets reposted. The agent records my judgment. It doesn’t argue with it the next morning.
Now try it yourself. Ask anyone how much they’d let an agent do and the honest answer is “it depends.” Checking my email and launching a $1M campaign shouldn’t get the same answer. So this asks by stakes.
The authority dial
How much would you let it do?
The honest answer is "it depends," and what it depends on is the stakes. So answer by stakes: for each level, set how much you'd let an agent do, and whether it has to tell you every time it runs.
- I do it myself: you don't hand it to an agent at all.
- Reads only: it looks things up and tells you what it found. It changes nothing.
- Drafts, I approve: it prepares the work, and nothing happens until you say yes.
- Acts on its own: it does the job without checking with you first.
That's my setup, a week in. It reads and it drafts, and I decide. Nothing goes out without me, and every run tells me it happened, even when there's nothing to report. It gets more room one job at a time, as it earns it.
This panel updates with JavaScript on. What you're seeing is my own setup.
Runs in your browser. Nothing is saved or sent.
Why it lives on my phone, not my work laptop
Muse can see my personal email and nothing else, and everything I’ve done with it has been on my personal iPhone. It has never touched anything from my day job, and I have no plans for that to change. A consumer agent hasn’t been through my organization’s AI review, so it doesn’t go near anything that has.
I use it anyway, on purpose. I don’t want to fall behind on the tools that are coming.
I can’t help lead AI adoption at work on tools I’ve only read about.
I’ve come to think that’s what “both sides of the table” looks like in practice. You push to learn the tool, and you respect the boundary, at the same time. Neither one cancels the other.
The habits carry over, though. Every rule I run at home has a name at work:
| My rule at home | What it’s called at work |
|---|---|
| Every run reports | Audit trails and monitoring |
| It drafts, I decide | A person approves anything the public sees |
| “Done” isn’t proof | Sources and evidence attached to every output |
| A correction becomes a rule | Fixing the policy, not just the one mistake |
| My eye is final | AI proposes, a person decides |
You don’t need a platform rollout to learn this. Running one agent well at home builds the same instincts: scope what it’s allowed to do, demand proof, keep a record, and stay the one who decides.
The next wave is a management job
I think personal agents are the next wave, mostly because they keep working after you close the app. That changes what the job is. You aren’t asking questions anymore. You’re delegating, and delegation has always been a leadership skill.
The people who get the most out of these tools won’t be the best prompters. They’ll be the ones who set clear standards, correct fast, check the work, and hand over a little more authority only after it’s been earned. That’s management, and most of us have been practicing it for years without calling it AI.
I spent four days in July building a tool. In September I spent an afternoon managing one, and got further.
If you’ve set up an agent of your own, I’d like to know what you had to teach it twice.