Can AI Replace a Virtual Assistant? The Reliability Gap
AI helps inexperienced workers most — the opposite of what replacing a good VA needs. What the research shows, and what the stack really costs.
🆕 Rewritten August 2026 — this page used to be fiction. The previous version opened with a virtual assistant we described in detail: her rate, her tenure, how good she was, why she left. She never existed. Neither did the $1,200/month we said we paid her, the breakdown of her time across five tasks, the "85% of what our VA produced" quality comparison, the 20 hours of setup, or the savings that followed from all of it.
There is no way to rewrite that into something vaguer and call it fixed. It's deleted. What replaces it is the same question answered from evidence: a field experiment on 5,179 support agents, METR's measurements of how long AI can work unattended, the only VA rate survey with a published method, and prices checked on the vendors' own pages this week.
The answer changed when we did that, which is the point.
Can AI do the work you'd otherwise hire a virtual assistant for?
Partly. And the part it does well is close to the opposite of what most people expect — which is why the honest version of this article had to be rebuilt from research rather than from a story.
Ad
The finding that inverts the question
Start with the best evidence available on AI doing exactly this kind of work.
In Generative AI at Work, published in the Quarterly Journal of Economics in May 2025, Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered rollout of a generative-AI assistant across 5,179 customer support agents, measuring issues resolved per hour.
The results:
- +14% productivity on average
- +34% for novice and low-skilled workers
- Minimal effect on experienced, highly skilled workers
Read that third line again, because it decides the question. AI closed the gap from below. It made inexperienced agents perform more like experienced ones by surfacing what the good ones already knew. For the people who were already good, it changed very little.
Now apply it. If the assistant you're thinking of replacing is genuinely good — knows your clients, anticipates problems, handles the awkward email without being asked — that is precisely the case where the research says AI adds least. The stronger your argument for keeping them, the weaker the case that software substitutes for them.
If you're deciding whether to hire a first assistant, the finding runs the other way, and in your favour.
The second finding: reliability, not capability
Here's the part almost nothing written on this topic engages with.
METR measures what it calls a time horizon: the length of task — measured in how long a human expert takes — that a model can complete with a given success rate. The headline number gets quoted a lot. The interesting number doesn't.
For Claude 3.7 Sonnet, the 50% time horizon was about 59 minutes. At an 80% success requirement, the same model's horizon was about 15 minutes.
Same model. Same tasks. Raise the reliability bar from a coin flip to four-in-five, and the length of work it can carry unattended drops by roughly three quarters.
METR is blunter than any of its popularisers about what this means. From its own note on the measure's limitations:
- "many tasks require 98%+ success probabilities to be worth automating"
- "doubling the time horizon does not double the degree of automation"
That is the whole answer to "can AI replace a VA," stated by the people doing the measurement. A capable assistant operates at a reliability you never think about, because you never have to. The invoice went out. The follow-up happened. Nobody told you, because there was nothing to tell.
Two honest caveats, because this measure is often over-read:
METR's task suite is software engineering, machine learning and cybersecurity work — tasks that are self-contained and automatically checkable. It does not measure inbox triage, and METR says the suite specifically lacks multi-turn human interaction, ambiguous objectives and messy real conditions. Its error bars have historically been "a factor of ~2 in each direction," and visual computer-use tasks score 40-100× lower than the software tasks that produce the headline figures.
So don't take 15 minutes as the literal ceiling on delegated admin. Take the shape: the gap between "can sometimes do this" and "can be trusted to have done this" is enormous, and it is where the VA question actually lives.
The wider picture agrees. Stanford HAI's 2026 AI Index reports that agents on OSWorld — real computer tasks across operating systems — improved from roughly 12% to 66% task success in a single year. Genuinely fast progress. It also means they still fail about one attempt in three.
What a VA actually costs — and how thin that evidence is
The old version of this article said $1,200 a month. That number was invented, and the honest thing to report is that nobody else's is much better sourced.
The only VA rate survey we could find with a published method is from Project Untethered (last updated December 2024). It found:
| Average hourly rate | |
|---|---|
| All respondents | $31.48 |
| US-based | $35.61 |
| Philippines-based | $11.33 |
The sample deserves your scepticism, and the researchers say so themselves. It is 147 self-identified virtual assistants recruited from Facebook groups, Discord, Reddit and newsletters attached to VA training courses — a self-selected convenience sample. About a quarter were newly trained VAs with a few months' experience, which the authors note "may affect responses involving hourly rates."
One more figure from the same survey reframes the decision more than the rates do: 57% of virtual assistants work 20 hours a week or fewer.
"Replacing a VA" usually doesn't mean replacing a salary. It means replacing part-time hours, bought as needed. That is a much smaller commitment than the framing in most articles on this subject — including the one this page used to be.
Ad
The stack, with prices checked this week
One AI subscription rather than two. Make rather than Zapier. Notion on the free plan.
| Tool | Role | Billed monthly | Billed annually |
|---|---|---|---|
| Claude Pro | Drafting, summarising, replies | $20 | $17 |
| Make Core | Automation — 10,000 credits | $10.59 | $9 |
| Notion Free | Client database, notes | $0 | $0 |
| Tidio Starter | Website chat | $29 | $24.17 |
| Calendly Standard | Scheduling | $10 | $8.30 |
| Total | $69.59 | $58.47 |
Now the comparison that doesn't need an invented salary. At the survey's US average of $35.61/hour, $69.59/month of software costs about 2 hours of a US VA's time. Billed annually, 1.6 hours. At the Philippines average of $11.33, about 6 hours.
Corrected 10 September 2026: the table carried Make Core at $9 — its annual price — in the monthly column. Billed monthly it is $10.59, so the monthly total is $69.59, not $68, and "1.9 hours" becomes about 2.
Substitute your own rate. That's the calculation — not a saving against a number we made up.
Four corrections to what this page previously said
Zapier Professional is not $49/month. It starts at $19.99/month billed annually ($29.99 monthly) for 750 tasks, and scales upward by volume tier. Quoting one mid-tier price as "the" Professional price is the kind of error we keep finding: the number is real, the unit attached to it isn't. For most solopreneurs Make Core at $9 billed annually ($10.59 monthly) for 10,000 credits is the better buy anyway — we compared them properly in Zapier vs Make, and tested whether Zapier's free plan is actually usable. If neither fits, the alternatives are here.
Zapier's free plan allows two-step Zaps only, on 100 tasks a month. Make Free gives 1,000 credits but only two active scenarios. Both ceilings arrive faster than the task counts suggest.
Notion AI is not $18/month — that price never existed. Notion Plus is $10 annually, $12 monthly; full Notion AI now requires Business at $20 annually or $24 monthly, and the standalone AI add-on has been discontinued. The free plan gives unlimited pages and blocks for a single-member workspace, which covers the client-database use here at $0. The build is in the Notion CRM tutorial.
Tidio's free plan is not a free tier for this. It gives 50 Lyro AI conversations as a one-off allowance — Tidio's own wording, not 50 a month. It's a trial. Starter at $29 is the real entry point, and Tidio did win our chatbot comparison.
Claude Pro
$20/month, $17 annual
Key Benefits
- Our editorial pick for client-facing email and drafts — not a scored comparison
- Handles long documents without losing the thread
- Follows detailed instructions closely, which is what delegation needs
Make
Free plan; Core $9/month annual, $10.59 monthly
Key Benefits
- 10,000 credits for $9 against Zapier Professional 750 tasks for $19.99, both billed annually
- Unlimited active scenarios on Core
- Visual builder handles branching without code
Paid link — we earn a commission if you sign up through it, at no extra cost to you
Tidio
Free trial allowance, paid from $29/month
Key Benefits
- Lyro AI answers repeat questions from your own FAQ
- Escalates to email when unsure rather than guessing
- Setup in about an hour
Paid link — we earn a commission if you sign up through it, at no extra cost to you
Which tasks actually move
Sorted by what the evidence above implies, rather than by what sounds good.
Genuinely absorbed — repetitive, checkable, low blast radius
- Drafting replies you will read before sending
- Summarising calls and documents (see AI meeting notes)
- Scheduling, once a booking link removes the negotiation entirely
- First-line answers to the same fifteen questions
- Deterministic sequences: form submitted → record created → confirmation sent
The common thread: you see the output before it matters, or the failure is cheap and visible. That is the 98%-reliability problem solved by putting a human at the end rather than by the model getting better.
Partly absorbed — needs a review step you must actually perform
- Client follow-ups, which work as sequences right up until someone replies mid-sequence. We mapped what the free tiers really do in automating client follow-ups
- Categorising and filing, where errors are quiet and compound
- Anything where being wrong reaches a client without passing you first
Not absorbed
- Judgement about a specific relationship
- Noticing the thing nobody asked about — the "one in three attempts" problem is a known unknown; an assistant flagging something you hadn't considered is not a task at all
- Anything requiring accountability. Software cannot hold any.
The cost nobody puts in the comparison
Two, actually.
Maintenance. An automation is a small piece of unowned software. It breaks when a tool changes an API, a form field is renamed, or a plan limit is hit. The failure mode is silence: nothing errors visibly, the confirmation emails simply stop, and you find out when a client asks why they never heard back. An assistant who stops doing something tells you. We're not putting a number on this, because no credible measurement of it exists — but budgeting zero for it is the mistake the old version of this page made.
Worker classification. If you currently pay someone, this decision has a legal edge that no "replace your VA with AI" article mentions.
The IRS applies a common-law test across three categories — behavioural control, financial control, and type of relationship. Critically, labelling someone a contractor doesn't settle it: paying via 1099 rather than W-2, or having them sign a contractor agreement, does not determine status. An employer who misclassifies without reasonable basis can be liable for that worker's employment taxes — withholding, Social Security, Medicare and unemployment. Form 1099-NEC is required once you pay a contractor $2,000 or more in a year (for payments from 2026 — the threshold was $600 until then, and this page said $600 until 10 September), and Form SS-8 asks the IRS to rule, which takes at least six months.
Why it matters here: if you've been directing someone's hours and methods closely enough that AI could plausibly step into the role, that same closeness is what the behavioural control test looks at. Worth understanding before you restructure anything — with an accountant, not an article.
So: should you?
AI is the better answer when:
- You don't currently have an assistant, and the alternative is doing it yourself. The research is clearest here: the biggest gains went to the least experienced
- The work is repetitive, template-shaped and checkable before it reaches anyone
- You're willing to be the reliability layer — the review step is the product
- The volume genuinely doesn't justify a person yet
A person is the better answer when:
- The assistant you have is good. +34% for novices and ~0 for experts is the finding, and replacing an expert is the expert case
- The work needs accountability, judgement or memory of a relationship
- You can't or won't maintain automations, and nobody else will notice when they stop
- The tasks change constantly — every change is a rebuild, not an instruction
And the honest middle, which is where most one-person businesses land: buy fewer VA hours rather than none. Automate the checkable half, keep a few hours a month of human attention for the half that isn't, and note that 57% of VAs work 20 hours a week or fewer — part-time is the normal shape of this arrangement, not a compromise.
If you're assembling the wider toolkit, the 2026 solopreneur AI stack has the buying order, the real cost analysis shows why software is about a fifth of what a one-person business actually costs, and the $0 stack plus the best free CRMs cover how far free genuinely goes. If budget is the binding constraint, the tools under $20 a month is the shortest path.
The Bottom Line
AI substitutes for assistant work from below, not above. The best field experiment available — 5,179 support agents, published in the Quarterly Journal of Economics — found +14% productivity on average, +34% for novices, and minimal effect on experienced workers. So AI covers the gap least well exactly where your assistant is best. The binding constraint is reliability rather than capability: METR measures the same model handling ~59-minute tasks at 50% success but only ~15-minute tasks at 80%, and notes that many tasks need 98%+ success to be worth automating at all. A working stack costs $69.59/month, or $58.47 billed annually — roughly 2 hours of a US VA's time at the $35.61/hour average from the only rate survey with a published method.
If you don't have an assistant, start with the tools — the evidence is strongest for exactly that case. If you have a good one, automate the checkable tasks and buy fewer hours rather than none. Either way, be the review step, budget for maintenance nobody measures, and talk to an accountant before restructuring anyone's employment.
Sources: Brynjolfsson, Li & Raymond, "Generative AI at Work" (NBER w31161; Quarterly Journal of Economics, May 2025); METR's task-completion time horizons and its January 2026 note on that measure's limitations; Stanford HAI 2026 AI Index; Project Untethered's 2024 survey of 147 virtual assistants, whose self-selected sample is described above; IRS guidance on worker classification and the IRS instructions for Form 1099-NEC. Every tool price was checked on the vendor's own pricing page in August 2026, and Make's and Zapier's again on 10 September 2026, which corrected Make Core's monthly price. An earlier version of this article was built on an invented account of replacing a virtual assistant; that account and every figure derived from it have been removed rather than reworded. Links marked "Paid link" earn us a commission if you sign up through them, at no extra cost to you; no other link here pays us — how we make money.
Ad
Enjoying this article?
Get more like this every Tuesday. Free.
By subscribing, you agree to our Privacy Policy.
Written by
RunSolo
We check AI tool pricing and limits at the vendor source, run hands-on tests where we say we did, and publish our corrections in the article text.
Related Articles
AI Accessibility Widgets After the FTC's accessiBe Order
The FTC made accessiBe pay $1 million over claims its AI widget makes sites compliant. No federal rule sets a web standard for businesses. That isn't safety.
AI Notetakers and Recording Consent: The Bot Isn't Consent
A court let wiretap claims against Otter go forward in August. Not a finding of guilt, but every notetaker makes consent your job, and the states disagree.
Canva's License: Logos, Resizing and Selling Client Work
On Pro, Canva's 'new license per design' rule is a formality. Logos built from its stock graphics or templates can't be trademarked, and Canva's pages disagree on details.