IntegrationsBlogBook a free AI assessment
AI for accounting firms

Doing bookkeeping faster is not the opportunity. Selling what it frees up is.

Every vendor sells accounting firms the same thing: the same work, cheaper. That protects a shrinking fee. The firms pulling ahead use automation to fund a service line they can sell at three to five times the price of bookkeeping, to the clients they already have.

Where the margin moves
Compliance only52
Automated delivery74
Repackaged as advisory79
gross margin per clientmodelled
In one sentence

AI for an accounting firm has two jobs. First, cut the cost of delivering compliance work so fee compression stops eating the practice. Second, turn the freed capacity into an advisory service line the firm sells to existing clients at a materially higher price.

Your stackQuickBooks, Xero, and the tools you already run
Human gatesOn filings and payments
You own itCode and runbook
Free AI assessment

Bring one messy workflow. We will show whether an agent, automation, SaaS product, or no build is the right next move.

Find your first agent workflow
01

The two jobs, and why one alone is a losing game

Cutting cost to serve is necessary and it is not a strategy. If you automate a $650 monthly engagement and keep charging $650, you have improved margin on a fee that competitors and software will keep pushing down. Automation is the funding mechanism. Repricing is the payoff.

  • Defend: cut the cost of compliance work you already do
  • Grow: sell the freed capacity back as advisory, at a higher price
  • Do only the first and you win a race to the bottom more slowly
Build vs. buy
Buy

Use a product when the workflow is standard and the data path is simple.

Fast startLess control
Build

Build when integration, compliance, or differentiation decide the outcome.

Your stackYour code
02

What this does to the economics of a single client

Take one $650 a month bookkeeping client. Automating delivery lifts gross margin by roughly 23 points. Repackaging that same client into an AI-enabled service at $1,150 lifts revenue 77 percent and adds another 5 points of margin on top. The second move is worth more than the first.

Model C, cost to serve one client per month

LineBeforeAfter agentsRepackaged
Monthly fee$650$650$1,150
Delivery hours7.53.14.4
Labor cost at $42 loaded$315$130$185
Agent and tooling cost$0$38$52
Total cost$315$168$237
Gross margin51.5%74.2%79.4%

Assumes a $42 fully loaded delivery hour and agents absorbing categorization, document chasing, and close prep. Illustrative model based on the assumptions shown. Not a guarantee of results. Individual firm results vary.

03

What you can sell, and what it should cost

Most firms price advisory by guesswork because there is no public benchmark. This is the ladder we see work. The goal is not to invent a new product. It is to move existing clients up one rung, which is a conversation you can have without winning a single new logo.

Model D, the service tier ladder

TierWhat the client getsMonthly priceTarget margin
1. ComplianceTax and annual close$450 to $90055 to 65%
2. AI-enabled bookkeepingAutomated categorization and reconciliation, monthly package$900 to $1,60070 to 78%
3. AI controllerTier 2 plus AP/AR agents, KPI dashboard, monthly review call$2,000 to $3,80072 to 80%
4. Fractional CFOTier 3 plus scenario modeling, cash flow agents, board pack$4,500 to $8,50065 to 75%

A realistic twelve month migration target is 25 percent of Tier 1 clients to Tier 2, and 15 percent of Tier 2 to Tier 3. Illustrative model based on the assumptions shown. Not a guarantee of results. Individual firm results vary.

04

The hours have to go somewhere, and you only get to spend them once

This is where most AI business cases quietly cheat. A nine person firm might free around 857 hours in year one. Those hours can become billable advisory work, or they can avoid a seasonal hire. They cannot do both. Any vendor who adds the two together and calls it total ROI is selling you a number, not a plan.

Model B, capacity reclaim for a nine person firm

LineValue
Firm-wide hours per year on categorization, chasing, cleanup, 1099s, close prep4,100
Agent-eligible share38%, or 1,558 hours
Realistic year-one capture55%, or 857 hours
Path 1: redeploy to billable advisory857 hrs x 70% conversion x $165 realized = $98,983
Path 2: avoid a seasonal hireOne staff accountant, fully loaded = $72,000

Path 1 and Path 2 are alternatives. Adding them together is the most common credibility failure in AI ROI marketing, and it is why most of these numbers should be read skeptically, including ours. Illustrative model based on the assumptions shown. Not a guarantee of results. Individual firm results vary.

05

What actually gets built

Not a chatbot bolted onto QuickBooks. Supervised agents that run inside your existing stack, take real actions in the ledger, and stop and ask a human when the situation is unclear. Every action is logged, because you will be asked to explain it during a review.

  • Runs in your cloud, against your data, under your access controls
  • Human approval gates on anything that touches a filing or a payment
  • Full audit trail of every action, retrievable during review
  • You own the code and can operate it without us
Deploy target
OpenAIOpenAIClaudeClaudeGeminiGeminiSalesforceSalesforceSnowflakeSnowflakePostgresPostgres
SSORBACAudit logCloud
06

The uncomfortable part of the twenty-four month picture

Firms that do this properly end up with fewer clients, not more. Capacity moves to the engagements that carry margin, and the bottom of the client list gets released or repriced. If a plan promises more revenue, higher margin, and a growing client count all at once, it has not been thought through.

Model E, firm-level trajectory for a $2.1M practice

MeasureBaselineMonth 12Month 24
Revenue$2.10M$2.51M$2.98M
Gross margin54%61%67%
Revenue per FTE$150K$179K$199K
Advisory share of revenue12%27%41%
Client count240236228

The falling client count is deliberate, not an error. Illustrative model based on the assumptions shown. Not a guarantee of results. Individual firm results vary.

07

The four things being sold as an "AI accounting assistant"

This page is scoped to tools that touch the ledger, the bank feed and the workpaper. Research tools like Checkpoint Edge are deliberately out of scope here.

Inside that scope, four different products share one label. A bank feed classifier that suggests a category. A document extractor that reads a W-2 into a field. A drafting tool that writes the client email. And an agent that executes a step in your close.

Ask which one you are looking at before the demo starts. The first two touch the ledger and the workpaper. The third touches the client's inbox. Only the fourth does work without a person pressing a key.

The categories fail differently, so they need different gates. A bad category suggestion is caught in review. A bad email is caught by the sender. A bad agent action posts.

Firms buy across all four without noticing. That is how a practice ends up with three overlapping subscriptions and no measured saving.

08

How do you tell a rules engine from an agent during the demo?

Bring a de-identified export from one of your messiest clients. Ask the rep to run it live, in front of you. A sandbox tuned by the vendor tells you nothing.

A rules engine matches a string and returns a category. An agent reads context, picks a treatment, and shows its reasoning. Then ask what happens on the second occurrence.

Rules repeat identically forever. A model that scores confidence behaves differently once a pattern builds in the client's own history.

Intuit describes its top confidence tier as exactly that kind of history match. It says QuickBooks has strong data behind the suggestion, with a clear pattern in your history.

That is pattern matching, honestly labeled. It also explains why a client onboarded in November gets weak suggestions on their first bank feed. There is no history to match against yet.

09

Does the tool tell you when it is guessing?

Some do, and it is the single most useful feature in the category. QuickBooks grades its own suggestions. The lowest grade tells the user plainly that QuickBooks has limited data behind the suggestion.

Put this on your evaluation sheet. A tool that returns one answer with no confidence signal gives the reviewer nothing to triage by. Every line then gets the same attention, which is the same as no automation.

Confidence should sit at field level, not document level. A K-1 can be read correctly in eleven places and wrong in one.

Ask to see the low confidence queue during the demo, not the happy path. A vendor that cannot show you its own failure queue has not built one.

010

Where do bank rules actually run out?

At 2,000 rules per file, and five conditions per rule. Intuit publishes the ceiling directly, stating you can create up to 2,000 bank rules.

The second limit bites first. A single rule is capped, because you can set a single rule with up to 5 conditions.

Five conditions run out faster than you expect. This vendor, above a dollar threshold, only on the operating account, only when the description carries a second string. You are already at the ceiling.

Then there is the maintenance problem. Rules live inside one company file. They do not travel between clients except by copying them. A 40 client book means 40 rule sets to build and keep current.

Use the numbers in your own diagnosis. Count rules on your three messiest clients before you price any tool.

011

How much of a 1040 document set can software finish on its own?

Sixty five percent of standard documents, by the incumbent's own published figure. Thomson Reuters states that 1040SCAN eliminates the need to verify OCR data for 65% of standard documents by combining patented text-layer matching and AI.

Read the other half of that sentence. Roughly a third of standard documents still go to a person. Non standard documents are not in the denominator at all.

That remainder has a shape. SurePrep's review step color codes fields. It marks a field where OCR is uncertain about that field and requires verification.

Use 65% as your baseline when a newer vendor quotes you a number. Ask what their denominator is. Ask whether it includes the documents that arrive sideways in March.

012

What has to be settled with the vendor before any client file moves?

Start with what the vendor is. Under the 7216 regulations it is very likely a tax return preparer itself. The definition reaches a person who develops software that is used to prepare or file a tax return. Vendor selection is not an IT purchase. It admits a second preparer into the client's return data.

Then ask where the software actually runs. If anyone receiving the return information sits abroad, the rule is explicit. It requires the taxpayer's consent under § 301.7216-3 prior to any disclosure. Remote viewing counts as disclosure. "The data never leaves our tenant" does not answer the question.

Then read the contract. The Safeguards Rule reaches accounting firms directly, and diligence alone does not satisfy it. The rule obliges you to go further, requiring your service providers by contract to implement and maintain such safeguards. Encryption in transit and at rest belongs in the same clause. So does multi factor authentication.

So three questions cover the demo. Where does inference run, where does support sit, and what does the contract commit the vendor to. Answers given verbally do not count. For the full treatment of 7216 consent, client disclosure and reviewer duties, see our AI agents for accounting firms page.

013

The questions to put to a vendor in writing

There is a published list, and it came from the profession rather than a marketing page.

Clients want to know how their information is stored; whether their data is used to further train the generative AI model; who has access; whether their data is shared outside of your organization; how data is retained, by both the firm and any external provider(s); and whether the generative AI has de-identification procedures for stored information.

Send those six to the vendor before the second call. Add the three from the section above, plus one more. What does the published document scope exclude?

Some vendors already answer the training question on the record. Karbon states that it does not use your firm's data to train AI models, and no data is shared with third parties or used to inform any other Karbon customer's experience.

Ask for firm level controls too. Karbon notes that account administrators can disable AI functionality through firm settings. That is the switch a partner needs when a client objects.

014

What a 97% accuracy claim does not tell you

It does not tell you the denominator. Xero's marketing page says Xero reconciles your transactions at 97% accuracy, so there's less manual matching and more time for your business. No methodology is published on the page.

Treat every accuracy figure in this category as a vendor claim until a method appears. Nobody publishes the document mix, the sample size, or what counted as an error.

If that 97% held on your files, a client with 900 monthly transactions would still leave about 27 lines for someone to find. Xero publishes no test set. Treat that as arithmetic on a marketing number, not a workload estimate. The real figure could be better, or much worse.

Ask three questions instead of comparing percentages. What was the test set, who labeled the correct answer, and does the figure include the documents the tool refuses.

015

Does automation cut review time, or just move it?

It moves it, and the evidence comes from standard setters rather than vendors. One note on scope first. The PCAOB material below governs audits of issuers. A firm doing private company work sits under AICPA standards instead. The pattern still transfers.

It is about what happens after a tool flags something. The PCAOB observed that because technology-assisted analysis may enable the auditor to examine all items in a population, it is possible that the analysis may return dozens or even hundreds of items within the population that meet one or more criteria established by the auditor.

The IAASB reaches the same place from the other direction. In its worked example, the procedures performed using ATT do not provide sufficiently persuasive audit evidence for Group A, rather it further informs the auditor's risk assessment.

So budget reviewer hours, not just preparer savings. A firm that cuts prep by a third and adds an exception queue nobody owns has traded cheap time for expensive time.

The obligation does not shrink either. The IAASB puts it in one line. Regardless of the tools and techniques used, the auditor is required to comply with the ISAs.

016

What a pilot should look like before busy season

Pick one workflow, one client segment, and a closed prior period. Run the tool over books you already closed. Then you have the right answer to compare against.

Choose work where a mistake is visible immediately. Document sorting, extraction and duplicate detection flag things. Coding and accepting bank feed matches post to the ledger, so those need a person on the gate. Reconciliation belongs with the review side.

Baseline three numbers before anything is switched on. Prep hours per return type, reviewer hours per file, and the count of follow up loops back to the client's inbox.

Name one owner who knows the chart of accounts, usually a manager rather than a partner. Start after the spring deadline and finish testing by late summer. Decide in September, then leave Q4 for training the people who will use it.

017

How you will know the pilot failed

Reviewer hours went up and nobody noticed. That is the most common failure. It is invisible unless you baselined the reviewer, not just the preparer.

Watch for a growing queue of files the tool will not take. Vendors publish scope limits, including orientation requirements and excluded account types. Those exclusions arrive as a pile at the worst moment.

Watch for silent coding to a catch all account. A close that looks clean because unknown merchants were swept into one line is worse than an obvious exception list.

And watch the measurement itself. In a survey of 1,073 firms, 40% said they had not yet figured out how to track efficiencies due to technological advancements. Most pilots end in opinion rather than evidence.

018

Buy the tool or build the agent: the questions that decide it

Buy where a product already fits the standard path. Indexed 1040 binders, extraction from recognized forms and tax software integration are solved. Rebuilding them wastes money.

Build where the work is yours and the data must stay inside your boundary. Odd business returns, firm specific leadsheets, client chase loops, and anything where offshore hosting would trigger a consent requirement.

Most practices have no internal build capacity today. The 2024 CPA.com and AICPA PCPS CAS Benchmark Survey drew 206 self-selected respondents. Only 13% build automation with an internal team, while more than 67% are partnering with software vendors to provide these kinds of tools.

The report concedes self-selection bias, so read it as directional. The stack you are automating is messier than the demo assumes. In the same survey, only 46% of respondents report using a specific set of software applications that are fully integrated.

Implementation is also unpriced at most firms. Only 55% charge separately for client technology setup, and only 42% have a team that does it, per the same benchmark.

019

Where firms like yours have actually put AI so far

At the edges of the work, not the middle. A 2026 survey of 486 bookkeeping and accounting professionals published a task table. The figures below sit on an AI-active base rather than the full sample.

Those two bases are not the same, which matters on a page about denominators. The reported rows run like this:

- Drafting client emails and communication, 75% - Summarizing documents or meetings, 71% - Research and answering technical questions, 69% - Transaction categorization or coding, 42% - Financial reporting and commentary, 40% - Bank reconciliation, 24% - Tax return preparation, 9%

Read the bottom of that list before the top. Bank reconciliation at 24% and return preparation at 9% tell you where the technology has not earned trust yet.

Trust is the limit, not availability. In the same survey, only 19% trust AI enough to use it with limited review.

Almost nobody wants the tool acting alone. In a separate Intuit commissioned survey of 725 US accounting professionals, only 6% want AI to execute autonomously. That survey is vendor funded, and its sampling method is published on the page.

The shape firms do want is draft plus review. Asked what role AI should play in client work, 40% want AI as a support tool and 34% want AI to draft, with a human reviewing and signing off.

021

What to leave in the client file

Three things, and none of them are hard. The same Aon column recommends that employees document the prompts used, how the outputs were verified, and who performed the review.

Store the named human approver, never a service account. If the log shows software approving software, the firm has automated away its own evidence.

Insurers are already asking to see the policy behind it.

Aon's risk control lead says carriers are going to ask a firm about their AI policy and procedures. They want to see firms are approaching the use of AI with the same basic risk management protocols that they would have in place for engagement letters or client acceptance and continuation or documentation.

Claims data does not exist yet, and that is worth saying plainly. The same source notes there really hasn't been a lot of claims or large dollar amounts paid on claims. Nobody can price this risk from experience.

022

The real blocker is time, not resistance from staff

Partners expect a fight from the team and get a calendar problem instead. In a survey of 1,073 firms, the biggest barrier to implementing emerging technologies was lack of time to explore or implement (41%). Staff resistance or fear of change registered at only 6%.

The tax side reports the same two constraints. In a separate survey of 639 tax and accounting professionals, almost half (47%) of respondents said the biggest reason they can't (or don't) pursue more automation is lack of time and resources, followed closely by the cost of implementation (45%).

Most firms are also less automated than the marketing suggests. In that same survey, about half (49%) of the respondents to this year's survey estimate that one-quarter of their tax workflows are automated, while 21% said that up to half are automated. And 18% said they use no automation at all.

So the practical answer is to scope small and staff it properly. A named owner with real hours beats a firm wide rollout nobody has time to run.

Where it pays off

Concrete places agents earn their keep.

01
ticket82% resolved
#4821Damaged ordernew
Agent

Policy matched. Refund ready for approval.

Lookup orderApprove refund
human-gated

Bank reconciliation

Match across feeds and the ledger, surface only the exceptions a person needs to judge.

02
ledger31 hrs saved
Stripe$18,240matched
Bank$18,240clear
audit-ready

Month-end close prep

Assemble the close package, chase the missing documents, flag what does not tie.

03
pipeline+18% coverage
LeadFitBrief
91

account score

CRM updated
crm synced

Client cleanup and onboarding

The wedge offer. Work through a messy back file fast enough to quote it as a fixed fee.

04
reviewHIPAA path
Credentialing packet3 checks passed
Human review required
review queue

AP and AR chasing

The follow-up nobody has time for, run on schedule with a human on approvals.

05
extract14 fields
Invoice no.TotalDue date
2 exceptions routed
exceptions out

Cash flow forecasting

The sellable advisory product: rolling 30, 60 and 90 day projections per client.

06
answerfresh docs
Answer drafted3 cited sources
HR policyOkta SOP
sources shown

1099 and filing prep

Seasonal volume absorbed without seasonal hiring.

FAQ

Common questions.

How much revenue can an accounting firm add with AI advisory services?+
A fourteen person firm converting 18 of its 240 clients to a $2,200 a month AI-enabled controller service adds about $475,000 of gross revenue and roughly $350,000 of gross profit a year. Modelled on a 30 percent attach rate among suitable clients, not a measured client result.
What should a ten person firm automate first?+
Bank reconciliation and client cleanup. Reconciliation is high volume, rule-heavy and easy to supervise, so it pays back fastest. Cleanup is the offer you can sell immediately, because prospects already know their books are a mess and will pay to fix it.
Is it safe to give an AI system access to client financial data?+
Only under conditions you control. Agents should run inside your own cloud or a dedicated tenant, use your access controls, log every action for review, and require human approval before anything touching a filing or a payment. Ask any vendor whether your client data trains their models.
Will AI replace accountants and bookkeepers?+
It replaces the categorization, chasing and reconciliation work that partners should not be doing anyway. Judgment, client relationships, planning and anything you sign your name to remain human. The firms that shrink are the ones that automate and keep selling the same low-margin compliance package.
AI or offshore staffing, which is cheaper for a small firm?+
Offshore staffing has a lower entry cost and scales linearly: twice the volume needs twice the people. Automation costs more up front and then scales cheaply. Below roughly 4,000 eligible hours a year, offshore usually wins on cost alone. Above it, the economics invert.
How long before a firm sees anything?+
A first working build on one scoped workflow can land in days. Meaningful margin change follows the client migration, not the build, so expect the picture to look like the ramp above: negative for two quarters, crossing over around month eight.
Production AI agents, shipped with an owner

Want agents like these in your stack?

Book a free assessment, we'll map where an AI agent creates real leverage in your workflows and scope the first one to ship.

Build, deploy, runYour cloudYou own the code