AI agents for business, explained and put to work.
What an AI agent actually is, when it beats a chatbot or SaaS tool, where it pays off, and how to get one from workflow document to supervised production.
Read ticket, check policy, look up order, decide next step.
Uses tools with approval gates where risk matters.
Customer notified, CRM updated, trace saved.
An AI agent is software that uses a large language model to plan and take multi-step actions toward a business goal, calling tools, reading your data, writing back to your systems, and escalating to a human when risk or uncertainty requires it.
Bring one messy workflow. We will show whether an agent, automation, SaaS product, or no build is the right next move.
Agent vs. chatbot vs. automation
A chatbot answers. Rule-based automation follows a fixed script. An agent decides: it reads the situation, chooses the next step, calls the right tool, and adapts when reality doesn’t match the happy path, escalating to a human when it should.
- Chatbot: responds in one turn
- RPA: fixed rules, breaks on exceptions
- Agent: plans, acts, and adapts across steps
Customer says the order arrived damaged and asks for a refund.
Source: ZendeskWhy now
Models crossed the threshold where multi-step tool use is reliable enough for real work, and standards like MCP make it practical to connect agents to your systems. The constraint is no longer the model. It is integration, evaluation, and getting from pilot to production.
- Reliable tool-calling and reasoning
- MCP plus APIs make integration tractable
- The bottleneck is deployment, not capability
1. Retrieved customer and order history
2. Matched refund policy with citations
3. Requested approval before issuing refund
4. Wrote outcome back to Zendesk
How an agent ships into production
Scope the workflow from your existing documents, build with evals and guardrails, connect the systems of record, test in a sandbox, then launch with human approval gates. This unglamorous middle is why many pilots never become operating systems.
- Workflow map before code
- Evals and guardrails before users
- Sandbox first, supervised production second
Inputs, systems, owners
Tools, prompts, permissions
Known cases and edge cases
Approvals, traces, rollback
What an AI agent does that a chatbot does not
A chatbot returns text. An agent takes actions inside your systems and then reports what it changed.
The difference is tool calling. The model is handed a list of functions it may invoke, each with a name, a description and a schema. It picks which one to call, reads the result, and decides the next step. Anthropic's documentation describes the `tools` parameter as carrying "tool names, descriptions, and schemas" on every request.
So the useful question in a scoping call is not "how smart is the model". It is "which four systems does this workflow touch, and what is the agent allowed to write to".
How an agent differs from RPA
RPA repeats a fixed path. An agent reads the situation each time and chooses a path, which is why it survives inputs the script writer never saw.
That flexibility is also the cost. A script that breaks is obvious. An agent that takes a wrong but plausible action is not, unless you gated it. Everything below about approvals, logging and metrics exists because of that single difference.
What it costs to run an AI agent
Inference is cheap and published. As of September 2026, Claude Sonnet 5 lists at $2 per million input tokens and $10 per million output tokens. Check the page before you build a case on it, because list prices move.
Anthropic publishes a worked example for support. At roughly 3,700 tokens per conversation on Claude Haiku 4.5, 10,000 tickets come to "~$37.00 per 10,000 tickets".
Read that number before you budget. For most workflows the model bill is not the expensive part. The integration, the review time and the person who owns the agent are.
Why a sprawling agent costs more per run than a scoped one
Tool definitions are billed as input tokens. Every tool name, description and schema rides along on every call, plus a system prompt surcharge for enabling tool use at all.
Anthropic states the extra tokens come from "the `tools` parameter in API requests (tool names, descriptions, and schemas)". An agent with forty tools pays for forty tools on every turn, including the turns that use one.
This is the practical argument for building per workflow rather than building one agent that does everything.
How to keep the run rate under control
Two published levers, both vendor documented, not promises. The Batch API and prompt caching.
The Batch API carries "a 50% discount on both input and output tokens" for work that does not need an answer this second. A prompt cache hit is billed at 10 percent of the standard input price, which matters when every run resends the same policy document or schema.
Ask any vendor which of the two your workload can use. If neither applies, the bill scales linearly with volume.
What it costs to build one
No published source measures this. We looked for an analyst study, a vendor document or a government figure giving a typical build cost or time to production for a custom agent. There is none.
Every number on the search results page for this topic is a platform list price, not a build cost. Treat any "average cost to build an AI agent" figure as unsourced until someone shows you the sample.
What we can do is publish ours. For accounting firms, Gaper puts the bands on the page: one production workflow runs $12,000 to $34,000 to build and $690 to $2,630 a month to run. That monthly figure counts the licensed person reviewing what the agent posts, not software alone.
Those are Gaper's own numbers for one sector, modelled from our delivery costs. They are not an industry average. A workflow somewhere else prices differently, and we would rather say so than publish a figure with no sample behind it.
On timing, one commitment and its limit. Gaper delivers a first working build inside 24 hours, for a single scoped workflow on a stack we already support. That is something you can run and criticise. It is not an integrated, gated, production agent.
The integrations, the approval boundaries, the evaluations and the quota planning come after it, and they are measured in weeks. Most of this page is about that second part, because that is the part agent projects fail at.
Build or buy, by the numbers
Platform pricing is per conversation. Raw model pricing is per token. At volume the gap is large enough to decide the question.
As of September 2026, Salesforce publishes Agentforce at "$2 USD/Per conversation", with Flex Credits at $500 per 100,000 and a $125 per user per month add-on. Ten thousand conversations is $20,000 at that rate.
Name the model on the other side, because it changes the size of the gap. Anthropic's $37 example runs on Claude Haiku 4.5. The same volume on Claude Sonnet 5, at twice Haiku's published input and output rates, is about $74 by arithmetic rather than by published example.
So the gap is roughly 540 times on the cheap tier and roughly 270 times on the mid tier. Either way there is room for a build inside it. Price the build and see whether it fits.
That comparison is not fair in one direction. The platform price includes the product, the integrations, the console and the support contract. The token price includes none of that.
The honest reading is that a build pays back at volume, and on workflows nobody sells a product for.
When buying software is the better call
Buy when a product already covers the workflow end to end and your data fits its model. Buy when the workflow is not a differentiator and you would be rebuilding a commodity.
Build when the agent has to write into systems the vendor does not integrate with. Build when residency or audit requirements rule out sending data out. Build when the workflow is specific to how your company operates.
A competitor on this page puts it well: "An agent that excels at qualifying inbound leads is built differently than one that triages support tickets or processes contracts."
If a product fits, we will tell you to buy it. A build you do not need is the most expensive thing on this page.
Where the agent runs, and who owns it
It runs in your cloud account, under your authentication, and you keep the code. That is a configuration choice you can verify, not a trust exercise.
Gaper deploys into the client's own Azure or AWS account. Those are the two clouds quoted on this page, so each commitment below is one you can check on your own subscription. For Google Cloud or on-prem, ask for the equivalent documentation page before you sign.
Azure is the worked example here. Microsoft's documentation states that stored agent data "is stored at rest in the Foundry resource in the customer's Azure tenant, within the same geography as the resource". The customer can delete it at any time.
One caveat worth naming. Residency has a setting attached. Prompts stay in your geography "unless you are using a Global or DataZone deployment type", so check the deployment type, not just the region.
Does our data train someone else's model?
On a named enterprise cloud, no, and the vendor says so in writing. Read it from the documentation rather than from a sales deck.
Microsoft states that customer prompts, completions, embeddings and training data "are NOT available to other customers", "are NOT available to OpenAI or other providers of Models sold by Azure". The same page says they are not used to improve those models.
There is an exception, and you should hear it here rather than find it later. Abuse monitoring is a separate retention and review path. When the system flags indicators of abuse, "a sample of customer's prompts and completions may be selected for review". Automated review runs first, with authorised Microsoft employees reviewing where needed.
Customers approved for modified abuse monitoring sit outside that path. Microsoft states that for them "the data storage and human review process described above is not performed". It is an application, not a checkbox, so raise it during scoping if your data is regulated.
Ask your vendor for the equivalent sentence on their own documentation page. If it exists only in an email, it is not a commitment.
What an agent can do in one run, using AWS Bedrock AgentCore as the worked example
There are hard ceilings on a managed runtime, and they are published. A synchronous request and a long-running job are two different products with two different limits.
AWS Bedrock AgentCore defaults, as published in September 2026:
| Limit | AWS Bedrock AgentCore default | Adjustable | |---|---|---| | Synchronous request timeout | 15 minutes | No | | Asynchronous job maximum duration | 8 hours | No | | Maximum payload size | 100 MB | No | | Idle session timeout | 15 minutes of inactivity | Yes, via `idleRuntimeSessionTimeout` | | Maximum session duration | 8 hours | Yes, via `maxLifetime` |
Those values come from the AWS Bedrock AgentCore quota documentation, which lists the request timeout at 15 minutes and asynchronous jobs at 8 hours.
None of these numbers is a property of agents. A self-hosted agent on your own container sets its own ceilings, and a different managed runtime publishes different ones. Check the runtime you are actually buying.
The design consequence holds either way. Work that cannot finish inside the runtime's request timeout has to be an asynchronous job with a resume point, decided during scoping, not after launch.
Will it scale to our volume?
Yes, within account quotas that you raise on request. Capacity planning is a real step, not a slide.
AWS defaults, as published in September 2026, cap active session workloads at "5,000 in US East (N. Virginia) and US West (Oregon), and 2,500 in other AWS Regions". Those can be raised via Service Quotas. New session creation is throttled separately, at 25 per second by default.
For most business workflows these ceilings are far above real demand. For a consumer-facing agent at peak, file the quota increase before launch week, not during it.
What supervision actually looks like
It looks like a written list of decisions the agent makes alone and decisions a person signs off. Most companies running agents do not have that list yet.
Deloitte surveyed 3,235 IT and business leaders across 24 countries, all directly involved in their organisations' AI programmes. Only 21 percent said they have a mature governance model for agentic AI.
The same report defines what is missing: roughly 80 percent lack "clear boundaries for agents that define which decisions they can make independently versus which require human approval".
You can write that list yourself, and you should try before anyone sells you one. List every action the agent can take. Beside each, put reversible or not, and who reverses it. Anything irreversible, anything a customer sees and anything that moves money gets a named approver. That page is the boundary document.
Gaper writes it with you during scoping and hands it over with the code. If you already have one, bring it, and we build to it.
Is human approval a legal requirement?
In the EU, for high-risk systems, yes. It is a design requirement written into the AI Act, not an optional extra.
Article 14 requires that high-risk systems "can be effectively overseen by natural persons during the period in which they are in use".
The same article says what oversight must let a person do. That includes "interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state". Read that as a build spec. The kill switch ships with the agent, and someone has to be able to reach it.
Do we have to tell customers they are talking to an AI?
If the agent interacts directly with people, the EU AI Act says yes, with one carve-out.
Article 50 requires providers to ensure that people "are informed that they are interacting with an AI system". The obligation does not apply where that is already "obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect".
Do not spend a meeting arguing about what counts as obvious. Disclose in the first message, not in a footer. It costs nothing and removes a whole category of complaint.
What happens when the agent is wrong
Three questions sit under this heading, and none of them is about the model. How do you find out, how do you undo it, and who is accountable.
You find out from the log. Every tool call is recorded with its inputs, its output and its timestamp, and someone reads the exceptions daily in the first month. An agent with no tool log cannot be debugged, only argued about.
You undo it because the design made the action reversible, or because a person approved it before it happened. Writes that cannot be reversed belong behind an approval gate from day one. That is a scoping decision, not a later hardening pass.
Accountability sits with a named person on your side, not with the vendor and not with the agent. They own the boundary document, they hold the stop control, and they decide when to pause the agent. Name them before launch.
When the attack arrives inside the data
This is the second failure type, and it needs a different answer. The agent is not making a bad judgement here. It is being instructed by something it read.
OWASP defines prompt injection as occurring when "user prompts alter the LLM's behavior or output in unintended ways". The indirect form is when the model accepts input from external sources such as websites or files.
Think about what that means for a support agent. It reads customer emails, so it reads whatever a customer chooses to put in one. Text in an inbound ticket can be written to instruct the agent.
The defences are boring, and they limit the damage rather than remove the risk. Narrow tool permissions, approval gates on anything irreversible, treating retrieved content as data rather than instructions, and a log of every tool call.
Injection is an open problem. There is no known complete defence, so the design assumption has to be that the agent will eventually be fooled. That is why irreversible actions sit behind a person, and why anyone telling you the risk is closed has not read the literature.
How to measure whether it worked
With standard telemetry, captured from day one, in a vendor neutral format you can take with you. Not with a slide of before and after anecdotes.
OpenTelemetry publishes named instruments for exactly this. `gen_ai.invoke_agent.tool_calls` tracks the "number of tool calls a GenAI agent makes during a single invocation".
The companion instrument covers speed. `gen_ai.invoke_agent.duration` captures the "end-to-end duration of a single in-process agent invocation".
Together with token usage per run, those give you cost per run and time per run on one dashboard. Ask for them at handover. They are portable, which is part of what owning the agent means.
Are other companies actually running agents in production?
Adoption is real and uneven, and trust lags behind deployment. Two separate samples say the same thing.
A Deloitte survey of 501 US senior leaders was fielded April to June 2026. It found 42 percent have tested or deployed AI agents, while "70% don't feel they can trust and govern agents".
Census data shows the size gradient. In the Business Trends and Outlook Survey, 37 percent of firms with at least 250 employees reported using AI, against roughly 17 to 20 percent of all businesses.
The blocker in both datasets is governance, not model quality. That is the work, and it is the part a roundup of twelve platforms will not do for you.
What you should get at handover
The repository, the deployment configuration, the evaluation suite, the boundary document and the runbook. If any of those five is missing, you are renting, not owning.
A practical checklist to take into any vendor conversation:
| Ask for | What good looks like | |---|---| | Source code and repo access | In your organisation, not the vendor's | | Deployment target | Your cloud account, your authentication | | Data handling statement | A vendor documentation page, not an email | | Approval boundaries | Written list of autonomous vs approved actions | | Telemetry | Standard OpenTelemetry metric names, exportable | | Stop control | A tested way to halt the agent safely | | Runbook | Who operates it, and how to roll back |
We hand over all seven. If you can run the agent without us after month one, the build did its job.
Concrete places agents earn their keep.
Policy matched. Refund ready for approval.
Customer support
Resolve tickets end to end, look up the order, issue the refund, update the case, not just answer FAQs.
Finance & accounting
Reconcile transactions, chase exceptions, and draft the close, wired into the ledger.
account score
Sales operations
Enrich leads, update the CRM, and prep the rep, the busywork that never gets done.
Healthcare ops
Credentialing, compliance reviews, scheduling, and appeals, deployed inside HIPAA boundaries with human review.
Document processing
Read messy PDFs and emails, extract the fields, and route the exceptions.
Internal knowledge
Answer employee questions from your real docs, with citations and freshness.
Common questions.
What is an AI agent for business?+
How is an AI agent different from a chatbot?+
Should we build an agent or buy a SaaS product?+
How long does it take to get a first agent live?+
Why do so many agent pilots fail to reach production?+
Want agents like these in your stack?
Book a free assessment, we'll map where an AI agent creates real leverage in your workflows and scope the first one to ship.