A chatbot responds. An agent acts.
Claude reads the objective, builds a plan, calls your real tools, checks whether it worked and retries if it failed. With hard limits and human approval where the risk warrants it.
12
tools
orchestrated by one agent
0
blind actions
everything verified or approved
L0–L3
autonomy levels
per tool, not per agent
8
industries
same agentic pattern
UNDER THE HOOD
It picks each tool and writes down why.
An agent that calls APIs without reasoning is a script with risk. Claude exposes the objective it received, which tool it chose at each step, what it verified and what it would do if something fails.
Sample goal
The goal comes in
Objective: "The client Distribuidora del Sur hasn't bought in 60 days, they were recurring. Recover them." The agent has access to: CRM (read/write), order history, WhatsApp and calendar.
Reads the real state first
crm.getAccount() + orders.history() → confirms 61 days without order, high average ticket, no open incidents
Chooses the right tool
doesn't open a support ticket (no complaint); decides WhatsApp channel because they responded there the last 5 times
Drafts with context, not generic
claude.draft() uses their last purchased product + reactivation discount within policy
Defines the limit before acting
sending message = low risk (auto); applying discount >10% = requires human OK → leaves it proposed
Score
91/100
Verdict from Claude
Message sent automatically. 12% discount stays in approval queue with the reasoning attached.
The goal comes in
Objective: "A transfer of $4,820 arrived without reference. Assign it." The agent has access to: bank (read), accounts receivable, ERP and email.
Searches before guessing
ar.openInvoices() → 3 candidate invoices with close amounts; none matches exactly
Crosses multiple signals
amount + value date + issuing bank → 1 candidate with 89% match, invoice #INV-2291
Verifies before writing
erp.checkBalance(client) confirms exact pending balance $4,820 → match rises to 99%
Knows when NOT to decide alone
match <95% would have escalated to accounting; with 99% it reconciles and notifies, no OK needed
Score
88/100
Verdict from Claude
Payment reconciled against INV-2291 automatically. Confirmation email to client triggered.
The goal comes in
Objective: "Order #8841 is marked as undelivered, the client is complaining via WhatsApp." The agent has access to: logistics, inventory, CRM, WhatsApp and approvals.
Reconstructs what happened
logistics.getPOD(#8841) → no delivery photo, driver marked "address not found"
Evaluates the least-friction option
inventory.check() confirms stock; calculates rescheduling vs refund → rescheduling is viable today
Acts on the safe, stops on the costly
responds to client and schedules re-delivery (auto); a refund would have required approval
Closes the loop
verifies the new route was assigned; if not, retries once and then escalates to human
Score
84/100
Verdict from Claude
Re-delivery scheduled for tomorrow, client notified, case marked as contained. No money touched.
WHAT AN AGENT DOES
Five things a chatbot could never do
The steps of the agentic loop: the agent plans, executes on your real systems, verifies and recovers from failures on its own.
agent.plan(goal)
Converts an objective into a sequence of steps
You don't tell it 'do step 1, then step 2'. You give it the expected result and Claude breaks down which tools it needs, in what order and with what dependencies.
› input
goal: "client claims double charge, resolve it"
› claude →
plan generated:
1 · payments.lookup(client) — confirm charges
2 · if duplicate → refund.create() [requires OK]
3 · crm.logCase() + whatsapp.reply()
→ 3 steps · 1 approval point
agent.callTool()
Uses your real APIs, not invented answers
Real tool use: the agent invokes your CRM, ERP, bank or WhatsApp with structured parameters and receives the result back to reason on it.
› input
tool: crm.getAccount · args: { id: "ACC-1182" }
› claude →
→ crm.getAccount("ACC-1182")
← { status: "active", last_order: "-61d",
avg_ticket: 4820, incidents: 0 }
the agent now reasons on real data,
not on what it 'thinks' happened
agent.verify(result)
Checks that the action actually worked
The difference between "I ran a command" and "the result is correct". The agent reads the state again after acting to confirm the real effect.
› input
after logistics.reschedule(#8841)
› claude →
verify → logistics.getRoute(#8841)
expected: status = "assigned"
received: status = "assigned" ✓
effect confirmed · moves to next step
if it did not match → retry()
agent.retry(step)
Recovers from failures without intervention
An API went down, a parameter was wrong, a timeout. The agent diagnoses the error, adjusts and retries — with a cap to avoid infinite loops.
› input
error: whatsapp.send() → 429 rate_limited
› claude →
failure detected: rate limit (attempt 1/3)
strategy: 8s backoff + retry
attempt 2 → sent ✓
if 3/3 fail → handoff to human
never fakes success that did not happen
agent.handoff()
Knows when it should NOT decide alone
The most useful agent is the one that knows its limit. High-risk or low-confidence actions are handed to a human with all the context already assembled.
› input
action: refund.create($4,820) · confidence 71%
› claude →
⚠︎ high-risk action + confidence < threshold
does not execute · assembles the full case:
evidence + proposed plan + 1-click approve
→ to human queue, not blindly
human decides in 10s with context, not 10 min
LIVE SYSTEM · Grupo Latitud
This is what it looks like inside an AI-operated system — and how it is governed.
Each agent works with three pieces: its tools, its cap and its trace. Here there are twenty-two, across four countries.
Modules in this tour
- 01Each agent with its purpose and its cap
- 02The level is set per tool
- 03Every run is replayable
- 04Whatever commits money waits
- 05Who wrote all of this
Boosty Standard for Operating with AISimulated AI · demo data
THE SAME PATTERN
One single agentic loop. Tools change, judgment does not.
We don't build a different agent per industry. The same perceive→plan→act→verify pattern operates on the tools specific to each business. This is already running in production.
Signals specific to the industry
Signals specific to the industry
Signals specific to the industry
Signals specific to the industry
Signals specific to the industry
Signals specific to the industry
THE AGENTIC LOOP
A loop that repeats until the goal is met.
It perceives, plans, acts and verifies, and starts again while the goal is still open. Every turn runs on the real state of your systems.
Perceives
agent.observe()
→Plans
agent.plan(goal)
→Acts
agent.callTool()
→Verifies
agent.verify()
› agent.observe()
Perceives
Reads the real state: crm.getAccount() · orders.history() — 61d with no order, high ticket
IN THE SYSTEM · TRACES AND CIRCUITS
Every run can be replayed. Step by step.
What triggered it, what context it read, which tool it called, what came back, where it waited for a person and what it recorded at the end. If something goes wrong, you can see where.
Real screenshot of the system · demonstration data
TOOL USE LIVE
Watch the agent choose and call tools
It thinks, calls your real API, reads the result and reasons over it until it has enough confidence to act. The case: reconciling a payment with no reference.
GUARDRAILS + HUMAN APPROVAL
Autonomy with a brake. Not all or nothing.
The same agent acts on its own where it is safe and stops where it is costly. Low risk and high confidence: it executes. Irreversible or below threshold: it hands the case to a person with the context ready. You define where the line is.
Requested action
Reply to the customer on WhatsApp
Risk level
Low · reversible
Agent confidence
98%
Why
Reversible action, no financial impact, high confidence.
Decision: execute on its own
Executed automatically — without waiting for anyone.
you define the threshold and the policy · the agent never crosses them
IN THE SYSTEM · TOOLS AND POLICIES
Autonomy is set per tool and per country.
Each tool has a kind, a risk and a maximum level from N0 to N3. The same agent can execute in one country and only propose in another, according to the policy and the history of that place.
- N0 observes · N1 prepares · N2 executes and notifies · N3 autonomous
- The risk of each tool, declared
- Which agents use each one
Real screenshot of the system · demonstration data
CONNECTED STACK
The agent acts on what you already use.
Each integration is a tool the agent can invoke. If your system has an API, the agent talks to it.
Claude · Anthropic
Claude PartnerThe agentic engine: planning, tool use, verification and recovery reasoning
Kommo CRM
PartnerRead/write tool: the agent reads the account and moves the stage on its own
Monday.com
PartnerThe agent creates projects and advances tasks when the condition is met
WhatsApp Business
Action channel: the agent responds and notifies through where the client responded
Make / n8n
Tool adapters for ERPs or legacy cores without a modern API
Supabase
Deno Edge Functions that execute the loop + agent state memory
Frequently asked questions about Agents with Claude
A chatbot responds with text. An agent takes actions: reads your systems, calls your APIs, verifies the result and retries if it fails. The chatbot tells you what to do; the agent does it and shows you what it did. The difference lies in the tool use and verification architecture around the model.
Only if you explicitly authorize it. By default we define guardrails by risk level: reading and notifying is automatic, moving stages or scheduling is usually automatic with verification, and everything that touches money or is irreversible stays in one-click human approval. You raise the autonomy level when you trust the judgment.
The agent detects the error (timeout, rate limit, 500), diagnoses the cause, adjusts and retries with backoff, up to a configurable cap. If it exhausts the retries, it doesn't fake success: it hands off to a human with the full context. It never reports an action as completed if it didn't verify it.
No. The agent doesn't replace anything: every system you already use becomes a tool it can invoke. If it has a REST API we integrate it natively; if it's legacy without an API, we expose it via Make/n8n. Your CRM, ERP and bank remain the source of truth.
We define an explicit tool catalog with permissions per action (read-only, write with verification, write with approval). The agent cannot invent or call anything outside that catalog. Every call is logged with its parameters and result for auditing.
Because an agent that acts without explaining is a risk your team won't approve. Exposing the plan, the chosen tools and the verification is what allows trust to build and autonomy to increase gradually. Black boxes that touch your operation don't get adopted.
In production we orchestrate processes with 8 to 12 tools with multiple decision and verification points. The practical limit is set by how clear the objective and the guardrails are; with that defined, the agent chains the steps on its own.
It is defined in the assessment. The scope — what gets built first and what waits — comes from what we see in your operation, not from a catalog. Book 30 minutes and we give you the range in writing.

A WORD FROM THE FOUNDER
“A useful agent acts, verifies what it did and interrupts you only when the decision is yours.”
Think of a new shift lead: before handing over the keys you explain the objective, give them the tools, tell them how much they may decide alone and ask them to write down what they did. An agent works the same way. What makes it useful in production is that frame: the system’s tools, a clear objective and the habit of confirming the effect instead of assuming it.
An agent with Claude receives an expected result, builds the plan, calls your CRM or your bank and confirms the change happened. Each tool has its autonomy level: whatever sits within the cap it executes and reports; whatever commits money it prepares and leaves for a person. That level is raised tool by tool, as evaluations and usage justify it.
For the business, that means a process that moves without anyone acting as glue between systems, with every run on record. Book 30 minutes with me: we take one of your processes and watch the agent decide and verify live. Which of your processes depends today on copying data from one system to another?

Gabriel Montiel
CEO · Boosty Digital
LET'S TALK
Ready for AI that finishes the job?
Schedule a 30-minute assessment. We bring one of your real processes and show you the agent planning, acting and verifying live. No corporate deck.