Agent washing: how to tell a real autonomous agent from a rebranded chatbot before you sign
Gartner estimates only about 130 of thousands of “agentic AI” vendors are the real thing. A buyer’s field guide to spotting rebranded chatbots and automation before procurement, with a vendor scorer and an RFP checklist.
Every vendor deck now says “agentic.” The label got popular faster than the technology did, and the gap between the two has a name: agent washing, rebranding chatbots, robotic process automation, and assistants as autonomous agents. Gartner put a number on it, estimating only about 130 of thousands of agentic-AI vendors are the genuine article. That is not a rounding error; it is the market. For a buyer, the risk is concrete. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 20271, citing escalating costs, unclear business value, and inadequate risk controls. Some of that is agent washing coming due, software that could never do what the label promised. This guide gives you a field test. It separates the four capabilities a real agent has from the ones a chatbot only imitates, then puts the sharpest question at the top: when the agent acts, can it prove what it did? An accountable agent produces a record you can check yourself, not a screenshot you have to trust.
What is agent washing, and why did Gartner count only about 130 real vendors?
Agent washing is the practice of dressing up existing software, chatbots, RPA scripts, virtual assistants, in the language of autonomous agents without adding the underlying capability. The word does the selling; the product stays the same. Gartner named the pattern directly and estimated that only about 130 of the thousands of vendors claiming to be agentic actually are (RCR Wireless3). The distinction matters because the two categories fail differently. A chatbot that overpromises frustrates users. An “agent” trusted with consequential work it cannot safely perform creates real exposure.
The cost of getting this wrong is already visible in the forecast. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 20271, pointing to escalating costs, unclear business value, and inadequate risk controls. Read that list as a buyer, not a headline: two of the three failure modes are things you can screen for before you sign. Unclear value and weak controls are questions you ask in procurement, and the vendors who cannot answer them are usually the ones who leaned hardest on the label.
What four capabilities does a real AI agent have that a chatbot doesn’t?
Strip away the marketing and an autonomous agent is defined by what it can do, not what it can say. A chatbot responds inside a conversation; an agent pursues a goal across steps, tools, and systems, and stays accountable while it does. Four capabilities separate the real thing from a rebrand. Use them as a checklist, because a genuine agent clears all four, and most agent-washed products fail at least one, usually the last.
- Goal-directed action: it plans and executes multi-step work toward an objective, rather than answering one prompt at a time.
- Tool use with real effect: it can call tools and take actual actions in your systems, not just describe what someone should do.
- Bounded, reversible autonomy: its actions have limits and an undo path, so a wrong move is contained rather than catastrophic.
- Verifiable accountability: it produces an independently checkable record of what it did, so you can confirm actions without taking the vendor’s word.
How do you test whether an AI agent can actually prove what it did?
The first three capabilities are about power, can it act? The fourth is about trust, can it prove it acted correctly? Put that question at the top of your evaluation, because it is the one agent washing cannot fake. A chatbot in a costume can be scripted to look goal-directed in a demo. What it cannot produce is a record of its actions that you can verify yourself, without trusting the dashboard that generated it. “Independently verifiable” has a precise meaning: a third party, or you, can confirm the agent did what it claims using evidence the vendor cannot quietly alter after the fact.
This is where most “agent” pitches quietly downgrade. Ask how you would prove, three months from now, that a specific action happened as recorded, and watch whether the answer is a verifiable receipt or a screenshot. Internal logs are better than nothing, but logs a vendor controls prove very little when something goes wrong. The stronger answer is a cryptographically signed, tamper-evident trail: each action carries its own proof, anchored so that after-the-fact edits are detectable. RankShield’s helix is built around exactly this, an agent that can prove what it did, not merely assert it. Make “can it prove its actions?” the first line of your buyer’s test, not the last. See the deeper treatment on verifiable AI security.
What does agent washing cost the buyer who gets it wrong?
It costs more than a disappointing demo, because the failure shows up after you have already built around the tool. The visible cost is the cancelled project: Gartner’s expectation that more than 40% of agentic initiatives will be scrapped by the end of 20271 is, in part, agent washing coming due, software that could never do the autonomous work its label promised, discovered only once real tasks and real budgets were riding on it. By then you have paid twice, once to adopt and integrate, and once to unwind and replace, plus the opportunity cost of the months in between.
The less visible cost is risk you cannot see until it fires. An agent-washed product trusted with consequential work, moving money, changing records, acting in your systems, without genuine bounds or a verifiable trail, is exposure disguised as productivity. When something goes wrong, the same missing capability that made it agent-washed, no independently checkable record, means you cannot even reconstruct what happened cleanly. That is why the verifiability question is not a nice-to-have at the bottom of an evaluation; it is the single screen that most reliably separates the tools that will survive contact with production from the ones that will become the cancelled 40%. Screening for it in procurement is far cheaper than discovering its absence during an incident.
Is the vendor a real agent or a chatbot in a costume?
Score a vendor against the same five checks a real evaluation uses. Answer honestly, because the bands are calibrated to Gartner’s finding that most self-described agentic vendors are not the real thing. The value of scoring rather than eyeballing is that it forces the uncomfortable question on each axis: not “does the demo look autonomous,” but “can it act, stay bounded, and prove it,” one capability at a time. A vendor can dazzle on one axis and quietly fail another, and it is almost always the last axis, verifiable accountability, where the costume slips.
What questions should you put in your agentic AI RFP?
Turn the field test into procurement language. Drop these ten questions into your RFP verbatim; the ones a vendor dodges tell you as much as the ones they answer. Weight the verifiability and control questions most heavily, because they are the hardest to agent-wash and the most expensive to discover you lack after go-live.
- Describe a task the product completes end to end, autonomously, and name the specific steps it plans and executes.
- Which real actions can it take in our systems, and which are read-only or human-in-the-loop?
- What bounds constrain its autonomy, and how do we configure or tighten them?
- When it takes a wrong action, what is the undo path and how fast does it reverse?
- How does it produce an independently verifiable record of each action, and can we check that record without your dashboard?
- Is the action trail tamper-evident, so an after-the-fact edit would be detectable by us or a third party?
- Can we watch the agent live and stop it mid-task with a kill switch?
- What happens to in-flight actions when we halt it, are they rolled back or left partial?
- What concretely distinguishes this from the chatbot, RPA, or assistant we may already own?
- What is the measurable business outcome, and how do we verify it rather than take it on faith?
Questions, answered.
What is agent washing?
Agent washing is marketing existing software, typically a chatbot, an RPA script, or a virtual assistant, as an autonomous “AI agent” without adding the underlying autonomous capability. The label changes; the product does not. Gartner named the pattern and estimated that only about 130 of the thousands of vendors claiming to be agentic genuinely are, which makes agent washing less an exception than the default state of the market. For a buyer, the danger is trusting consequential work to software that was never built to do it safely.
How can I tell a real AI agent from a rebranded chatbot?
Check for four capabilities a chatbot cannot fake: goal-directed action across multiple steps, tool use that takes real effect in your systems, bounded and reversible autonomy with limits and an undo path, and an independently verifiable record of what it did. A genuine agent clears all four; agent-washed products usually fail at least one, most often the last. Of the four, verifiable accountability is the sharpest test, because a demo can be scripted to look autonomous but cannot produce proof of its actions that you can check yourself.
Why is “can it prove what it did” the most important question?
Because it is the one capability agent washing cannot fake, and the one that matters most when something goes wrong. A chatbot in a costume can look goal-directed in a demo, but it cannot produce a record of its actions that you can verify without trusting its own dashboard. Independently verifiable means you or a third party can confirm what the agent did using evidence the vendor cannot quietly alter after the fact, ideally a cryptographically signed, tamper-evident trail. Ask how you would prove a specific action three months later, and see whether the answer is a verifiable receipt or a screenshot.
Are internal logs good enough to prove an agent’s actions?
They are better than nothing, but they prove little when it matters most, because logs a vendor controls can be incomplete or altered, and you are trusting the same party whose product is in question. The stronger standard is a tamper-evident, independently checkable trail where each action carries its own proof and after-the-fact edits are detectable. The difference is the gap between “our records say this happened” and “here is proof anyone can verify,” which is exactly the difference that decides disputed cases and audits.
What does agent washing have to do with the 40% of projects Gartner expects to be canceled?
A meaningful share of that predicted cancellation is agent washing coming due: software adopted on the strength of the label that could not actually perform the autonomous work, discovered only after teams built around it. Gartner cites escalating costs, unclear value, and inadequate risk controls as the drivers, and two of those, value and controls, are exactly what agent-washed products cannot substantiate. Screening for genuine capability and verifiable controls in procurement is how you avoid joining the 40%.
What should I put in an RFP to screen for agent washing?
Ask the vendor to describe a task the product completes end to end autonomously and name the steps; which real actions it can take versus read-only; what bounds constrain it and how you configure them; the undo path and reversal speed for a wrong action; how it produces an independently verifiable record you can check without their dashboard; whether that trail is tamper-evident; whether you can watch and stop it live; what happens to in-flight actions on halt; how it differs from tools you already own; and the measurable outcome and how you verify it. Weight the verifiability and control questions most heavily.
References
Jamie Kloncz
Founder & CEO, RankShield
Jamie Kloncz is the founder and CEO of RankShield, the verifiable AI and quantum security platform. He started the company after two attacks landed in a single week: his phone was cloned, and his business was hit by a click-fraud campaign. One targeted him as a person, the other his livelihood, and no single tool defended both. That experience, together with surviving an AI voice-clone scam, shaped RankShield’s core belief: the threats of the AI age are personal first, and trust should be something you can check, not just extend.
Make every AI action provable.
RankShield is the verifiable, quantum-safe AI security platform — protection you can check, not just trust.