Prove whether Agentforce earns its place.
One real workflow, built in your sandbox on your data, and measured against a scored scenario set. The engagement ends with a recommendation you can take to a board, including the recommendation not to proceed.
- Duration
- Four to six weeks
- Scope
- One workflow, built and evaluated in a sandbox
- Ends with
- Scored results and a go or no-go recommendation
When the demonstration is not enough.
Agentforce demonstrates well. The gap is between a curated demonstration and your permissions model, your knowledge base and the cases your team finds difficult.
You are here if any of these apply.
The pilot is designed for organisations that need evidence before committing, not after.
- Your executive has asked what you are doing about Agentforce, and you need an answer with evidence behind it.
- A service or sales team handles high volumes of repetitive, rules-driven work inside Salesforce.
- You have seen the demonstration and cannot tell whether it survives your data and your permissions model.
- Agentforce licences are on the table and you need to know what they would actually buy.
- A previous AI initiative stalled because nobody defined what success meant.
- You need a defensible position, including the position that now is not the time.
Pilots fail for the same three reasons.
Almost none of them fail because the model was not clever enough. They fail on the engineering around it, which is the part a demonstration never shows.
No definition of success
The pilot runs, everyone is impressed, and nobody can say whether it worked. Without numbers agreed up front, the decision falls back to opinion.
Nothing to ground on
The agent can only answer from what it can reach. Where knowledge is thin, contradictory or locked behind permissions, the output degrades in ways the demonstration data hides.
No boundary
An agent with no defined limits improvises. In front of a customer, on a record it should not have seen, that is the incident that ends the programme.
Six stages,
built to produce evidence.
The build is the middle of the engagement, not the whole of it. Selection and evaluation are what make the result trustworthy.
Use case selection
A working session to pick the one workflow worth testing, and to define what success means before anything is built.
- Candidate workflows scored on volume, rules clarity and risk
- Success criteria agreed in numbers, not adjectives
- The failure conditions that would end the pilot
- Current-state baseline captured for comparison
Data and knowledge readiness
An agent is only as good as what it can ground itself in. We establish that before building.
- Records, fields and knowledge articles the workflow depends on
- Gaps and contradictions in existing knowledge content
- Data Cloud and retrieval options assessed against your licences
- Sandbox prepared with representative, permissioned data
Agent build
The agent itself: topics, actions, instructions and the boundary between them.
- Topic and action design against the real workflow
- Apex and Flow actions where standard actions fall short
- Prompt and instruction engineering with version control
- Conversation design, including what the agent refuses to do
Guardrails and escalation
What happens at the edges decides whether an agent can be trusted in front of customers.
- Permission and sharing enforcement verified, not assumed
- Escalation and human handover paths
- Scope limits so the agent declines rather than improvises
- Audit trail of every action the agent takes
Evaluation
The part most pilots skip. A scored scenario set, run repeatedly, with the results written down.
- Scenario set built from your real cases, including the awkward ones
- Scored runs with pass rates by scenario category
- Human review of a sample by your subject matter experts
- Regression runs after each change, so improvement is measurable
Production readiness
An honest assessment of what standing this up properly would take.
- Licence and platform prerequisites identified
- Deployment, monitoring and rollback approach
- Operating model: who owns the agent once it is live
- The work required before a real customer sees it
Evidence, not a demonstration.
Everything produced stays in your environment and your document set, whichever way the decision goes.
A working agent in your sandbox
Configured in your own environment against your own data and permissions, not a demonstration org. It stays there whatever the decision, so your team can keep exercising it.
Evaluation results
Pass rates by scenario category across the scored runs, with the failing cases listed in full and the reason for each failure identified.
Knowledge and data gap report
What the agent could not ground itself in, and what would have to be written, corrected or exposed for the results to improve. This is usually the most actionable document of the six.
Guardrail and escalation design
The documented boundary: what the agent is permitted to do, what it must escalate, and how a human takes over mid-conversation.
Production readiness assessment
The licences, platform work, integration work and operating model required to run this for real, with the risks named.
Go or no-go recommendation
A written recommendation with the evidence behind it, presented to your team and your sponsor. No-go is a result we deliver when the evidence supports it.
Six weeks, with a checkpoint each Friday.
A narrow workflow with clean knowledge completes in four. Every week ends with a short written update, including the weeks where the results are not what anyone hoped.
- Week 1
Select and define
Workshops with the business owner and the platform team. We choose the workflow, agree the success criteria in numbers, capture the current-state baseline and confirm sandbox and licence prerequisites.
- Weeks 2 to 3
Build and ground
Topics, actions and instructions built in the sandbox. Grounding wired to the records and knowledge the workflow depends on. Permission and sharing behaviour verified early, because it is the most common cause of a pilot quietly failing.
- Week 4
Evaluate
The scenario set is run and scored. Failures are analysed, changes are made, and the set is run again. Your subject matter experts review a sample independently of us.
- Weeks 5 to 6
Harden and decide
Guardrails, escalation and audit finalised. We write the readiness assessment and the recommendation, then present both to the people who own the budget and the customer experience.
What this engagement is not.
A pilot that expands mid-flight stops being a pilot. The boundary below is what keeps the timeline and the evidence honest.
- No production deployment. The pilot lives in a sandbox for its whole duration.
- One workflow only. A second use case is a separate engagement, not a scope change.
- No remediation of underlying org problems. Where they block the pilot we report them.
- No licence procurement or contract negotiation with Salesforce.
- No custom model training. Agentforce and the platform models are what is tested.
- Not a data platform build. We work with the data you can expose inside the pilot window.
The ones we are asked most.
What if the pilot says no?
Then it has done its job. A documented no-go in six weeks is a far better outcome than a rollout that quietly fails in front of customers. The evaluation results tell you what would have to change for the answer to become yes.
Do we need Data Cloud?
Not always. Some workflows ground perfectly well on records and knowledge articles. We assess it in the first week and tell you plainly whether your use case needs it, rather than assuming it does.
What licences do we need to start?
Agentforce enabled in a sandbox, and the relevant platform features for the workflow. We confirm the exact prerequisites before the engagement begins so there is no mid-pilot procurement delay.
Does this work outside Service Cloud?
Yes. Service workflows are the most common starting point because the volume is visible, but we have scoped pilots across sales operations and internal support workflows on the same basis.
How do you protect customer data?
The pilot runs in a sandbox under a mutual non-disclosure agreement, with masked or representative data where production records are sensitive. Access is least-privilege and revoked in writing at the end.
What happens if we decide to proceed?
The readiness assessment is written to be scoped from directly. You can take it to us, to your existing partner, or to your internal team. Nothing in the pilot is built to lock you in.
Bring us the workflow you are least sure about.
The first conversation is with the engineer who would run the pilot. If your use case is not a good candidate, we will say so before you spend anything.