The five decisions, and 23 frameworks
A professional-services firm that starts delivering its work through AI agents plus human service managers has to decide five things: who owns the launch, how humans and agents split the work, what the price meters, what the offering is called, and what gates the quality. This page maps the published thinking on all five, with one firm's live practice as the worked example.
The map: five decisions
Everything the research found sorts under five headings. Each one is a real decision with at least two defensible answers, which is what makes this a framework rather than a checklist. The worked example below shows one firm's answers; the library further down shows how 23 published frameworks carve the same territory.
Who owns the launch
The documented successes give a new AI service line to a small embedded team with a named senior sponsor, and hand it to the scaled organisation only after it clears a measured bar. Harvey's law-firm rollouts run a three-phase pilot of six to nine months owned by forward-deployed engineers plus a partner-level sponsor, and the bar is a gold-standard evaluation set benchmarked against the partner's own past output. Salesforce ran itself as Customer Zero for 18 months before generalising the playbook.
The gate is a measurement, never a date.
Harvey gates on evaluation sets against expert ground truth. Salesforce gated on CSAT climbing from 76 to 90. Zendesk went furthest and made the gate the billing event itself: an interaction bills only as a verified resolution. A pilot that graduates on a calendar instead of a threshold is how the 95%-no-return statistic below happens.
How humans and agents split the work
The convergent pattern is humans above the loop, not out of it: agents absorb the routine volume, humans hold judgment, escalation and the relationship. McKinsey's agentic-organization work formalises it, TSIA splits delivery into human-led for strategic accounts and AI-augmented for the rest, and coaching platforms shipped hybrid AI-plus-human coaching rather than either extreme, citing mixed client preference.
Crossing the line costs you in public.
Klarna pushed to AI-only service in 2024 and its CEO conceded the reversal by May 2025: the efficiency push bought lower quality. The pattern that survives is the draft-then-approve primitive: the agent produces, a human releases, and the human's time moves up a level instead of disappearing.
What the price meters, and how the bill lands
Live pricing for agent-delivered work runs a spectrum: per seat, per action, per conversation, per work unit, per outcome. The published poles: Intercom's Fin at $0.99 per resolution with a guarantee tied to a resolution floor, Salesforce cycling through three models in a year (per conversation, credits, per user), and Sierra's contract-negotiated outcomes where an unresolved conversation is free. Kyle Poyar's survey shows hybrid, a base fee plus a usage or outcome layer, rose from 25% to 37% adoption in one year and is now the modal answer.
The margin heuristic has inverted.
Bessemer benchmarks AI-delivered services at 50 to 60% gross margin against 80 to 90% for classic SaaS, and a16z's Sarah Wang argues an AI product showing SaaS-grade margins is itself suspicious: it implies too little real model usage inside the product. The named protection lever is model routing, cheap models for routine calls, frontier models where the work is hard. And a hidden COGS line: McKinsey's FinOps survey found most agentic spend goes to correcting outputs, not producing them.
What the offering is called
The market splits cleanly by what is being sold. A single agent sold to end users gets a human first name: Ava, Alice, Piper, Fin, Amelia. A multi-agent platform sold to an enterprise stays functional and firm-prefixed: Zora AI by Deloitte, EY.ai, KPMG Workbench, agent OS by PwC. Incumbent professional-services firms uniformly choose the endorsed sub-brand: a distinct name with the parent firm visibly attached, spending existing trust rather than building new.
Anthropomorphism is a two-edged instrument, and there is a red line.
A human name lowers the barrier in casual service contexts and backfires in high-stakes professional ones, where a failure by a named character collapses more trust than a failure by a tool. The hard red line is literal employment status: Lattice gave AI agents employee records and managers in 2024 and reversed within three days under near-universal backlash. Softer "digital worker" language has drawn no comparable fire. Disclosure is getting a legal floor too: EU AI Act Article 50, binding August 2026, requires telling consumers they are talking to an AI; professional services runs a softer test, a human reviews and owns the final deliverable.
What gates the quality
The cheapest durable quality gate is one the organisation already had. Harvey's insight was that law firms already review every associate's work, so the agent only has to clear an existing gate, not a new bureaucracy. The most mature launches then fuse the gate with the price: Intercom backs Fin with a guarantee of up to $1M tied to a 65% resolution floor, and Zendesk bills only what an independent check verifies as resolved. The pricing model and the service-level agreement become the same mechanism.
The worked example: CauseMatch
CauseMatch is a fundraising platform for nonprofits with a client base of roughly a thousand organisations, a customer-success and coaching team, and a board decision from May 2026 to grow through AI-enabled expansion of its services. It has been answering the five questions in live operation through 2026. The deep material, with the real numbers, sits on its own gated pages: the client service operating model and the pricing and billing exploration. What follows is the mechanism level.
Decision 01R&D owns a new service until it is proven
The product team resolved the launch-ownership question explicitly in June 2026: R&D owns a new AI service until it has been sold to and delivered for initial clients, and only then hands it to customer success and sales. The first service through the model ran as a structured pilot: pitch broadly, sell to several clients, remove friction found in the offer page, then hand off on a date with training scheduled and adoption incentives considered. Idea intake got its own machinery, an async vote across the team on a shared backlog. This is the piece the outside literature is thinnest on: the research pass found rollout and pricing richly documented but frontline enablement mostly absent from public sources, so a firm that writes its handoff model down is publishing something the market has not.
Decision 02Deterministic playbooks, and the agent only where it adds value
The operating thesis: not every decision needs a model. Service delivery is encoded as playbooks whose routing is deterministic, the possible outcomes of each step and the next step for each outcome are data, selected by a human in a purpose-built approval surface. The LLM is reserved for the zones where generation earns its cost: analysing client data to recommend a path, and composing the personalised client communication. The stack decision followed the same logic, own the control plane, rent the transport: capabilities exposed through the company's own MCP, thin custom interfaces for the service managers, and a draft-then-approve primitive so nothing reaches a client unreviewed.
The playbook library has its own architecture, researched separately.
One playbook equals one engagement equals one department's complete pass at one commitment. Four layers: base playbook, shared fragment, client overlay, seniority render, where seniority is a render-time key rather than a forked file, so a senior variant is physically incapable of dropping a check. That came out of a 22-agent research sweep across CS platforms, incident runbooks, BPMN engines and the internal tooling of large engineering organisations; it lives at the operating-model exploration, v12.
Decision 03Pricing is two independent axes
The pricing work reframed a stalled conversation by splitting it: axis A is what you meter, the number the price scales with, which decides whether the largest client pays in proportion; axis B is how the bill lands, how much appears on one invoice, under which name, on which date, which decides whether the total is felt as an event or as ordinary operating cost. Nothing on one axis fixes the other, and the two cost differently: the good axis B moves are cheap and reversible, the good axis A moves need a story and a renewal. So the sequence, not the menu, is the recommendation: fix how it lands first, then fix what it meters where a price change has a natural occasion.
Fundraising got to outcome pricing a decade early.
The AI world's newest pricing idea, charge for the outcome, is charity crowdfunding's oldest one: a success-based percent of funds raised. Which means this firm already knows what the vendors above are discovering, that pure outcome pricing concentrates the bill on your best client and invites a cap, and a cap is an axis A instrument solving an axis B problem. The two-axes analysis, with the arithmetic, is on the pricing exploration.
Decision 04The sub-brand already exists, the persona does not yet
CauseMatch already runs one endorsed sub-brand: its in-house creative studio carries its own name on the offer, presented as part of the company. The open question its CEO raised, whether services should bill as separately named practices on their own schedules, is the same architecture the Big Four chose, an endorsed house, one step further. The next step on that road is the persona question below: not a named practice but a named agent.
Decision 05The human gate was there first
The service model inherits its quality gate the way Harvey does: coached campaigns already run through named service managers who review what goes to a client, so agent output enters an existing review, and the approval surface makes the review a recorded step rather than a habit. Value-based presentation of the offer was a deliberate choice in the first launch, to protect perceived quality while delivery got faster.
Where the practice meets the research
Pilot ownership sits with an embedded team plus a named sponsor, and hands off on a measured bar.
R&D owns the service until proven with real clients, then a dated handoff to CS and sales with training attached. Convergent, arrived at independently.
Outcome pricing needs a mechanically detectable outcome, and drifts to hybrid wherever the vendor does not control the result.
Has run percent-of-raised for years and already lives the failure mode the vendors are just meeting: the cap. Its answer is the two-axes sequence rather than a new meter. Ahead of the literature.
Frontline enablement for selling AI services is the undocumented layer: no named certification or incentive playbook surfaced in public sources.
Training scheduled at handoff, adoption incentives debated, transcript analysis used to sharpen the pitch. A gap in the market this framework can fill.
Back-office automation is where measured AI returns concentrate; customer-facing agents are the hard case.
Puts the agent behind a human service manager rather than in front of the client, which keeps the client-facing surface human while the agent does the back-office share of the work. Consistent with the evidence.
The frameworks, side by side
Both sweeps together verified 23 published frameworks: 11 from consultancies, analysts and academia, 12 from vendors and investors. Each is drawn the same way, the framework at the root, its pieces beneath, and the actual text behind every node, tap a row to open it. Provenance sits on every card: who published it, when, and the link. Where the source publishes its own diagram, the card says so and links it rather than redrawing it. Filter by what each framework is for.
The framework library above renders from the research data. To comment on a specific framework, attach the note to its card. Full JSON, with every node text and source, is in the vault record.
Persona-based pricing: the verdict
The specific question: has anyone given the agent a persona and then priced, billed and branded that persona as its own semi-separate offering, a named agent a client hires, with its own brand and its own line on the invoice? A dedicated research pass tested 18 candidates against the four parts of that claim.
| Company | Persona | Priced as unit | Own invoice line | Semi-separate brand | What actually happens |
|---|---|---|---|---|---|
| Sintra | Twelve named helpers with faces, each sold individually at $39 a month. The persona is the product taxonomy itself. | ||||
| Mesaya | "Hire the role. We run the employee." From $400 a month per named agent, priced by deliverables, billed like payroll. The closest existing thing to the idea. | ||||
| 11x | Alice and Julian each carry their own product page and pricing page, from $3,750 and $5,333 a month; "Add Alice to any plan" is an add-on line. Inside the SKU the meter is volume. | ||||
| Qualified | The instructive reversal: Piper launched as a paid add-on SKU and was deliberately un-SKU'd, folded into every plan. The persona sets the price level, then vanishes from the price structure. | ||||
| Artisan | Spent about $2M on "Fire Steve. Hire Ava." billboards, never priced per Ava, retired the slogan in August 2026 and hired human BDRs. | ||||
| Intercom / Fin | Own domain, own brand, sells with or without the parent product, so strong on brand and line; but the unit is the resolution, not the persona. | ||||
| Sierra | The inverse: the client names the agent (ThirdLove's Barbra, SiriusXM's Harmony), so at enterprise scale the persona flips to the buyer's brand and cannot be the vendor's offering. | ||||
| Garfield Law | The structural alternative: the SRA authorised the AI as its own regulated law firm with its own bills, £2 chaser, £7.50 letter before action. The persona became the firm. |
Why the market stops where it stops
The persona sells the deal; the meter bills the deal. The human-salary comparison anchors the price level and then a usage unit takes over the structure, because agent cost is variable compute and a flat per-persona price is margin risk the vendor eats. Vendors who tried the persona SKU walked it back: sold as an add-on, a persona reads as optional. Enterprise procurement has no purchase order shaped like a person, and at enterprise scale the buyer wants the agent wearing the buyer's brand. Where full matches exist, the persona is the whole catalogue, not a wrapper on one product.
The open position
The billing plumbing already exists in the agency channel, white-label platforms rebill an "AI Employee" line per client at healthy margins. What is missing is not the invoice line. It is the name on it. In professional services the space is empty because the pricing practice itself is immature: the emerging convention itemises deliverables or discounts hours, and none of it carries an agent's name. A firm that already runs an endorsed sub-brand, already bills through its own rails, and already has a trusted human layer in front of clients is holding the pieces nobody has assembled. The four-part test above is the spec; the risk ledger is the Qualified and Artisan rows.
What fails, on the record
Promising the agent instead of the service
Klarna's AI-only service push was publicly walked back within a year, its CEO conceding it traded quality for cost. Artisan retired "stop hiring humans" and hired humans. The reversal pattern is consistent: the provocation buys attention and then bills it back with interest.
Disclaiming your own agent
Air Canada argued its chatbot was responsible for its own false statement and lost in tribunal, February 2024. The output of your agent is your firm's output from the first day; every quality gate above exists because of this.
Pilots with no P&L
MIT's Project NANDA found 95% of enterprise gen-AI pilots deliver no measurable return, and the successes concentrate in back-office automation rather than customer-facing agents. The implication for a services firm is sequencing: prove the agent behind the human first.
Agent washing
Gartner estimates only about 130 of the thousands of vendors marketing agentic AI have genuine agentic capability, and predicts over 40% of agentic projects cancelled by end of 2027. The trust tax spills across the category: every overclaimed launch raises the proof bar for the next honest one.
Open for v2
Where this could go next, depending on which thread pulls: 1. fold the 23 frameworks onto one shared spine, the five decisions, so the side-by-side becomes a real comparison rather than a parallel display; 2. draft the framework article itself, the five decisions with the CauseMatch worked example and the persona verdict as its sharpest section; 3. go deeper on the empty quadrant, what a persona-billed service line would concretely need, the CFO test, the EU disclosure floor, the un-SKU risk; 4. the enablement gap, a publishable handoff-and-training playbook, since the research says nobody has written one.
Research: six parallel agent passes, 1 September 2026, roughly 130 sourced findings and 23 verified frameworks; raw JSON in the vault record. Companion explorations: CauseMatch: Client Service Operating Model and CauseMatch: Pricing and Billing.
The brief this page answers, verbatim.
"Build a NEW Creative Exploration on LEVERAGING AI FOR PROFESSIONAL SERVICES, to be iterated with me and published as a FRAMEWORK ARTICLE. ... Mine the CauseMatch material in gbrain for OUR approach: how we launch services for our client base using AI plus our customer success managers. ... PRICING of AI-driven services (dedicated agents) ... BRANDING of AI-driven services ... PERSONA-BASED PRICING specifically. Has anyone actually given the AGENT a persona, and then priced, billed and branded that persona as its own semi-separate offering? A named agent that a client 'hires', with its own brand and its own line on the invoice. Find real examples, or prove the space is empty and say so. ... FRAMEWORKS. Search hard for any published framework for launching, servicing and making AI-driven services successful. Bring in EVERY framework you find, each one organised the same way. ... I want to understand MULTIPLE different frameworks that exist today, side by side. My starting idea: a zoomable tree per framework, framework at the root, pieces under it, sub-pieces under those, and clicking a node reveals the actual text."