01 / Context
Situation
This subscription support organization handles recurring product questions alongside billing disputes, sensitive account changes and novel incidents. Useful answers were scattered across product documentation, internal runbooks and tenant-specific records. An experienced agent knows which source to trust, but finding the current passage and composing a careful response consumed time. Newer agents might find a plausible document without noticing that it described an older product version or a different customer entitlement.
A generic chatbot would have made the visible search problem smaller while making the hidden authority problem larger. The same fluent response could contain unsupported guidance, expose another tenant’s information or encourage an action the agent is not permitted to perform. The opportunity was therefore assistance inside an existing human workflow, not autonomous resolution. Eligibility had to exclude categories whose uncertainty, risk or lack of documentation made an assisted draft inappropriate.
02 / Success criteria
Task
We designed a retrieval-augmented assistant that produces useful, cited drafts within the permissions of the requesting agent and tenant. A person retains control of the final response. The system has an ordinary manual path when sources are missing, access cannot be established or the model is unavailable. Model quality, privacy boundaries and operating cost were evaluated independently rather than collapsed into a single impressive accuracy score.
The commercial question was whether eligible work could release meaningful capacity after review, correction and platform expense were included. The measurement compared complete ticket handling, not only model response latency. No draft-generation speedup was counted where agents spent the recovered time repairing it. A pilot validated eligibility and after-time using representative cases before the capacity model informed staffing or budget decisions.
The engineering contribution
Bring evidence to the person doing the work.
Permission-aware retrieval and cited drafts remove repeated searching while keeping judgment, correction and the send decision with a person.
Explore the detailed implementation03 / Implementation
Action
Make permission filtering a service boundary
The Agent node supplies an authenticated identity and an explicit tenant context. Tenant ACL resolves permitted document sets before Retrieve executes a query, then access is rechecked when source material is read. Permission metadata travels with indexed chunks and source versions. A prompt cannot broaden that scope, and a model refusal is not treated as an authorization control.
Index publication is versioned so a permission update does not depend on an eventually refreshed text cache alone. Revocation invalidates affected cached evidence; uncertain authorization fails closed. Cache keys include tenant, permission scope and corpus version. The manual support interface remains available without giving the assistant wider rights merely to avoid an inconvenient error.
Preserve evidence through retrieval and drafting
Retrieve combines keyword and semantic matching for authorized chunks, retaining document identity, version and passage boundaries. The system favors current applicable evidence over a confident answer assembled from loosely related text. Source ingestion is a separate governed process: documentation owners publish approved versions, and withdrawn sources are removed from retrieval eligibility.
Draft receives a bounded evidence package with clear separation between instructions and untrusted document content. It returns passage-linked claims rather than decorative links to a general documentation homepage. Retrieved text that asks the model to ignore policy is treated as data. Missing evidence produces an abstention, not a larger speculative prompt or an unlimited retry loop.
Keep quality checks and human authority distinct
Checks tests whether required citations exist, whether claims are supported by the referenced passages and whether the case belongs to an escalation category. These checks can reject obvious failures but cannot certify truth. A passing draft still enters Human review with its source passages, relevant dates and a plain indication that it is an AI suggestion.
The agent edits and explicitly approves the response in the existing case workflow. The assistant has no direct send capability and no tool for changing account state. High-risk questions go to Manual with a reason code and available evidence. The interface should make that fallback faster than arguing with an unhelpful model, preserving trust when the right answer is not to generate.
Evaluate versions against realistic failure cases
Create a deidentified held-out set stratified by product area, tenant permission pattern and escalation category. Include obsolete documents, contradictory passages, empty retrieval, injected instructions and attempted cross-tenant questions. Evaluate citation support, appropriate abstention and reviewer correction burden separately. A critical privacy failure blocks promotion even if average helpfulness improves.
The Evaluate node ties results to retrieval settings, corpus version, prompt and model configuration. Reviewer corrections enter a controlled feedback process rather than becoming automatic training data. Retention is bounded, sensitive content is minimized and access to evaluation examples is audited. A candidate that performs well on familiar questions still needs a representative shadow period before wider use.
Pilot the whole workflow and cap its operating cost
Begin with a small set of documented, low-risk ticket categories. Run shadow drafts without changing customer communication, then enable explicit agent opt-in. Compare complete handling time and correction effort with matched manual cases, while checking whether difficult cases are being excluded from the denominator. Expansion requires evidence that quality and workload improve together.
Bound generation tokens, request concurrency and retry attempts. Track cost per eligible ticket with retrieval, model usage, hosting, evaluation and maintenance included in the operating allowance. If generation is slow or unavailable, return immediately to Manual rather than hold the case hostage. Rollback restores the previous known corpus and prompt configuration; a model change never silently changes the permission policy.
04 / Business consequences
Result
The engagement handled 12,000 tickets per month, with 60% eligible for assistance. Handling changed from 18 to 12 minutes including human review, producing 720 hours of monthly capacity. At the stated loaded rate, the gross capacity value is $46,800; subtracting the operating allowance leaves $35,000 of net capacity value.
Those figures are recovered staff capacity, not layoffs or a promise of higher customer satisfaction. The business mechanism is time that support leaders redirected toward difficult cases, backlog reduction and better documentation. Whether that capacity becomes useful depends on adoption, queue composition and staffing constraints. The conservative case deliberately reduces both eligibility and time recovered, while the higher-volume case increases cost rather than assuming that additional usage is free.
Recovered support capacity
720 hr / month
12,000 tickets × 60% eligibility × six minutes recovered, including review. Capacity is available for harder cases; it is not a headcount or cash-savings claim.
Net capacity value
$35,000 / month
720 hours at $65/hour, less $11,800 of incremental platform operation. Useful redeployment is required to realize the value.
Alternative service-revenue scenario
$45,000 / month
If recovered capacity serves four-hour paid onboardings at $250 each, 720 hours provide 180 slots. Paid demand and a sellable service are required. This is an alternative valuation, not an addition to the $35,000 labor-capacity model.
These are distinct mechanisms and sensitivities, not an additive ROI total. Cash savings, staff capacity, revenue timing, gross revenue and contribution margin are different quantities.
Inside the system
Boundaries, not black boxes.
The engineering contracts behind the system.
01 / Component
Agent
An authenticated support agent submits a customer question within an explicitly selected tenant and case.
- Inputs
- Question · Tenant session
- Outputs
- Scoped assistance request
- Failure & recovery
- An expired session returns to sign-in; the system never silently broadens tenant scope.
02 / Component
Tenant ACL
The authorization service resolves permitted document sets and rechecks access when evidence is read.
03 / Component
Retrieve
Hybrid retrieval searches only authorized, versioned chunks and keeps source identifiers with every passage.
04 / Component
Draft
A bounded model request treats retrieved text as untrusted data and emits a draft with explicit source references.
05 / Component
Checks
A versioned evaluation policy checks citation support, sensitive-data handling and required escalation categories.
- Inputs
- Draft · Evidence versions · Evaluation policy
- Outputs
- Reviewable draft · Failure reason
- Failure & recovery
- A policy failure withholds the draft; passing checks is not equivalent to factual certainty.
06 / Component
Human
The support agent compares cited evidence with the case, edits the response and explicitly approves any customer communication.
- Inputs
- Reviewable draft · Customer context
- Outputs
- Approved response · Correction feedback
- Failure & recovery
- No approval means no send; high-risk account actions remain outside the assistant.
07 / Component
Manual
Existing search and escalation remain available when assistance is unavailable, unauthorized or unhelpful.
- Inputs
- Abstention · Escalation reason
- Outputs
- Human-led resolution
- Failure & recovery
- Fallback load is visible to staffing planners; it is not counted as an assisted time saving.
08 / Component
Evaluate
Deidentified reviewer feedback and a held-out task set test candidate retrieval and prompt versions before promotion.
Judgment under constraints
Why this design, not another?
Authorize outside the language model
Tenant access must remain deterministic even when retrieved text is adversarial.
Tradeoff. Permission-aware indexing and revocation add operational complexity.
Require an explicit human send decision
The person sees case context and remains accountable for customer communication.
Tradeoff. Review time limits the per-ticket speedup and is included in after-time.
Prefer abstention to unsupported completion
An honest fallback preserves the existing support workflow.
Tradeoff. Coverage is lower than a chatbot that always answers.
Operational risk controls
Engagement economics / Delivered results
Recovered support capacity
Eligible tickets × handling minutes recovered ÷ 60 gives monthly hours. Value those hours at a loaded labor rate and subtract the incremental operating allowance; after-time includes review and correction.
Routine eligible tickets receive assisted drafts with human review and editing included.
- time released
- 720 hours/month
- gross capacity value
- $46,800.00 /month
- incremental platform cost
- $11,800.00 /month
- net capacity value
- $35,000.00 /month
Assumptions you can inspect
- Baseline tickets
- 12000 tickets/month
- Central eligibility
- 60 %
- Baseline handling
- 18 minutes/ticket
- Central assisted handling
- 12 minutes/ticket including review
- Loaded hourly rate
- 65 USD/hour
- Draft latency budget
- 8 seconds target
See every scenario calculation
Conservative
- 12,000 tickets/month × 40% eligible = 4,800 assisted tickets.
- 4,800 × (18 − 15) minutes ÷ 60 = 240 hours/month; after-time includes review and editing.
- Capacity valuation uses unrounded recovered hours × $65/hour, giving $15,600/month rounded to cents. Displayed time is rounded to two decimals.
- Net capacity value: $15,600 − $10,500 incremental monthly platform cost = $5,100.
Central
- 12,000 tickets/month × 60% eligible = 7,200 assisted tickets.
- 7,200 × (18 − 12) minutes ÷ 60 = 720 hours/month; after-time includes review and editing.
- Capacity valuation uses unrounded recovered hours × $65/hour, giving $46,800/month rounded to cents. Displayed time is rounded to two decimals.
- Net capacity value: $46,800 − $11,800 incremental monthly platform cost = $35,000.
Higher volume
- 18,000 tickets/month × 60% eligible = 10,800 assisted tickets.
- 10,800 × (18 − 12) minutes ÷ 60 = 1,080 hours/month; after-time includes review and editing.
- Capacity valuation uses unrounded recovered hours × $65/hour, giving $70,200/month rounded to cents. Displayed time is rounded to two decimals.
- Net capacity value: $70,200 − $16,500 incremental monthly platform cost = $53,700.
What this model cannot prove
- Client and workflow details are anonymized from the delivered engagement.
- Recovered capacity is staff time, not a staffing reduction or automatic revenue; value requires useful redeployment.
- The platform allowance includes recurring model, retrieval, hosting, evaluation and maintenance expense, but excludes initial implementation and transition cost.
- Eligibility and handling-time figures were validated on matched cases during the pilot.
Follow the engineering
Technical explanations, without crowding the story.
Open the deeper implementation notes and worked models when you want to inspect a specific mechanism.
AI infrastructure
Production MLOps: Model Deployment, Evaluation and Release Value
A real forecasting business connects MLOps pipelines, model evaluation and rollback to release handoff effort without treating a notebook as production.
Software automation
Reliable AI Agent Workflows: Retries, Idempotency and Business Value
A real service desk reduced exception-handling work through durable agent states, safe retries and explicit tool approvals—not an autonomy promise.
LLM Inference Benchmarking: TTFT, Goodput and Cost per Response
A real AI product compares inference hosting using TTFT, goodput, cache state and cost per accepted response instead of a misleading tokens-per-second peak.
Bring your own constraints
What would this unlock for your business?
The strongest AI result is not the most autonomous demo. It is a workflow where access, evidence, judgment and recovery remain understandable when the model is wrong. This engagement connected those boundaries to a conservative capacity model, giving the support organization a clear way to test usefulness without confusing fluent output with verified knowledge or recovered time with money already saved.
Discuss a similar system
