What model-assisted analysis actually costs, what it actually produces, and what the demo does not prove.
In April I wrote that I had built a contact centre intelligence dashboard in thirty minutes from an airport departure lounge, on a phone, for a retailer that does not exist.
That post did numbers. It also had three problems, and they are the reason I am writing it again rather than leaving it up as it was.
The first is that I never showed the dashboard. A post whose entire argument rests on an artefact should contain the artefact. The second is that the comparison I leaned on — fourteen weeks, six to eight people, £200k to £400k — was a figure I had absorbed from a decade of watching BI programmes rather than one I could source. The third is that I named the specific models I used, which turned an argument about method into something that read like a product endorsement, and which was out of date within a quarter.
So: the artefact is below, the numbers are sourced, and the models are unnamed on purpose.
1. The artefact
Bridworth Group — Contact Centre Intelligence
847,000 contacts · 7 sites · 5 countries · rolling 26 weeks
Contact volume by week
Weekly handled contacts. Weeks 19–22 sit above the trend line. Hover for values.
Emerging issues
Clusters detected below the reporting category level, ranked by rate of change rather than volume.
| Issue | Severity | Contacts / wk | 26-wk total | Δ vs 4-wk avg |
|---|
Top contact drivers
Share of all 847,000 contacts. Darker = higher volume.
View as table
| Driver | Contacts | Share |
|---|
First contact resolution — in-house vs outsourced
Weekly FCR. Both models dip in weeks 19–22; the outsourced estate dips further.
Site performance
All seven sites, ranked by volume. AHT in seconds; FCR and CSAT as measured.
| Site | Model | Contacts | AHT (s) | FCR | CSAT |
|---|
Interactive — switch tabs, hover the charts, toggle dark mode. All figures synthetic.
This is Bridworth Group. Seven sites, five countries, 847,000 contacts across a rolling twenty-six weeks.
None of it is real. Bridworth does not exist, the sites do not exist, and every figure on that page came out of a seeded script. I want to be blunter about this than I was in April, because the disclosure is not a footnote — it determines what the post can and cannot claim.
What the artefact demonstrates is form: that a specific analytical shape — volume trend, an emerging-issues table ranked by rate of change rather than by volume, driver mix, a delivery-model split, a site table — can be specified and rendered quickly. That is a real and checkable claim, and you are looking at the evidence for it.
What it does not demonstrate is accuracy. Synthetic data cannot be wrong. There is no ground truth to miss, no mislabelled disposition, no site that codes “resolved” differently from the other six. The hardest part of real contact centre analysis is absent from this demo by construction, and any post that shows you a synthetic dashboard and implies it has proven something about real transcripts is selling you something.
The single most useful thing in the layout is the emerging-issues table. Not because it is clever, but because it is ranked by rate of change against a four-week baseline instead of by volume. A payment failure clustered to one acquirer BIN range runs at 412 contacts a week — it would sit nowhere near the top of a volume report, and it is the most expensive thing on the page.
2. What it costs, with the arithmetic shown
In April I said the model cost was “a rounding error.” That is the kind of phrase that sounds like analysis and is actually just a vibe. Here is the arithmetic instead.
Assume a full pass over all 847,000 contact transcripts. Assume roughly 1,200 input tokens per transcript — a mid-length voice contact, transcribed — and 150 output tokens for a structured triage response. That is about 1.02 billion input tokens and 127 million output tokens.
At July 2026 published list prices:
| Model tier | $/M in | $/M out | Full pass over 847k contacts | Per contact |
|---|---|---|---|---|
| Hosted open-weight, small | 0.03 | 0.12 | $46 | $0.00005 |
| Cheapest proprietary “lite” tier | 0.10 | 0.40 | $152 | $0.0002 |
| Small proprietary | 0.75 | 4.50 | $1,334 | $0.0016 |
| Small proprietary (alt vendor) | 1.00 | 5.00 | $1,652 | $0.0020 |
| Mid-tier reasoning | 1.25 | 10.00 | $2,541 | $0.0030 |
| Mid-tier reasoning (alt vendor) | 2.00 | 10.00 | $3,303 | $0.0039 |
| Frontier | 2.50 | 15.00 | $4,447 | $0.0053 |
I have deliberately not attached vendor names to those rows. The point is the shape of the curve, not who is currently sitting on which rung, and the rungs move every few weeks. Substitute your own provider’s current rate card and your own token estimate; the method survives the substitution, which is the test of whether it was a method or an advert.
Two things the table hides, and they matter more than the headline:
Transcription is not in it. If you are starting from audio rather than text, speech-to-text is very often the larger line item, and it does not fall at the same rate as inference. Anyone quoting you a per-contact analysis cost that excludes ASR is quoting you half a number.
One pass is not one pass. In practice you iterate — reprompt, re-cluster, re-run a subset. Three to five effective passes over the corpus is a more honest planning assumption than one, which moves the frontier row from $4,447 to something between $13k and $22k. Still not fourteen weeks of a programme. But not $46 either, and the gap between those two numbers is where most of the overclaiming in this space lives.
3. The comparison I should have sourced
The honest version of my April claim is narrower than the one I made.
Published 2026 implementation guidance for a large enterprise BI deployment (500–5,000 users) puts the deployment and setup phase at $690,000 to over $2.1 million, with data setup alone at $150,000–$500,000 and report development at $250,000–$750,000. Mid-sized deployments of 100–500 users put the planning and assessment phase at $40,000–$120,000. Large migrations of 150–500+ dashboards are given as 6 to 12 months. (Entrans, Power BI Implementation Cost Breakdown, 2026)
I will flag the obvious: that source is an implementation consultancy, and implementation consultancies have an interest in both the largeness of the number and the manageability of the number. Treat it as an order of magnitude, not a benchmark. I could not find a vendor-independent figure I would stand behind, and I would rather say that than dress up a plausible one.
The comparison also is not like-for-like, and in April I let the reader assume it was. A BI programme delivers a governed, permissioned, refreshing, auditable system with lineage back to source. Thirty minutes on a phone delivers a picture. Those are different objects. The correct claim is not “this replaces that.” It is: the exploratory phase — the part where you are still working out which questions are worth building a system around — has become nearly free, and the build phase has not.
That is a smaller claim than the one my April headline made. It is also the one I can defend.
4. Method, without the model names
The pipeline shape has held up over three months, which is more than the model names did:
- Triage at volume. The cheapest model that can reliably follow a structured output schema, run over every transcript, producing one small record each: intent, sentiment, resolution state, entities. This is a classification job, not a reasoning job, and paying frontier prices for it is the most common cost mistake I see.
- Cluster and correlate. A mid-tier model over the records, not the raw transcripts. This is where the emerging-issues table comes from — grouping below the level of the existing category taxonomy, which is usually where the interesting things hide, because the taxonomy was designed before the problem existed.
- Synthesise. A frontier model, over the clusters, once. This is the smallest token spend of the three and the one worth paying for, because narrative judgement is what you cannot get from the cheap tier.
That shape is vendor-agnostic. It works on any provider with three price tiers and structured output, and it works on self-hosted open-weight models at the bottom tier, which is where the volume cost sits. If a piece of advice only works on one company’s stack, it was not advice.
5. Where this fails
Three months of using this on real engagements rather than airport demos, and the failure modes are consistent:
Silent miscategorisation. The model produces a confident, well-structured, plausible record for a transcript it misread. At 847,000 records nobody is checking, and there is no error bar on the output. You need a human-labelled sample — a few hundred contacts — to measure the triage step against, and if you have not built that, you do not know your error rate. You have a dashboard that cannot tell you it is wrong.
Taxonomy drift. Re-run the clustering next month and the clusters come back slightly different. Excellent for discovery. Useless for a trended KPI. These are different jobs, and the moment someone puts a model-generated cluster on a board pack as a tracked metric, you have a problem.
The governance gap. Gartner predicted in February 2024 that 80% of data and analytics governance initiatives would fail by 2027, on the basis that governance which does not enable prioritised business outcomes fails. (Gartner, February 2024) Making analysis cheap does not touch that. If anything it makes it worse, because the constraint that used to force prioritisation — cost — has been removed, and nothing has replaced it.
Data access, which is the actual bottleneck. Nothing above helps if the transcripts sit behind a vendor API with a rate limit, a retention policy that deletes at ninety days, and a contract that is ambiguous about sending them to a third-party model. In most real estates this is the binding constraint, and it is a commercial problem, not a technical one.
A note on the statistics I am not using. The “85% of big data projects fail” figure that circulates in every deck of this kind traces back to a 2017 Gartner statement and has been recycled without provenance ever since; there is a whole genre of write-ups documenting how thin the sourcing is on it and its cousins. It would have supported my argument nicely. That is precisely why I left it out.
6. Where I have landed
The April version of this post was true in its direction and loose in its detail, which is a description of most writing about AI capability right now, including the writing that is trying hard to be careful.
What has actually changed in my work is narrower than “everything.” The cost of asking has collapsed. The cost of knowing whether the answer is right has not moved at all — it is still a human-labelled sample, still a domain expert reading a hundred transcripts, still someone who knows that Glasgow codes escalations differently to Lisbon. The scarce resource has moved one step down the chain, from analysis to verification.
The dashboard at the top of this page took about thirty minutes. Establishing that a version of it built on real Bridworth data was correct would take considerably longer than thirty minutes, and that is the honest headline. It is a worse headline. It happens to be the one I can show you the working for.
Mark McLeish writes about contact centre architecture, enterprise AI and customer experience. Views are his own. Not Accenture’s. Practitioner signal. No vendor PR.
Figures in the cost table are July 2026 published list prices and will date; the arithmetic is shown so you can re-run it against current rates. The dashboard is synthetic throughout.