The Brief The AI Layer · Last updated July 2026
Speech analytics: what CallMiner, Verint, and the rest actually surface.
Speech analytics has been sold to contact centers since 2010. In 2026, CCaaS-native QA and real-time agent assist have absorbed most of the per-call use-case. What the category still uniquely delivers is aggregation — cross-call pattern detection, fleet-scale compliance monitoring, and structured discovery data for teams outside the contact center.
7 min · The AI Layer · Series piece 7
Questions this article answers
- What does standalone speech analytics do in 2026 that CCaaS-native QA and real-time agent assist do not?
- Where do the categories overlap, and where has native closed the gap?
- Where does a specialist still win outright for a 30-500-seat operator?
- How does per-audio-hour pricing compare to per-agent-per-month once you plug in real volume?
- Which vendors are in the lineup, and what does each one actually do?
- What breaks after signature?
The category predates the AI Layer, and that is the whole problem. What speech analytics still uniquely delivers is not per-call insight. It is aggregation.
Speech analytics platforms — CallMiner Eureka, Verint Speech Analytics, NICE Interaction Analytics, Genesys Speech and Text Analytics, Observe.AI, and a long tail behind them — were sold as a discrete category from roughly 2010 onward. The pitch was mining recorded calls for coaching signal a supervisor could never surface by hand. Fifteen years later, that pitch is largely obsolete. The CCaaS-native quality module ships with automated scoring. The real-time agent assist layer catches the coachable moment before the call ends. DMG Consulting, which has covered this market for two decades, renamed the annual report from “Speech and Text Analytics” to Conversation Analytics for the Enterprise — an analyst-side signal that the standalone label no longer describes the buy. What speech analytics still uniquely delivers is aggregation: cross-call pattern detection over millions of interactions, compliance keyword monitoring at fleet scale, root-cause correlation into churn and CSAT, and structured discovery data pushed back to product, marketing, and compliance teams that never touch the contact center.
What per-call agent assist cannot do is aggregate
Real-time agent assist and CCaaS-native QA operate on the interaction in front of them. Both are excellent at what they do — the next-best-response prompt during the call, the automated evaluation form after it — and both are the right tool when the question is “how did that specific call go.” Neither produces the artifact a product manager wants, which is: across 400,000 calls last quarter, which three phrases correlate with the churn signal, and which coaching cohort produces them. Aggregation is a different engineering problem. It requires 100 percent transcription at usable accuracy, a taxonomy that survives category drift, and a query surface a non-analyst can drive. That is the speech-analytics category's actual home ground, and it is the one thing native modules mostly still stub.
Where the categories overlap — and where native has quietly won
The overlap is the coaching use case. If your requirement is score every call, flag the outliers, generate a summary for the supervisor, and hand the agent a prompt on the next one — the NICE Interaction Analytics module inside CXone, the Genesys Cloud native tools, and the Observe.AI post-call review will all clear the bar. For a 30-500-seat operator whose top requirement is per-agent coaching, the CCaaS-native module in a tier you likely already pay for is enough. Where native has not closed the gap is depth on regulated content, cross-vertical benchmarking, and the analyst workbench a business analyst uses to build a new category from scratch on Monday morning and ship a finding to marketing on Friday. That is still a specialist workload, and it is why Verint's installed base is dominated by aggregate-view use cases — compliance, churn correlation, product signal — not per-call coaching.
Where standalone speech analytics still wins outright
Three workloads still justify the specialist premium. Regulated compliance keyword monitoring at fleet scale — the collections script disclosure, the financial-advice disclaimer, the healthcare pre-authorization language — where the fine for a false negative outruns the annual license by an order of magnitude. Root-cause discovery that connects contact reason to downstream churn, CSAT, or repeat-contact rate — the analysis a BI team cannot run because the source data lives as audio, not rows. Structured discovery data for teams outside the contact center — product managers looking for objection language, marketing looking for competitive mentions, risk looking for fraud patterns. In each case the buyer is not the WFO team. It is compliance, product, or executive leadership, and the deliverable is a dataset the CCaaS-native module was not designed to produce.
The pricing math — per audio-hour or per agent, and never both
Speech-analytics vendors and CCaaS-native modules do not price on the same unit, and that is the single most consequential fact in the category. Specialists — CallMiner, Verint, NICE Interaction Analytics as a standalone SKU, Sestek, VoiceBase — historically price per hour of audio ingested, or per named seat. CCaaS-native modules and the newer per-agent conversation-intelligence buys — Genesys native, Observe.AI, Level AI, Cresta — price per agent per month. Cresta and Level AI extend into auto-QA on a per-seat basis. Deepgram sits underneath everyone as raw ASR, priced per minute. A per-audio-hour quote at $1.20 per hour looks trivial next to a $60 per-agent-per-month subscription — until you multiply the audio-hour rate by 100 agents × 6 productive hours × 21 working days and land at a number that is triple the per-agent quote. The bill inverts again for shops with high part-time or seasonal headcount, where per-agent looks punitive and per-hour looks cheap. DMG's 2025-2026 report benchmarks a 250-seat cloud deployment as the reference pricing case — worth reading before any RFP.
The vendor lineup, at capability
CallMiner (Eureka) — the reference platform for aggregate discovery and analyst workbench; named a Leader in the Forrester Wave Conversation Intelligence Solutions for Contact Centers Q2 2025; priced historically per audio-hour; strong on category creation and BI integration. Verint (Speech Analytics) — largest installed base per DMG; deepest compliance and financial-services tooling; Genie Bot GenAI layer on top; priced enterprise-negotiated per hour or per seat. NICE (Interaction Analytics, inside Enlighten) — tightly integrated into CXone; the default choice when NICE is already the CCaaS; available as a WEM module tier. Observe.AI — per-agent-per-month conversation intelligence, strongest for mid-market coaching-plus-QA bundles under 300 seats. Genesys (Speech and Text Analytics, native) — bundled inside Genesys Cloud tiers; the “already paying for it” option for Genesys shops. Cresta — real-time agent assist and post-call review; per-seat. Level AI — auto-QA plus speech analytics on a per-seat model; oriented to sub-500-seat operators. Sestek — multilingual specialist; DMG-covered; strong in EMEA and Turkish-language deployments. VoiceBase — cloud-native ASR plus analytics API, often embedded upstream. Deepgram — the raw ASR backbone many of the above (and many CCaaS platforms) rent from underneath, per-minute pricing.
The buyer artifact — a 6-question fit-to-shape rubric
The speech-analytics fit-to-shape rubric
- Coverage gap. Does the CCaaS-native module or agent-assist tool we already pay for cover per-call QA and coaching adequately? If yes, you are not buying a coaching tool — you are buying aggregation.
- Compliance table stakes. What keyword-and-phrase monitoring is required for our vertical (collections FDCPA, HIPAA, MiFID II, TCPA)? Write the list. If the list is short, native scoring may clear it.
- Volume floor. Do we ingest more than roughly 1 million voice minutes per month? Below that, aggregate pattern detection is statistically thin and the ROI narrative collapses.
- Insight consumer. Who reads the output — WFO, product, compliance, marketing, exec? If the only consumer is the WFO team, native is almost always the right buy.
- Integration surface. Which CRM, BI, and data-warehouse targets need the output, and what native connectors exist? An export-only vendor multiplies your data-engineering cost.
- Pricing unit. Are we quoted per audio-hour ingested or per agent per month? Model both against 2× current volume. The unit that looks cheapest at pilot is rarely cheapest at production.
What breaks
Wrong tool for the actual need. The most common failure is buying speech analytics when the ops team's real complaint is per-call coaching latency. The right buy there is real-time agent assist, not a $200k aggregation platform whose insight the WFO team will not have the analyst headcount to consume.
The keyword-list decay problem. Compliance categories drift. Product names change, disclosure scripts get rewritten, agents route around flagged phrases. Without an owner who curates the taxonomy quarterly, false-positive rates climb, supervisors stop trusting the alerts, and adoption collapses within twelve months. Every DMG-covered vendor has a story about this. Solve it in the SOW, not the demo.
Audio-hour pricing at growth. Per-audio-hour scales linearly with call volume, and call volume is a function of your business's success. A per-hour rate that penciled at last year's volume can double the invoice against a surge quarter or a new product launch — with no corresponding license reset. Cap the rate at 2× projected volume before signature or convert to a committed-use annual pool.
Three rules
Buy for aggregation, not for coaching. If the deliverable is a per-call score, the CCaaS-native module or per-agent assist tool wins on cost and integration every time. Reserve the specialist premium for cross-call pattern detection and fleet-scale compliance.
The pricing unit is the whole decision. Per audio-hour and per agent per month never compare on their quoted numbers. Model both at real volume against a 2× growth scenario before scoring a single feature.
Name the consumer before the RFP. If you cannot name the compliance officer, product manager, or executive who reads Monday's report, standalone speech analytics is a tool without a user. Buy the seat.
In short
- Speech analytics no longer owns per-call coaching. CCaaS-native QA and real-time agent assist have absorbed that use case. What the category still owns is aggregation.
- The unique deliverable is fleet-scale pattern detection, regulated compliance monitoring, and structured discovery data for teams outside the contact center.
- Pricing splits per audio-hour (CallMiner, Verint, NICE, Sestek, VoiceBase) versus per agent per month (Observe.AI, Cresta, Level AI, Genesys native). Modeling both against real volume is non-optional.
- Below roughly 1M voice minutes per month, the aggregate view is too thin to justify the standalone spend.
- DMG renamed the annual category report from “Speech and Text Analytics” to “Conversation Analytics for the Enterprise” — the analyst-side signal that the standalone label no longer describes the buy.
Sources
- CallMiner, Eureka product overview. callminer.com
- Verint, Speech Analytics product page. verint.com
- NICE, Interaction Analytics (part of Enlighten). nice.com
- Genesys, Speech and Text Analytics. genesys.com
- DMG Consulting, “2025-2026 Conversation Analytics for the Enterprise” (formerly Speech and Text Analytics). dmgconsult.com
All linked sources were live at time of publish (July 2026). Verify before quoting in a procurement document.
Deciding between a specialist buy and the CCaaS-native module?
Run the Tier 1 benchmark.
Submit your CCaaS contract and a month's worth of voice minute counts by queue. We return a benchmark PDF in five business days showing whether native scoring clears your compliance table stakes, where a specialist aggregation platform actually pays back, and which pricing unit survives at 2× your projected volume. Free. No follow-up sales drip.
Run the Tier 1 benchmark →Series · The AI Layer
Read the anchor, Where AI actually fits in your tech stack. See the sibling piece on real-time agent assist, and the platform framing in The AI layer on top of CCaaS. Category hub: CCaaS vendor selection.