Case 09
AI Infrastructure Optimization for a Voice AI Product
Voice AI startup · production LLM & speech stack · 1 month
Where it stands
Measured and live- 75.6%lower memory-generation cost
- ~26%premium-model cost reduction from caching alone
- ~23%further validated opportunity in speaker processing
The challenge
A production voice-AI stack was generating conversational memories at scale, but its inference architecture had grown organically. Premium models were used at stages where they were not always necessary, prompt caching was misconfigured, speech and speaker processing created separate cost lines, and AI traffic was spread across multiple backend paths with limited visibility into cost per feature.
What we built
04 parts- 01
Simplified summarization pipeline
Restructured a fourteen-plus-step processing path into a tighter three-stage summarization flow, eliminating unnecessary AI calls before changing models.
- 02
Corrected prompt caching
Fixed caching behavior and added monitoring, cutting premium-model cost per call before larger model-routing changes were required.
- 03
Model routing based on quality
Moved appropriate workloads to lower-cost and open-source models only where measured quality held.
- 04
Single AI gateway and observability
Centralized AI traffic through one gateway so cost per feature remains visible after handover and savings do not quietly erode.
The results
- Memory-generation cost reduced by 75.6% on matched production windows.
- Approximately 26% reduction in premium-model cost per call from caching alone.
- A further ~23% cost opportunity validated in speaker processing.
- No user-facing product change was required to realize the primary savings.
AI costs growing faster than usage economics can support?
Scope yours in a 30-minute call48-hour reply
Let's build something that actually works.
Tell us where you are and what you need. We’ll come back with a clear, honest plan within 48 hours.