Skip to content

Case 09

AI Infrastructure Optimization for a Voice AI Product

Voice AI startup · production LLM & speech stack · 1 month

Where it stands

Measured and live
  • 75.6%lower memory-generation cost
  • ~26%premium-model cost reduction from caching alone
  • ~23%further validated opportunity in speaker processing

The challenge

A production voice-AI stack was generating conversational memories at scale, but its inference architecture had grown organically. Premium models were used at stages where they were not always necessary, prompt caching was misconfigured, speech and speaker processing created separate cost lines, and AI traffic was spread across multiple backend paths with limited visibility into cost per feature.

What we built

04 parts
  1. 01

    Simplified summarization pipeline

    Restructured a fourteen-plus-step processing path into a tighter three-stage summarization flow, eliminating unnecessary AI calls before changing models.

  2. 02

    Corrected prompt caching

    Fixed caching behavior and added monitoring, cutting premium-model cost per call before larger model-routing changes were required.

  3. 03

    Model routing based on quality

    Moved appropriate workloads to lower-cost and open-source models only where measured quality held.

  4. 04

    Single AI gateway and observability

    Centralized AI traffic through one gateway so cost per feature remains visible after handover and savings do not quietly erode.

The results

  • Memory-generation cost reduced by 75.6% on matched production windows.
  • Approximately 26% reduction in premium-model cost per call from caching alone.
  • A further ~23% cost opportunity validated in speaker processing.
  • No user-facing product change was required to realize the primary savings.

AI costs growing faster than usage economics can support?

Scope yours in a 30-minute call

Next caseUnified Client Data & AI-Powered Session Prep

48-hour reply

Let's build something that actually works.

Tell us where you are and what you need. We’ll come back with a clear, honest plan within 48 hours.