Call Transcription to Agent Memory: A Reliable Pipeline
Turn completed transcript segments into scoped, traceable memory. Handle live audio, revisions, retries, speaker identity, and costs separately.

Turning call transcripts into agent memory requires a reliable path from audio to completed text, then from text to scoped records the application can retrieve later. Treat transcription, extraction, persistence, and recall as separate stages. A live audio connection is not itself a durable memory store.
The central design choice is when a statement becomes eligible for reuse. Partial speech, revised transcripts, uncertain speakers, and duplicate events should not silently become permanent customer facts.
Separate live interaction from recorded-audio processing
Google's Live API capabilities guide documents streaming audio and transcription of input and output. Choose a currently supported model and configuration for that interface. Do not assume a recorded-audio example for a named Flash model can be used unchanged as a live-call pipeline.
A recorded-call workflow can wait for a complete file, transcribe it, review the result, and index the accepted text. A live agent needs to handle incremental events, interruptions, reconnects, and the timing of its next response. Those are different latency and state-management problems.
Session resumption can help continue a live connection, but application memory must survive beyond that connection's lifecycle. Keep the durable call identity in your application.
Use identities that survive reconnects
Map an authenticated tenant and customer to a call ID, then to transcript segments and revisions. A provider connection ID may change when the stream reconnects. If that becomes your memory identity, one call can fragment into unrelated records.
Speaker labels require care. “Speaker 2” is not automatically the same person across two calls. Resolve participants through trusted meeting or account information where possible, and preserve uncertainty when you cannot identify the speaker.
For downstream memory, attach the call, segment, revision, speaker attribution, and time reference to each extracted claim. A future answer about a commitment should be able to return to the supporting utterance.
Accept completed revisions before extraction
The following Python example operates on application-normalized events. Your adapter supplies final, revision and segment after interpreting the transcription provider's events.
def apply_segment(state, event, authorized_tenant):
if event["tenant"] != authorized_tenant:
raise PermissionError("Wrong tenant")
if not event["final"]:
return False
key = (event["tenant"], event["call"], event["segment"])
previous = state.get(key)
if previous and event["revision"] <= previous["revision"]:
return False
state[key] = {
"revision": event["revision"],
"text": event["text"],
"speaker": event["speaker"],
}
return True
state = {}
base = {
"tenant": "acme", "call": "call-1", "segment": "s1",
"revision": 1, "final": True, "speaker": "customer-7",
"text": "Please send the draft on Friday.",
}
assert apply_segment(state, base, "acme")
assert not apply_segment(state, base, "acme")
assert apply_segment(state, {**base, "revision": 2,
"text": "Please send the draft on Thursday."}, "acme")
assert state[("acme", "call-1", "s1")]["revision"] == 2
The reducer accepts newer completed revisions and suppresses repeated deliveries. It uses in-memory state for clarity. A production worker needs durable storage and a transaction or conditional write so concurrent consumers cannot overwrite a newer revision with an older one. Re-extract or invalidate derived claims when an accepted segment changes.
Reject malformed events at the adapter boundary. Treat an empty corrected segment according to an explicit removal policy rather than leaving its previous extracted fact active.
Keep statements, commitments, and summaries distinct
“Could you send the proposal Friday?” is a request. “I will send it Friday” is a commitment. A summary that converts the first into the second has changed the business meaning even if its prose reads smoothly.
Store the original segment reference alongside the extracted type, participants, and any resolved date. Preserve the call's time zone when turning “tomorrow” into a date. When attribution or wording is uncertain, keep the claim provisional instead of turning it into a confident customer profile fact.
If a statement triggers an external action, use the application's action rules. A transcript is evidence of speech, not blanket authorization for an agent to send messages or modify an account.
Retrieve before the next relevant turn
Use the verified customer and account scope to load relevant prior context before generating a response. Keep recent live dialogue in working context while durable memory supplies selected prior facts. Avoid reloading every old transcript for every utterance.
Measure the stages separately: audio arrival to usable transcript, transcript acceptance to searchable memory, retrieval latency, and time to the agent's audible response. A retrieval benchmark does not establish the total delay a caller experiences.
For a later support call, retrieve the previous commitment and its source, then check current fulfillment status in the owning system. See the customer-support architecture for that separation.
Model cost from the actual workload
For a fictional workload of 1,000 calls averaging 12 minutes, the input contains 12,000 audio minutes. At an illustrative transcription rate of $0.01 per minute, that component would be $120.
Add extraction, storage, retrieval, answer generation, retries, and any retained audio. If you budget two retrievals per call, that is 2,000 requests; a design that retrieves on every turn may have a very different total. Use your provider's current billing units rather than assuming all audio models charge per minute or share the same audio-token conversion.
Summarization can reduce later context, but measure what it removes. An apparent token saving that drops exact amounts, conditions, or commitments can increase support effort. The operating-cost guide provides a broader cost model.
Test the call-to-memory boundary
Replay synthetic events with partial transcripts, duplicate deliveries, a corrected final segment, a reconnect, an unknown speaker, and a different tenant. Verify the accepted record before evaluating an answer.
Then run a controlled audio pilot with representative accents, terminology, interruptions, and recording conditions. Check exact dates, amounts, and speaker assignments rather than relying only on a readable summary. Define retention and deletion across audio, transcript, extracted records, and caches using the memory lifecycle guide.
The first milestone is one call whose accepted commitment can be retrieved in a later session, traced to the right speaker and passage, and corrected without the old version quietly returning.
For a Pipecat application, use the voice-memory integration guide to place retrieval before answer generation.
To test the persistent-memory stage, start with Supermemory and one synthetic call transcript. Save an accepted commitment, retrieve it in a new session, and correct it. Once that path works, connect it to your audio pilot and measure the complete experience.