Skip to content
Back to Blog

Your Call Logs Are the Only Training Data That Actually Matters

Bashar AyyashJuly 9, 20267 min read1,387 words
Your Call Logs Are the Only Training Data That Actually Matters
TL;DR

Real call recordings are the essential training data for effective enterprise AI agents, revealing the messy reality of how humans actually speak.

7 min read · 1,387 words

The Dirty Secret of Enterprise AI

Every AI agent demo looks perfect.

Clean transcripts. Happy paths. Customers who speak in complete sentences and never interrupt themselves. I've sat through fifty of these pitches at Alrajhi Bank. The gap between demo and production is a canyon.

Here's what actually happens when your banking AI agent hits real users in Riyadh, Amman, or Cairo: code-switching between Arabic dialects, background noise from busy streets, customers who don't know their own account types, edge cases that break your intent classification in ways you couldn't imagine.

Your synthetic training data? It's architectural fiction.

Why I Started Recording Everything

Two years ago we built our first voice agent for balance inquiries and transfers. Standard stuff. Laravel backend, Next.js admin panel, React Native for the mobile interface.

We trained on 10,000 synthetic conversations. Beautiful dataset. Grammatically perfect. Covered every feature flag in our spec.

Week one in production: 34% containment rate. The agent couldn't handle a customer who said "habibi I need my money" instead of "transfer funds." Couldn't parse the Jordanian dialect "shu" as a question marker. Crashed on overlapping speech when someone's kid grabbed the phone.

We were flying blind. No recordings, no telemetry, just logs saying "intent_classification_failed."

That failure taught me something I've applied to every AI project since: the conversation is the infrastructure. Not the model. Not the prompt engineering. The raw, messy, recorded reality of how humans actually talk to machines.

What "Recording Everything" Actually Means

I'm not talking about compliance recording for legal teams. That's table stakes in banking.

I mean engineering-grade capture:

  • Full audio streams with speaker diarization (who spoke when)
  • Real-time transcriptions with confidence scores per word
  • Latency metrics for every turn in the conversation
  • Context state snapshots at each decision point
  • Model outputs with their probability distributions
  • Fallback triggers and the exact user utterance that caused them

In our Laravel backend, we built a pipeline that treats every call as a structured event stream. Audio chunks flow through Redis streams, get processed by Whisper for transcription, then land in ClickHouse for analysis. The React Native client uploads compressed audio with session metadata. Our Next.js dashboard lets engineers replay conversations with full context reconstruction.

This isn't Big Brother surveillance. It's observability for conversational systems.

The MENA Reality Check

Working from Jordan, I see how Western AI assumptions crumble in emerging markets.

Your average Saudi customer might start a conversation in formal Arabic, switch to Najdi dialect when frustrated, drop English technical terms, then code-mix with Urdu if they're South Asian. Try training your standard OpenAI or Anthropic model on that without real recordings.

Network conditions matter too. Our React Native app sees call quality ranging from pristine 5G in Riyadh financial district to degraded audio on 3G in rural Jordan. Without recorded samples across this spectrum, your noise robustness is theoretical.

And don't get me started on cultural friction patterns. In some MENA markets, customers repeat security information slowly because they don't trust the system understood. In others, they speak fast to test if the agent is actually listening. These behaviors don't appear in synthetic data. They only emerge when you have walls of real calls to analyze.

How We Use Recordings to Build Better Agents

Our current workflow at Alrajhi looks like this:

Week 1-2: Shadow Mode

New agent runs alongside human agents. We record everything but don't automate responses. The AI generates replies, we compare against what humans actually said, measure divergence. This catches intent misclassification before customers see it.

Week 3-4: Graduated Rollout

We enable the agent for 5% of traffic, but with aggressive escalation triggers. Every escalation gets tagged with the audio snippet that caused it. Our Laravel pipeline automatically surfaces clusters: "agent failed on balance inquiries over 1M SAR," "confusion between 'transfer' and 'payment' verbs."

Month 2+: Continuous Mining

We run weekly jobs that:

  • Extract high-entropy utterances (low confidence, multiple intent matches)
  • Find conversation patterns that led to human takeover
  • Identify successful repair sequences (where agent recovered from misunderstanding)
  • Build synthetic training data from real patterns, not imagined ones

This last point matters. We don't use raw recordings for training (privacy, compliance, PII). We use them to understand what to synthesize. The recordings tell us where our synthetic data was wrong.

The Technical Architecture

For engineers who want specifics, here's our stack:

Ingestion Layer

  • Twilio/Media Streams for telephony audio
  • React Native SDK with custom audio compression (Opus, 24kbps)
  • Laravel queue workers for stream routing

Processing Pipeline

  • Whisper v3 for transcription with fine-tuned Arabic dialect support
  • Speaker diarization via pyannote.audio
  • Real-time intent classification with our own fine-tuned BERT variant

Storage & Retrieval

  • Audio: S3 with lifecycle policies (30 days hot, 1 year glacier)
  • Metadata: ClickHouse for fast analytical queries
  • Search: Elasticsearch for transcript full-text + phonetic matching

Analysis Tools

  • Next.js dashboard with conversation replay
  • Jupyter notebooks for pattern mining
  • Automated alerting on containment rate drops

The whole system adds about 200ms latency to calls, which is acceptable for banking use cases. For high-frequency trading or real-time payments, we'd need edge deployment. Trade-offs everywhere.

What You'll Actually Find

After six months of recording, here's what we discovered about our "simple" balance inquiry agent:

  • 23% of calls included at least one non-task utterance (greeting, complaint about previous service, small talk)
  • Edge case explosion: 147 distinct ways customers asked for "recent transactions," none appearing in our original training set
  • Recovery patterns: The most successful agent recoveries used acknowledgment phrases specific to Saudi Arabic politeness norms, which our synthetic data completely missed
  • Silent failures: 8% of calls where the agent gave technically correct but practically wrong answers (customer asked for "last transfer," agent gave last incoming transfer, customer wanted last outgoing)

These aren't bugs you find in testing. They're gravity. They're the weight of real usage pulling your system toward its actual shape.

The Privacy Engineering Problem

I know what you're thinking. Banking. Recordings. Compliance nightmare.

Yes. And it's solvable.

We built PII detection directly into the Laravel pipeline. Audio streams get scanned for credit card numbers, account numbers, national IDs before storage. Transcripts are redacted automatically. Access controls are granular: engineers can see patterns, not individual calls, unless specifically authorized for debugging.

In Jordan and Saudi Arabia, regulatory frameworks are evolving fast. We work with legal teams to ensure consent flows are clear. The React Native app explicitly asks for recording permission with specific use case explanations, not blanket "improve our service" language.

The alternative — not recording, not knowing — is worse. You ship blind. You discover failures through customer complaints and regulatory fines rather than engineering analysis.

Before You Ship That Agent

If you're building AI agents for enterprise — especially in regulated industries, especially in complex linguistic environments — here's my checklist:

  1. Recording infrastructure first. Not after. First. Before you train your model, before you write your prompts, you need the pipeline to capture reality.

  2. Synthetic data is a starting point, not a substitute. Use it to bootstrap. Replace it with distributions derived from real recordings as fast as possible.

  3. Measure containment, but also measure confidence calibration. An agent that escalates appropriately is better than one that's confidently wrong.

  4. Build for dialect, noise, and interruption from day one. These aren't edge cases. They're the center of mass in emerging markets.

  5. Create feedback loops where engineers hear the failures. Not dashboards. Actual audio. The emotional weight of a frustrated customer teaches more than any metric.

The Real Question

Most teams building AI agents are optimizing for demo day. Clean transcripts, happy paths, metrics that look good in pitch decks.

The teams that survive production are optimizing for friction visibility. They want to see where the system grinds against reality. They build recording infrastructure not despite the complexity, but because it reveals truth.

We're about to deploy our third-generation agent at Alrajhi. It handles complex multi-turn conversations, integrates with our core banking systems through Laravel APIs, serves customers across three countries. The difference between this generation and our first attempt isn't better models or bigger training sets.

It's 400,000 recorded conversations showing us exactly where we were wrong.

So here's my question for you: What's your plan for capturing reality? Or are you still building in the clean room, hoping the world will cooperate with your assumptions?

Bashar Ayyash
AUTHOR

Bashar Ayyash (Yabasha)

AI Systems Architect for regulated industries — evals, harness design, AI security.

Bashar Ayyash is an AI engineer and dev lead in Amman, Jordan. 20 years shipping software, 4 years inside a regulated bank building production RAG and agent systems with evals, guardrails and monitoring — in Arabic and English. He writes at yabasha.dev and builds open-source tooling for AI-assisted development.

Newsletter

Practical AI + full-stack insights for MENA builders. No spam.

Read more on the blog

Browse the latest articles or explore the full archive.