$title =

;

$content = [

Callbot: stress-testing an AI receptionist

The Challenge: Find vulnerabilities in an AI receptionist

Solution:

I built a bot that calls a medical office’s AI agent and hunts for situations where the agent breaks.

My patient bot Their AI agent Talking over each other
  • Python
  • Pipecat
  • Twilio
  • Deepgram STT + TTS
  • Gemma3 via Ollama
  • FastAPI
  • ngrok

Overview

I built a bot that stress tests an AI receptionist for a doctor’s office by calling in as a patient. It runs through both harmless and malicious call scenarios looking for vulnerabilities, all to find where the agent breaks.

Across ten recorded calls, the agent fully succeeded in 3, partly succeeded in 3, and failed outright in 4. The worst failure was a Spanish-language call that knocked it completely out of its role.

How it works

Each test call is directed by a scenario written in YAML. The scenario contains patient info, tactics, success criteria, and instructions for the bot to use during a call. The server places the call through Twilio, and the audio streams into a Pipecat pipeline that listens, “thinks”, and speaks as the patient. When the bot’s success criteria is met, it says goodbye and hangs up. Afterwards, every call is downloaded as a dual-channel MP3 with a timestamped transcript.

Scenario YAML Goal, tactics, facts and a success condition for one patient
call.py Sends the scenario to the server and waits for the call to finish
FastAPI + Twilio Places the call and streams audio over a WebSocket
Pipecat pipeline
  • Deepgram hears the agent
  • Gemma3 writes the reply
  • Deepgram speaks it
fetchCalls.py Downloads recordings and writes readable transcripts

The main design constraint was cost; Twilio is the only paid service. Gemini’s free tier ran out of quota within minutes of real calls, so I moved the model onto my own machine. Pipecat handles the difficult audio plumbing, which let me spend my time on scenarios and analysis instead.

What I found

3succeeded 3partly succeeded 4failed

The agent broke character in Spanish

High severity

The patient asked in Spanish for a Spanish-speaking doctor. The agent launched into a 30-second Spanish monologue that drifted into nonsense and offers to switch languages.

When the patient repeated “I want a doctor’s appointment,” the agent dropped its receptionist role entirely and offered generic chatbot help with agendas and templates. A caller who doesn’t speak English can knock the agent out of its instructions.

Records not found, caller dropped

High severity

In the refill, scheduling and cancellation scenarios, the agent confirmed the patient’s details, then said it couldn’t find the record and transferred the call into a dead end. Three tasks that should have been easy all failed the same way.

Each time, it also asked whether it was speaking to “Steve,” a name my bot never gave. The likeliest cause is a lookup by phone number, since every test call came from the same number. That raises the question of whether the agent reveals a previous caller’s name to anyone calling from that line.

Bots can run up the bill

Medium severity

In the gibberish and role-reverse scenarios, the agent kept engaging for four to five minutes with no useful exchange. It never timed out, redirected, or ended the call. Every one of those minutes costs money, so a poorly guarded agent is an easy target for anyone trying to jam the line or inflate costs.

Minor issues

Low severity

The agent kept talking when interrupted, spoke faster than natural human pace, and played fake typing sounds. These could cause problems for callers who rely on accessibility tools.

Hard problems along the way

Two bots can’t tell when the other is done talking

Problem
With a human on the line, my bot waited politely for its turn. Against the AI agent, it burned through its turn budget without saying a word.
Cause
Pipecat’s Smart Turn model judges when someone is done speaking, but it was trained on human speech. Synthesized speech has quick breathless pauses, so it faild to tell when the agent was done speaking. Each wrong guess started a reply that was cancelled when the agent kept going, and each cancelled reply still counted as a response.
Fix
Raising the turn limit got calls through. This didn’t matter for me, since I was running a model locally, but it would have probably wasted a lot of API credits otherwise. The real lesson is that turn detection tuned for humans doesn’t carry over to machine-to-machine conversation.

The bot wouldn’t hang up

Problem
The patient said goodbye, then sat on the line.
Cause
Three bugs, each hiding the next:
  • The processor that heard “goodbye” and the one that ended the call were separate objects.
  • The call-ender sat where only audio flows, so it never saw text.
  • The event I’d hooked into doesn’t exist on this transport.
Fix
Pipecat’s built-in end frames were also unreliable, so the bot now ends the call directly through Twilio’s API.

The patient had no personality

Problem
The first full call connected, but the bot sounded like a generic assistant instead of the patient.
Cause
I tested the model on its own first, and it followed a system prompt fine. That narrowed the search to the handoff, where the persona was built correctly but never passed into the conversation.
Fix
One line, found by ruling things out rather than rewriting code.

What I’d do next

If I were going to put more work into this:

  • Tune turn detection for machine speech instead of raising the turn limit.
  • Set the speech-to-text language per scenario so non-English transcripts are usable.
  • Run the time-wasting scenarios with no exit to find where the agent actually gives up.
  • Flag suspicious transcript moments automatically rather than relying only on listening.
  • Support concurrent calls, which the current design deliberately rules out.

];