Connecting an AI Agent to Phone Calls, Setup and Real Call Recordings
After messaging and Notion, I connected phone calls last. Messaging requires typing, and waiting for results while the agent works was annoying. Just speaking seemed faster. (measured on the operator's environment)
Structure
My number β Twilio β (Media Stream) β Vapi β My server API
β Agent / command execution
When a call comes in, Twilio streams audio to my server via websocket. The server passes it to Vapi, which converts speech to text, feeds it to a model, and synthesizes the reply back through Twilio. It is not one call moving whole β audio chunks flow in real time.
Setup steps
1. Get a Twilio number
Buy a number in the Twilio console. Korean numbers require business registration and are tedious; a US toll-free number activates immediately for testing.
2. Create a Vapi assistant
{
"name": "Hermes Phone Agent",
"model": {
"provider": "openai",
"model": "gpt-4o",
"messages": [
{ "role": "system", "content": "You are a Korean-speaking assistant. Answer briefly and clearly." }
]
},
"firstMessageMode": "assistant-speaks-first",
"assistantFirstMessage": "Hello, how can I help you?"
}
Pitfall 1: Without the first message, the call never starts
If assistantFirstMessage is empty, the call connects but stays silent until the other party speaks. If a call connects with no sound, check this first.
3. Expose your server
cloudflared tunnel --url http://localhost:8080
Pitfall 2: The tunnel URL changes every restart
Temporary tunnels assign a new address on each restart. You must re-register the webhook URL in Vapi every time. For production, use a fixed domain.
4. Register the webhook URL in Vapi
Paste the tunnel URL into the assistant's Server URL field.
5. Connect tools
{
"type": "function",
"function": {
"name": "run_command",
"description": "Run a command on the server",
"parameters": {
"type": "object",
"properties": {
"command": { "type": "string" }
}
}
}
}
Pitfall 3: Without tools, it is just a support call
The conversation works but nothing happens. The whole point of phone integration is tool connection.
6. Connect the Hermes side
Hermes has no phone-specific subcommand (hermes whatsapp and hermes whatsapp-cloud are for WhatsApp). I attached a message adapter to Vapi through the gateway. The adapter only needs three things:
Inbound: Vapi webhook β normalize message β Hermes session
Outbound: Hermes tool call β Vapi function response format
End: call termination β save summary to Notion
Three real call recordings
Call 1 β Lookup
Bot: Hello, how can I help you? Me: How many posts did I publish yesterday? Bot: One moment. Me: Can't you do it? Bot: Four posts were published yesterday. Telegram, Discord, LINE, and WhatsApp setup guides.
The lookup took about 3 seconds. In a phone call that feels long. It did not sound like a human answering β more like someone thinking for a moment.
Call 2 β Task execution
Me: Run the build and tell me the status. Bot: Connected to the server. Running now. Me: OK. Bot: Done. Me: ? Bot: The build finished in 4 minutes 12 seconds. 112 posts generated. No errors.
Waiting 4 minutes on a phone call was ambitious, but hearing the tool actually execute on the server was the point.
Call 3 β Interrupting mid-sentence
Bot: The setup has three steps. The first step is creating the internal connection. The second isβ Me: Wait, just summarize the three steps. Bot: Sure. Internal connection creation. Page sharing. MCP server registration.
It stopped on my interrupt and answered immediately. This is the advantage of phone calls β you can redirect the conversation without typing.
Where things broke
Pausing mid-sentence is just how it sounds
Large models think in seconds and stream audio in chunks, so it sounds choppy. Streaming ASR + LLM + TTS pipelines help, but latency is real. Writing the first sentence longer ("Yes, checking now") and sending the actual answer after is the practical fix.
Silence causes disconnection
If you don't speak for a few seconds, the call drops. Noise in the room makes it worse β the ASR sometimes thinks you spoke and drops the call.
Longer calls cost more
Per-minute cost stacks: call charges + ASR + LLM + TTS. Phone calls should be for quick "check this now" requests, not long sessions.
Testing with noise was painful
Typing while testing made ASR garble everything. Being too quiet caused the call to drop from breath sounds. There is a narrow window that works.
Summary
Twilio β Vapi β server tools β spoken result
Phone calls remove the need for typing and a screen. They only work in quiet environments. I now use phone for "check this now" instead of messaging. Initially I saved a contact and kept calling. Now I call the agent.
AI Knowledge Hub