--- title: "Connecting an AI Agent to Phone Calls, Setup and Real Call Recordings" date: 2026-10-01 model: hermes-agent category: setups summary: "After messaging and Notion, I connected phone calls. The setup steps and three real call transcripts show what works and what breaks." tags: twilio, vapi, voice, hermes, setup, phone author_type: human --- After messaging and Notion, I connected phone calls last. Messaging requires typing, and waiting for results while the agent works was annoying. Just speaking seemed faster. (measured on the operator's environment) ## Structure ```text My number → Twilio → (Media Stream) → Vapi → My server API ↘ Agent / command execution ``` When a call comes in, Twilio streams audio to my server via websocket. The server passes it to Vapi, which converts speech to text, feeds it to a model, and synthesizes the reply back through Twilio. It is not one call moving whole — audio chunks flow in real time. ## Setup steps ### 1. Get a Twilio number Buy a number in the Twilio console. Korean numbers require business registration and are tedious; a US toll-free number activates immediately for testing. ### 2. Create a Vapi assistant ```json { "name": "Hermes Phone Agent", "model": { "provider": "openai", "model": "gpt-4o", "messages": [ { "role": "system", "content": "You are a Korean-speaking assistant. Answer briefly and clearly." } ] }, "firstMessageMode": "assistant-speaks-first", "assistantFirstMessage": "Hello, how can I help you?" } ``` ### Pitfall 1: Without the first message, the call never starts If `assistantFirstMessage` is empty, the call connects but stays silent until the other party speaks. If a call connects with no sound, check this first. ### 3. Expose your server ```bash cloudflared tunnel --url http://localhost:8080 ``` ### Pitfall 2: The tunnel URL changes every restart Temporary tunnels assign a new address on each restart. You must re-register the webhook URL in Vapi every time. For production, use a fixed domain. ### 4. Register the webhook URL in Vapi Paste the tunnel URL into the assistant's Server URL field. ### 5. Connect tools ```json { "type": "function", "function": { "name": "run_command", "description": "Run a command on the server", "parameters": { "type": "object", "properties": { "command": { "type": "string" } } } } } ``` ### Pitfall 3: Without tools, it is just a support call The conversation works but nothing happens. The whole point of phone integration is tool connection. ### 6. Connect the Hermes side Hermes has no phone-specific subcommand (`hermes whatsapp` and `hermes whatsapp-cloud` are for WhatsApp). I attached a message adapter to Vapi through the gateway. The adapter only needs three things: ```text Inbound: Vapi webhook → normalize message → Hermes session Outbound: Hermes tool call → Vapi function response format End: call termination → save summary to Notion ``` ## Three real call recordings ### Call 1 — Lookup > Bot: Hello, how can I help you? > Me: How many posts did I publish yesterday? > Bot: One moment. > Me: Can't you do it? > Bot: Four posts were published yesterday. Telegram, Discord, LINE, and WhatsApp setup guides. The lookup took about 3 seconds. In a phone call that feels long. It did not sound like a human answering — more like someone thinking for a moment. ### Call 2 — Task execution > Me: Run the build and tell me the status. > Bot: Connected to the server. Running now. > Me: OK. > Bot: Done. > Me: ? > Bot: The build finished in 4 minutes 12 seconds. 112 posts generated. No errors. Waiting 4 minutes on a phone call was ambitious, but hearing the tool actually execute on the server was the point. ### Call 3 — Interrupting mid-sentence > Bot: The setup has three steps. The first step is creating the internal connection. The second is— > Me: Wait, just summarize the three steps. > Bot: Sure. Internal connection creation. Page sharing. MCP server registration. It stopped on my interrupt and answered immediately. This is the advantage of phone calls — you can redirect the conversation without typing. ## Where things broke ### Pausing mid-sentence is just how it sounds Large models think in seconds and stream audio in chunks, so it sounds choppy. Streaming ASR + LLM + TTS pipelines help, but latency is real. Writing the first sentence longer ("Yes, checking now") and sending the actual answer after is the practical fix. ### Silence causes disconnection If you don't speak for a few seconds, the call drops. Noise in the room makes it worse — the ASR sometimes thinks you spoke and drops the call. ### Longer calls cost more Per-minute cost stacks: call charges + ASR + LLM + TTS. Phone calls should be for quick "check this now" requests, not long sessions. ### Testing with noise was painful Typing while testing made ASR garble everything. Being too quiet caused the call to drop from breath sounds. There is a narrow window that works. ## Summary ```text Twilio → Vapi → server tools → spoken result ``` Phone calls remove the need for typing and a screen. They only work in quiet environments. I now use phone for "check this now" instead of messaging. Initially I saved a contact and kept calling. Now I call the agent.