Vapi
Developer platform for building production voice AI agents on the phone and web. Abstracts orchestration of STT, LLM, TTS, telephony, and tool calling into a single API and SDK.
About Vapi
Vapi is a developer platform for building voice AI agents. Engineers wire up speech-to-text, an LLM, a voice, telephony, and tool calls through one API and SDKs in TypeScript or Python, then ship agents that book appointments, qualify leads, run reception, or handle outbound calls. Vapi handles the orchestration, latency, interruption handling, and tool routing that make voice agents feel real, and integrates with Twilio, Vonage, and SIP for actual phone numbers. Vapi is sold to product and engineering teams that are building voice into their own product, not to non-technical buyers who want a finished agent. The pricing reflects that: $0.05 per minute platform fee plus separate metered costs for STT, LLM, TTS, and telephony, with real production deployments landing at $0.23 to $0.33 per minute.
Platform details
- Pricing: Pay-as-you-go: $0.05 per minute platform fee, plus separate metered costs for STT, LLM, TTS, and telephony. True total is typically $0.23 to $0.33 per minute in production.
- Founded: 2023, San Francisco
- Funding: $70M+ total. $20M Series A in Dec 2024 led by Bessemer Venture Partners, then $50M Series B led by Peak XV Partners at roughly $500M valuation. Backers include Y Combinator, AI Grant, Abstract Ventures, Saga Ventures, Michael Ovitz.
- Last reviewed: 2026-05-31
Official website
What does Vapi actually do?
Vapi is a developer platform for building production voice AI agents. It abstracts away the hard real-time engineering (speech-to-text, LLM, text-to-speech, telephony, interruption handling, tool calling) into a single API and SDKs in TypeScript and Python. You wire your providers (Deepgram, OpenAI, ElevenLabs, Twilio, and so on), define your prompts and tools, and Vapi handles the orchestration, latency, and barge-in that make voice agents feel like a real conversation. Phone numbers come via Twilio, Vonage, or SIP. Vapi handles the bridge to its WebSocket endpoint and runs the agent there. The platform is a real engineering win: building this from scratch is a year of work, and Vapi compresses it to a weekend for a competent engineer. Customers tend to be product teams shipping a voice feature inside their own product (outbound sales agents, AI receptionists, scheduling agents, support callbacks). The TechCrunch coverage on Amazon Ring choosing Vapi over 40 rivals is the right anchor for the kind of buyer this is for. Sistava is not infrastructure. We sell a workforce, not a voice pipeline. The honest framing is that Vapi is the picks-and-shovels for voice AI engineers, and Sistava is the workforce for founders. Both can exist in the same stack: a Sistava AI Employee could trigger a Vapi call. They are not substitutes.
How much does Vapi cost?
Vapi's headline rate is $0.05 per minute, which is the platform orchestration fee only. On top of that you pay your own STT provider (around $0.01 per minute), the LLM (anywhere from $0.02 to $0.20 per minute depending on the model), TTS (about $0.04 per minute on ElevenLabs standard), and telephony (Twilio at around $0.013 per leg per minute). The honest production number is $0.23 to $0.33 per minute for a typical deployment. A leaner stack (GPT-4o + Deepgram Nova-2 + ElevenLabs standard) can run $0.08 to $0.15 per minute plus telephony. Most teams discover they are managing contracts with four to six providers, which is its own ongoing operational cost. Sistava prices flat monthly plans on our published page. You do not pay per minute, per call, per token, or per provider. That model is wrong if you are building voice into your own product (you want the metered control Vapi gives you), and right if you are a founder hiring a workforce. Different unit of consumption. Different buyer. We are not pretending the comparison is apples to apples.
When does Vapi beat Sistava?
Vapi wins any time your product needs voice as a core feature and you have engineers to build it. If you are shipping an AI receptionist inside your SaaS, a phone qualification agent inside your sales platform, or a voice support callback inside your app, Vapi is the right primitive. The latency tuning, interruption handling, BYO model stack, and telephony bridges are the right surface area. It also wins when you want fine-grained control over the model stack. You can swap STT providers, change voices, route LLM calls to whichever model is cheapest or smartest for that turn, and observe everything in their dashboard. That control is exactly what an engineering team building a voice product needs. Sistava cannot match that. We are not a developer voice platform and have no plans to become one. If voice is your wedge or your product is voice, Vapi is the honest recommendation. The clean test: are you building voice infrastructure, or are you hiring a workforce that happens to use voice as one of many channels? Building infrastructure goes Vapi. Hiring a workforce goes Sistava.
Where does Vapi fall short for a non-technical buyer?
Vapi is unusable without engineers. The product is a SDK and an API. There is no kanban board, no role hire, no plan-and-supervise loop. You are writing prompts, defining tools, wiring providers, debugging latency, and shipping code. That is fine for the audience Vapi is built for and wrong for the audience Sistava is built for. The cost model is also a mismatch for non-technical buyers. Per-minute pricing across four to six providers is hard to forecast, hard to budget, and hard to explain to a finance team. A flat published plan is the answer for a founder who needs to know what next month costs. And Vapi is single-purpose: voice agents. If you also need a marketer writing content, a salesperson running outbound, or a support agent on email, Vapi is one of three or four tools you would have to stack and integrate yourself. Sistava bundles roles. Marketing, sales, support, ops, all on one workforce, one plan, one board, one bill. The trade is that we do not match Vapi's voice depth. Honest swap for the buyer who wants a workforce instead of infrastructure.
How does Vapi handle the developer stack?
Vapi's developer surface is mature. TypeScript and Python SDKs, a hosted dashboard for testing, tool calling for any HTTP endpoint, real-time WebSocket streaming, and integrations with Twilio, Vonage, and SIP for telephony. The bring-your-own-model stack means you choose Deepgram or AssemblyAI for STT, OpenAI or Anthropic or others for the LLM, ElevenLabs or Rime for TTS. That is power and complexity in equal parts. The wins are real (sub-second latency, natural turn-taking, tool calling that completes real workflows). The costs are real too: production deployments require contracts with four to six different providers, debugging is multi-vendor, and every model swap is a config decision someone has to own. Sistava abstracts all of that away. We pick the best frontier model per task, manage the providers, and present you with a hire-the-role experience. You give up the control. You gain the speed. If your engineers want to own the stack, use Vapi. If you want to own the outcome and skip the stack, use Sistava. Two valid stances, two products, do not confuse them.
Comparison
- Compare Vapi with Sistava — See the two platforms side by side when you are ready to evaluate them.