Adding peer-to-peer communications to an application is relatively straight-forward. Developers can leverage WebRTC APIs or a CPaaS service to quickly add real time voice and video to their web or mobile app. But, what if you want to hold a meeting with more than two people? How can you leverage powerful WebRTC APIs to build a multi party conferencing application?
A REST API is a simple, standardized method of communication between web clients and servers. The main building blocks of the REST API are the request and the response. Learn about the REST API and how to issue requests and receive response data.
With webhooks, your app always knows what happens on the server-side in real time. This makes webhooks ideal for integrating communications apps with events and data from other systems.
Recently, we published a blog post describing why WebSockets are great for real-time services. In this article, we describe the process of establishing, maintaining and closing the WebSockets connection.
Voice Recognition API captures human speech in real-time, transcribes it, and returns it via text. By converting speech to text, you can process live or prerecorded audio, and receive transcriptions and summaries/interpretations with high speed and precision.
Connect any OpenAI-compatible text LLM to your voice AI pipeline in VoxEngine via Chat Completions or Responses API. Set baseUrl to any provider or your own deployment, including EU endpoints for data residency
Voximplant now includes a native Deepgram module that connects any Voximplant call to Deepgram’s Voice Agent API for real-time, speech‑to‑speech conversations. You can stream audio from phone numbers, SIP trunks, WhatsApp, or WebRTC into Deepgram’s unified agent environment—combining STT, LLM reasoning, and TTS—and play responses via Voximplant’s serverless runtime with minimal latency.
A practical guide to running a compliant voice AI stack on Voximplant: EU data residency, compliance documentation, and what to verify at each layer — telephony, LLM, STT, and TTS
Voximplant now includes a native Cartesia module for streaming, low-latency text-to-speech (TTS). You can use a single VoxEngine API to synthesize speech in real time, connect it to any call (PSTN, SIP, WebRTC, WhatsApp) and control playback from a Large Language Model (LLM) or other source, all inside VoxEngine.
Voximplant now supports Inworld's Realtime API, so you can bring Inworld's expressive, conversation-aware agents into real phone calls, SIP, and WhatsApp without custom media infrastructure
Voximplant now includes a native MCP Client for VoxEngine, giving developers direct connectivity to any MCP server and full control over every tool call