Agentic VoiceDSP is a VoxEngine module that handles turn-taking in cascade voice AI pipelines. It detects when a caller starts speaking, finishes, interrupts the agent or goes silent, and turns each moment into an event your agent or call flow can act on. Voice agents on your own STT, LLM and TTS get natural conversation flow in the same runtime that runs the call. 

Turn-taking is where voice agents often break in production. An agent can run on the best models available and still jump in while the caller pauses to think, or cut off mid-sentence because someone said "uh-huh." Callers notice this long before they notice anything about the model. 

How does Agentic VoiceDSP work?

Agentic VoiceDSP runs inside VoxEngine, Voximplant's serverless runtime for call logic, on the same media server that processes the call. There is no separate framework to deploy and no extra service between the caller and your agent.

The module decides when a caller's turn starts and ends. Your scenario decides what the agent does with that, for example, stopping playback when the caller interrupts or following up when the caller goes quiet.

Why do cascade voice pipelines need turn-taking control?

Turn-taking is the logic that decides who speaks when in a conversation. Speech-to-speech models handle it natively. A cascade pipeline connects separate speech-to-text, LLM and text-to-speech components chosen by the developer, and has no built-in turn-taking, so the developer has to build it. 

The cascade approach gives you full control over the stack: a speech recognition engine optimized for a specific language, an LLM fine-tuned on your data, a TTS voice that matches your brand. That flexibility used to cost you conversation quality.

In March 2026, Voximplant added voice activity detection and end-of-turn detection to VoxEngine as separate modules. They provided the signals, but turning those signals into conversation logic stayed with the developer. Agentic VoiceDSP brings both modules and noise suppression under one interface, with turn-taking logic built in.

Turn-taking options for voice agents on Voximplant

  • Speech-to-speech connector: one provider model handles speech and turn-taking inside the model.
  • Cascade pipeline with Agentic VoiceDSP: any combination of STT, LLM and TTS, with turn-taking built into the module and configured in your scenario.

What conversation problems does Agentic VoiceDSP solve?

  1. What happens when a caller pauses mid-sentence? The caller says "I need to change my..." and stops to think. When Agentic VoiceDSP detects the pause, an AI turn detection model checks whether the phrase sounds finished. If it doesn't, the agent waits, by default up to 3 seconds of silence. If the caller starts talking again before the model responds, the verdict is discarded and the turn continues.
  2. How do you keep the agent talking when a caller says "uh-huh"? Set a minimum word count for the start of a caller's turn. Agentic VoiceDSP then starts the turn only after the caller says that many words, so short acknowledgements like "uh-huh" don't interrupt the agent and it finishes its sentence. The word count comes from transcribed speech, so this setting needs speech recognition connected. 
  3. What happens when a caller goes silent? Agentic VoiceDSP has an idle timer that starts when the agent finishes speaking. The timer is off by default, and you set its duration. If the caller doesn't start speaking before it expires, your scenario receives an event, and the agent repeats the question or checks whether the caller is still on the line. If the caller still doesn't respond, set up a polite goodbye. 
  4. Does turn detection work on noisy phone lines? Yes, Agentic VoiceDSP includes noise suppression, a new addition to VoxEngine. Audio is cleaned on the server before it reaches voice activity and turn detection, so both work on cleaner input. Noise suppression is in beta and stays off unless you configure it.

Which turn-taking settings can you configure? 

Agentic VoiceDSP lets you choose how each caller's turn starts and ends, so the same module fits a support line and an outbound survey.

Start strategies

  • Voice activity: the turn starts as soon as the caller's voice is detected.
  • First transcribed words: the turn starts when speech recognition returns text.
  • Minimum word count: the turn starts after the caller says a set number of words.
    Stop strategies

Stop strategies

  • Turn detection model: an AI model decides whether the caller has finished. If the phrase doesn't sound finished, the agent waits through a short silence of the length you set.
  • Silence timeout: use this instead of the model when you want the turn to end after a fixed silence. It doesn’t override the model while the model is on.

Additional settings

  • Interruptions: allow or block the caller interrupting the agent, and switch this at any point during the call.
  • Idle follow-up: how long the agent waits before checking in with a silent caller.
  • Fallback timeout: if nothing else closes a turn, Agentic VoiceDSP ends it after 5 seconds by default, so the conversation never stalls.

An agent reading a legal disclaimer or payment terms can block interruptions for that part and allow them again once it's done.

When speech recognition is connected, Agentic VoiceDSP signals the moment it has enough information to start generating a response, before the final transcript arrives. Your LLM starts working without waiting for speech recognition to finish.

What does a basic Agentic VoiceDSP setup include?

  1. Create an Agentic VoiceDSP instance with your start and stop strategies.
  2. Connect speech recognition if you use transcription-based start strategies.
  3. Route the call audio to Agentic VoiceDSP.
  4. Signal when the agent starts and stops speaking, so the module can detect interruptions and silence.
  5. Handle turn events in your scenario logic.

Voice activity detection and the turn detection model work on audio alone.

Does Voice AI turn detection work in every language? 

Agentic VoiceDSP has no language restrictions. Turn detection relies partly on intonation, and in some languages a pause and the end of a phrase sound alike. For those cases, adjust how long the agent waits after a pause and how confident the model needs to be before it ends the turn.

Who should use Agentic VoiceDSP?

Agentic VoiceDSP is built for developers running cascade pipelines on Voximplant. Use it when your case requires a specific speech recognition engine, your own model or a particular voice, and the call still needs to sound like a natural conversation.

How do I get started with Agentic VoiceDSP?

Agentic VoiceDSP is available now in VoxEngine. The API reference covers every parameter and event.

Voice activity and turn detection are billed as one connection. Noise suppression is in beta and free during that period. 

Resources

Agentic VoiceDSP Documentation 
Agentic VoiceDSP API reference
VAD and turn detection
Sign up for Voximplant