What direct speech-to-speech means

Real-time voice translation begins with streaming audio. The application reads small microphone frames continuously instead of waiting for a complete recording. A compatible provider accepts that audio stream and returns translated speech as audio. The application then plays or routes those returned frames while the conversation continues.

Polyglot-Audio deliberately requires direct streaming audio input and translated audio output. It does not silently substitute a speech recognition, text translation, and speech synthesis chain. That distinction matters because each additional stage can add buffering, change failure behavior, and expose intermediate transcripts.

Direct route

Microphone → streaming translation provider → translated audio → virtual microphone or headphones

“Direct” does not mean local or offline. A cloud provider may still process the stream, require an account, charge for usage, and apply its own retention terms. Check the provider documentation before sending sensitive conversations.

Why two-way translation needs two routes

A natural conversation has two independent directions. The outgoing or TX route translates what you say and exposes it as an input to a calling, conferencing, streaming, or voice-chat application. The incoming or RX route captures the other participant’s application audio, translates it, and plays the result through your headphones.

TX: your translated voice goes out

The TX route reads your physical microphone. Translated output is written to a virtual microphone that the target voice application can select. The remote participant hears that translated signal.

RX: their translated voice comes back

The voice application sends the remote participant’s audio to a separate virtual route. Polyglot-Audio reads that route, translates it in the opposite direction, and plays the result through a physical output such as headphones.

TX and RX therefore need separate provider sessions, stream state, device selection, controls, buffers, and failure handling. Treating them as one pipeline makes it harder to stop one direction, diagnose a device problem, or recover from one failed stream without interrupting the other.

What affects translation latency

End-to-end latency is the combined delay of audio capture, local buffering, network travel, provider processing, returned audio buffering, and playback. No honest application can promise one fixed latency across every computer, network, language pair, and provider.

  • Frame and buffer size: large buffers resist dropouts but take longer to fill.
  • Network quality: distance, congestion, jitter, and packet loss affect a cloud session.
  • Provider behavior: models differ in how much context they wait for before producing speech.
  • Audio format conversion: incompatible sample rates or channel layouts may require extra work.
  • Conversation style: clear turns and short pauses are easier than overlapping speakers.

Practical expectation: optimize for conversational continuity, not an artificial zero-latency claim. Headphones prevent translated output from feeding back into the microphone.

Why virtual audio devices are needed

Most voice applications accept a microphone and produce speaker output, but they do not offer a translation interface. Virtual audio devices bridge that boundary. They appear as selectable inputs or outputs and let translated audio move between independent applications.

On macOS, a setup can use separate BlackHole routes. Windows can use two distinct VB-CABLE route families. Linux users can create separate PipeWire or PulseAudio-compatible virtual routes. Polyglot-Audio does not bundle or install these third-party drivers. Keeping installation separate makes device ownership and system changes explicit.

Route names differ by platform, but the rule stays the same: never reuse one virtual route for both TX and RX. Separate routes reduce feedback and make the direction of every signal auditable.

Privacy, accuracy, and honest limits

Polyglot-Audio keeps credentials in the operating system keychain and excludes raw audio, transcripts, authorization headers, device names, and device identifiers from its own diagnostics. The open-source engine can be inspected, and a local mock provider can test routing without a network connection.

Those application boundaries do not replace provider due diligence. Translation accuracy depends on language support, accents, noise, terminology, and model availability. Cloud processing depends on provider authorization, billing, regional availability, terms, and retention policy. Do not use automated translation as the sole authority for medical, legal, emergency, or safety-critical decisions.

A reliable setup checklist

  1. Install two separate virtual audio routes suitable for your operating system.
  2. Select the physical microphone and TX virtual output in Polyglot-Audio.
  3. Select that TX virtual output as the microphone in the voice application.
  4. Send the voice application’s speaker output to the separate RX virtual route.
  5. Select the RX virtual route as input and physical headphones as output in Polyglot-Audio.
  6. Choose supported source and target languages for each direction.
  7. Test one direction at a time before enabling full duplex.
  8. Confirm provider pricing, privacy, and retention terms before a real conversation.

The project is early-stage open-source software. Use the official installation tutorials and obtain builds only from the canonical GitHub Releases page.