WebRTC vs WebSocket: Which to Use for Realtime AI

You are picking a transport for a voice agent or avatar. Decide by what travels on each hop, then wire the two together.

Michael Trehan

Founder, Protoface

Published

July 7, 2026

Updated

October 2, 2026

Cover comparing WebRTC over UDP for live audio and video with WebSocket over TCP for ordered data, and using both together
On this page

WebRTC vs WebSocket comes down to what you are sending. Use WebRTC for live audio and video between a user's device and your agent: it runs over UDP and handles packet loss, jitter and echo for you. Use WebSocket for ordered data over TCP: signaling, transcripts, events and server-to-server audio.

WebRTC vs WebSocket: the short answer

WebRTC is a media stack. It carries encrypted audio and video over UDP, drops what arrives too late and keeps playing. WebSocket is a message pipe. It carries ordered text or binary frames over one TCP connection and never drops anything. A voice agent or avatar usually needs both: WebRTC on the hop that touches the user, WebSocket everywhere else.

What you are building

Use

Why

Voice agent in a browser or mobile app

WebRTC

Microphone and speaker on a network you do not control

Talking avatar with lip sync

WebRTC

Audio and video share one clock, so the mouth stays on the words

Audio between your server and an STT, TTS or speech model API

WebSocket

Data-center links lose few packets and need no NAT traversal

Phone audio handed to you by a telephony provider

WebSocket

The provider sets the format, and the link is server to server

Transcripts, tool calls, UI state, session events

WebSocket

Every message must arrive, in order

A production browser app with a live agent

Both

WebSocket sets up and controls the call, WebRTC carries it

The other peer in a realtime AI call is a server, not a second person, so peer-to-peer WebRTC is rarely what you run. The browser connects to a media server, and your agent connects to the same server as another participant. The article on WebRTC SFU, P2P and MCU architectures covers that choice.

How WebRTC and WebSocket differ

They sit at different layers. WebSocket is a transport for bytes. WebRTC is a transport plus codecs, timing, encryption and network adaptation for media.

Property

WebRTC

WebSocket

Transport

RTP over UDP by default, TCP or TLS relay as fallback

One TCP connection, usually inside TLS

Connection setup

Offer, answer and ICE candidates exchanged over a channel you provide

One HTTP Upgrade request

Media handling

Codecs, jitter buffer, audio and video sync, bitrate adaptation built in

None. You send bytes and build the rest

Lost packets

Concealed, corrected or skipped so playback continues

Retransmitted. Everything behind them waits

Encryption

Always on

On when you use wss://

NAT and firewalls

Needs ICE, STUN and sometimes TURN

Works wherever HTTPS works

Browser support

All current major browsers

All current major browsers

Transport. RFC 6455 defines WebSocket as an independent TCP-based protocol whose handshake is an HTTP Upgrade request, on port 80 or on port 443 under TLS. WebRTC media is RTP, normally sent over UDP.

Connection setup. A WebSocket opens with one request. A WebRTC connection needs both sides to trade a session description and network candidates before any media flows, and the standard leaves that exchange to you.

Media handling. RFC 7874 requires every WebRTC endpoint to implement the Opus and G.711 audio codecs, so two endpoints always share a codec. A WebSocket has no idea its payload is audio.

Reliability. TCP guarantees order and delivery. RFC 8834 gives WebRTC the opposite toolkit: RTCP feedback, optional retransmission, forward error correction and mandatory congestion control, all tuned to keep media on time instead of complete.

Encryption. RFC 8834 also forbids unencrypted RTP: endpoints must use the secure profile for every packet. A WebSocket is encrypted only if you serve it over TLS.

Latency: what each transport adds

On a clean link the two are close. The gap opens when packets are lost: TCP holds back everything behind the missing segment until it is resent, and WebRTC plays on. Wi-Fi and mobile networks lose packets routinely, which is why the user-facing hop decides the choice.

Head-of-line blocking on a WebSocket

Voice is sent in small frames, typically 20 milliseconds each. When one TCP segment is lost, the receiver's operating system keeps every later frame in its buffer until the retransmission lands. Your code sees silence, then a burst. A deeper playback buffer hides the burst, but every frame of every call then pays that delay.

A WebSocket also has no way to drop stale audio. If the sender produces faster than the network delivers, frames queue. Watch bufferedAmount on the sending socket; a number that keeps growing is latency accumulating.

Loss handling in WebRTC

WebRTC treats a late packet as a lost packet. The jitter buffer holds just enough audio to absorb variation in arrival time and resizes itself as the network changes. A missing audio frame is concealed by the decoder, and a missing video frame can be requested again or skipped. Congestion control lowers the bitrate before queues build.

One caveat: when a network blocks UDP and the call falls back to a TURN relay over TCP or TLS, WebRTC inherits TCP's blocking on that leg.

Is WebRTC faster than WebSocket for audio?

Not on a good network. Over a short, low-loss path, such as two services in one cloud region, raw audio on a WebSocket arrives about as fast as RTP would. WebRTC wins on the paths you do not control, because its delay stays bounded when packets drop.

Measure it on your own call

No vendor's number will match your network, model and region, so read the figures from a live session. The browser reports them through getStats(). MDN documents the inbound RTP statistics, including jitter buffer delay and packets lost.

async function sample(pc) {
  const report = await pc.getStats();
  report.forEach((s) => {
    if (s.type === "inbound-rtp" && s.kind === "audio") {
      if (!s.jitterBufferEmittedCount) return;
      const bufferMs = (s.jitterBufferDelay / s.jitterBufferEmittedCount) * 1000;
      console.log("jitter buffer ms", bufferMs, "packets lost", s.packetsLost);
    }
    if (s.type === "candidate-pair" && s.nominated && s.state === "succeeded") {
      console.log("round trip ms", s.currentRoundTripTime * 1000);
    }
  });
}
setInterval(() => sample(pc), 2000);
async function sample(pc) {
  const report = await pc.getStats();
  report.forEach((s) => {
    if (s.type === "inbound-rtp" && s.kind === "audio") {
      if (!s.jitterBufferEmittedCount) return;
      const bufferMs = (s.jitterBufferDelay / s.jitterBufferEmittedCount) * 1000;
      console.log("jitter buffer ms", bufferMs, "packets lost", s.packetsLost);
    }
    if (s.type === "candidate-pair" && s.nominated && s.state === "succeeded") {
      console.log("round trip ms", s.currentRoundTripTime * 1000);
    }
  });
}
setInterval(() => sample(pc), 2000);
async function sample(pc) {
  const report = await pc.getStats();
  report.forEach((s) => {
    if (s.type === "inbound-rtp" && s.kind === "audio") {
      if (!s.jitterBufferEmittedCount) return;
      const bufferMs = (s.jitterBufferDelay / s.jitterBufferEmittedCount) * 1000;
      console.log("jitter buffer ms", bufferMs, "packets lost", s.packetsLost);
    }
    if (s.type === "candidate-pair" && s.nominated && s.state === "succeeded") {
      console.log("round trip ms", s.currentRoundTripTime * 1000);
    }
  });
}
setInterval(() => sample(pc), 2000);

The function prints the average time audio waits in the jitter buffer and the network round trip on the active candidate pair. Both jitter buffer fields are running totals, so subtract the previous sample for a per-interval value. Run it on office Wi-Fi, then on a throttled mobile profile. The article on measuring and reducing WebRTC latency walks through the full budget from microphone to first avatar frame.

Audio and video handling

WebRTC gives you a working call the moment tracks are attached. Over a WebSocket you receive bytes and write the audio engine yourself.

What WebRTC does for you

  • Codecs. Opus for audio and a negotiated video codec, with encoder settings that follow the available bitrate.

  • Jitter buffer and sync. RTP timestamps let the receiver play audio at a steady pace and line video up with it. For an avatar, that shared timeline is what keeps the lips on the words.

  • Echo cancellation. RFC 7874 says a WebRTC endpoint should include echo control, and browsers ship echo cancellation, noise suppression and gain control for microphone tracks. Leave them on, and the agent does not hear its own voice from the user's speakers and interrupt itself. See how WebRTC echo cancellation works and where it fails.

  • Bandwidth adaptation. The sender lowers bitrate or resolution when the path narrows. With a media server you can add simulcast for per-viewer bandwidth adaptation.

What you build by hand over a WebSocket

Sending raw PCM chunks over a socket means writing the playback scheduler yourself. A minimal browser version queues each chunk right after the previous one:

const RATE = 24000; // must match the rate your server sends
const ctx = new AudioContext({ sampleRate: RATE });
let playhead = 0;
ws.binaryType = "arraybuffer";
ws.onmessage = ({ data }) => {
  const pcm = new Int16Array(data);
  const buffer = ctx.createBuffer(1, pcm.length, RATE);
  const channel = buffer.getChannelData(0);
  for (let i = 0; i < pcm.length; i++) channel[i] = pcm[i] / 32768;
  const source = ctx.createBufferSource();
  source.buffer = buffer;
  source.connect(ctx.destination);
  playhead = Math.max(playhead, ctx.currentTime + 0.05);
  source.start(playhead);
  playhead += buffer.duration;
};
const RATE = 24000; // must match the rate your server sends
const ctx = new AudioContext({ sampleRate: RATE });
let playhead = 0;
ws.binaryType = "arraybuffer";
ws.onmessage = ({ data }) => {
  const pcm = new Int16Array(data);
  const buffer = ctx.createBuffer(1, pcm.length, RATE);
  const channel = buffer.getChannelData(0);
  for (let i = 0; i < pcm.length; i++) channel[i] = pcm[i] / 32768;
  const source = ctx.createBufferSource();
  source.buffer = buffer;
  source.connect(ctx.destination);
  playhead = Math.max(playhead, ctx.currentTime + 0.05);
  source.start(playhead);
  playhead += buffer.duration;
};
const RATE = 24000; // must match the rate your server sends
const ctx = new AudioContext({ sampleRate: RATE });
let playhead = 0;
ws.binaryType = "arraybuffer";
ws.onmessage = ({ data }) => {
  const pcm = new Int16Array(data);
  const buffer = ctx.createBuffer(1, pcm.length, RATE);
  const channel = buffer.getChannelData(0);
  for (let i = 0; i < pcm.length; i++) channel[i] = pcm[i] / 32768;
  const source = ctx.createBufferSource();
  source.buffer = buffer;
  source.connect(ctx.destination);
  playhead = Math.max(playhead, ctx.currentTime + 0.05);
  source.start(playhead);
  playhead += buffer.duration;
};

The code converts 16-bit mono PCM to floats and schedules it with a fixed 50 millisecond margin. That margin is a jitter buffer with one setting. It does not adapt, conceal gaps or resync after a stall. You still have to cancel queued audio the instant the user interrupts and capture microphone audio upstream. Uncompressed PCM also costs far more bandwidth than Opus, and video on the same socket competes with the audio for one ordered stream.

Voice activity detection behaves differently on each path. A WebRTC track arrives as a steady, echo-canceled stream, so turn detection sees clean timing. Socket audio arrives in whatever rhythm TCP allows, so a stall can look like the end of a sentence.

NAT traversal, STUN and TURN

WebRTC has to find a UDP path between two endpoints that may both sit behind NATs, and that takes ICE plus servers you run or rent. A WebSocket is one outbound TLS connection on port 443, which passes almost any firewall or proxy.

MDN's WebRTC protocols reference describes the parts. ICE is the framework that tests candidate paths. A STUN server tells a client its public address. A TURN server relays all media when no direct path works, and MDN notes that this comes with overhead, so it is used only when there are no alternatives.

  • Setup time. Candidate gathering and connectivity checks run before the first packet of media.

  • Relay bandwidth. Every relayed call sends its full audio and video through your TURN servers, and you pay for that traffic.

  • Operations. TURN needs credentials, regional placement and monitoring. Media servers hold state per call, so scaling them takes more than a round-robin balancer: see load balancing WebRTC sessions across servers.

A media server with a public address makes this easier, because only the user's side is behind a NAT. You still need TURN over TLS on port 443 for corporate networks that block UDP. If your users sit on locked-down networks and you cannot run relays, use a managed WebRTC service. When calls fail to connect, start with the guide to troubleshooting ICE, autoplay and audio failures in WebRTC.

WebSockets have their own network trap: idle proxies close quiet connections. Send ping frames from the server, since browser code cannot send them, or add an application heartbeat. Reconnect with backoff.

When to use WebSocket for realtime AI

Use a WebSocket when both ends are servers, or when the payload is data that must arrive complete.

  • Server-to-server pipelines. Streaming audio from your agent process to a speech-to-text, text-to-speech or speech-to-speech API. OpenAI's Realtime API guide draws the same line: the session connects over WebRTC in the browser or WebSocket on the server.

  • Telephony bridges. The caller's audio reaches the provider over the phone network. The provider forwards it to your server over a socket in a fixed telephony format.

  • Text and control. Partial transcripts, tool calls, "user interrupted" events, avatar state and session lifecycle.

When to use WebRTC for realtime AI

Use WebRTC whenever a person's microphone, speaker or screen is on one end and the exchange has to feel like a conversation.

  • Browser and mobile voice agents. The user's network is the weakest link, and echo cancellation is required as soon as they are not wearing headphones.

  • Avatars with lip sync. Video multiplies the bitrate and makes any audio and video drift visible.

  • Interruptions. Barge-in only feels natural when the user's speech reaches the agent with low, stable delay.

  • Kiosks and long sessions. A screen in a lobby runs for hours on venue Wi-Fi. Congestion control and ICE restarts keep a call alive through conditions that would stall a socket.

Raw WebRTC or a platform such as LiveKit?

LiveKit, Daily and similar platforms are WebRTC, with the signaling, media servers, TURN and client SDKs already built. Choose raw RTCPeerConnection when you need full control of a one-to-one link and can run the infrastructure. Choose a platform when your agent has to join the call as a participant, which is the normal shape for voice AI. The comparison of LiveKit vs WebSocket for voice and video apps covers when a managed room beats a raw socket.

Avatars follow that shape. Protoface Realtime adds a face to a voice agent, driven by the audio the agent already produces. With the LiveKit plugin, the avatar joins your room as a participant and publishes audio and video, so the browser receives ordinary WebRTC tracks:

from livekit.plugins import protoface

avatar = protoface.AvatarSession(avatar_id="av_stock_001")
await avatar.start(session, room=ctx.room)

await session.start(agent=agent, room=ctx.room)
from livekit.plugins import protoface

avatar = protoface.AvatarSession(avatar_id="av_stock_001")
await avatar.start(session, room=ctx.room)

await session.start(agent=agent, room=ctx.room)
from livekit.plugins import protoface

avatar = protoface.AvatarSession(avatar_id="av_stock_001")
await avatar.start(session, room=ctx.room)

await session.start(agent=agent, room=ctx.room)

The avatar session starts before the agent session, and the plugin then routes the agent's audio to the avatar participant. The Pipecat integration shows the split from the other side. POST /v1/pipecat/sessions returns short-lived WebSocket media credentials for the pipecat-protoface package, which runs in your Pipecat worker. The service then emits synchronized audio and video frames to your pipeline's output transport.

When you do not want to own the transport at all

If the avatar is a contained feature on a website, an embed removes the decision. Protoface hosts the conversation, and the page carries only a public embed ID, no API key. You can use a share link, the <protoface-avatar> element, or your own interface built on protoface-client.

Using both: WebSocket for signaling, WebRTC for media

The standard design uses a WebSocket to set up and steer the call and a peer connection to carry it. MDN's signaling and video calling guide states that WebRTC does not specify a transport for signaling. A WebSocket fits because either side can send at any time, and the guide's own example uses one.

WebSocket carries signaling and events, WebRTC carries the audio and video

The WebSocket sets up and controls the call. WebRTC carries the media through a media server, where the agent and the avatar join as participants.

  1. The browser opens a WebSocket to your signaling server and authenticates.

  2. The browser creates an offer describing its microphone track and the video it wants to receive, and sends it over the socket.

  3. The server, or the media server behind it, replies with an answer.

  4. Both sides trade ICE candidates over the socket as they are discovered.

  5. ICE picks a working path. Encrypted audio and video now flow over UDP, outside the socket.

  6. The socket stays open for transcripts, state and control. A WebRTC data channel can take over that role if you prefer one connection.

const ws = new WebSocket("wss://example.com/signal");
const pc = new RTCPeerConnection({
  iceServers: [{ urls: "stun:stun.example.com:3478" }],
});
const send = (msg) => ws.send(JSON.stringify(msg));

pc.onicecandidate = ({ candidate }) => {
  if (candidate) send({ type: "candidate", candidate });
};
pc.ontrack = ({ streams }) => {
  document.querySelector("video").srcObject = streams[0];
};

ws.onmessage = async ({ data }) => {
  const msg = JSON.parse(data);
  if (msg.type === "answer") await pc.setRemoteDescription(msg.answer);
  if (msg.type === "candidate") await pc.addIceCandidate(msg.candidate);
};

ws.onopen = async () => {
  const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
  mic.getTracks().forEach((track) => pc.addTrack(track, mic));
  pc.addTransceiver("video", { direction: "recvonly" });
  await pc.setLocalDescription(await pc.createOffer());
  send({ type: "offer", offer: pc.localDescription });
};
const ws = new WebSocket("wss://example.com/signal");
const pc = new RTCPeerConnection({
  iceServers: [{ urls: "stun:stun.example.com:3478" }],
});
const send = (msg) => ws.send(JSON.stringify(msg));

pc.onicecandidate = ({ candidate }) => {
  if (candidate) send({ type: "candidate", candidate });
};
pc.ontrack = ({ streams }) => {
  document.querySelector("video").srcObject = streams[0];
};

ws.onmessage = async ({ data }) => {
  const msg = JSON.parse(data);
  if (msg.type === "answer") await pc.setRemoteDescription(msg.answer);
  if (msg.type === "candidate") await pc.addIceCandidate(msg.candidate);
};

ws.onopen = async () => {
  const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
  mic.getTracks().forEach((track) => pc.addTrack(track, mic));
  pc.addTransceiver("video", { direction: "recvonly" });
  await pc.setLocalDescription(await pc.createOffer());
  send({ type: "offer", offer: pc.localDescription });
};
const ws = new WebSocket("wss://example.com/signal");
const pc = new RTCPeerConnection({
  iceServers: [{ urls: "stun:stun.example.com:3478" }],
});
const send = (msg) => ws.send(JSON.stringify(msg));

pc.onicecandidate = ({ candidate }) => {
  if (candidate) send({ type: "candidate", candidate });
};
pc.ontrack = ({ streams }) => {
  document.querySelector("video").srcObject = streams[0];
};

ws.onmessage = async ({ data }) => {
  const msg = JSON.parse(data);
  if (msg.type === "answer") await pc.setRemoteDescription(msg.answer);
  if (msg.type === "candidate") await pc.addIceCandidate(msg.candidate);
};

ws.onopen = async () => {
  const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
  mic.getTracks().forEach((track) => pc.addTrack(track, mic));
  pc.addTransceiver("video", { direction: "recvonly" });
  await pc.setLocalDescription(await pc.createOffer());
  send({ type: "offer", offer: pc.localDescription });
};

The browser publishes the microphone, asks to receive one video stream, and attaches the remote stream to a <video> element when it arrives. The message names are yours, and the two URLs are placeholders for your own signaling and STUN or TURN servers. The server must answer the offer, then send its candidates in the same format: a candidate added before the answer is set is rejected.

Keep keys off the socket. Authenticate the signaling connection with a short-lived token minted by your backend. Provider API keys and media server secrets stay on the server.

The same choice on your framework or platform

The transport decision does not change with the framework. What changes is where the WebRTC implementation comes from and where you clean it up.

Platform

WebRTC

WebSocket

React, Angular, Vue, SvelteKit, plain HTML

Built into the browser

Built into the browser

iOS and Android native

A WebRTC library or a platform SDK you bundle

System or common networking libraries

Unity and game engines

An engine WebRTC package or a platform SDK

A socket client for game state and events

Python, Node.js, Rust back ends

An agent framework or a server-side WebRTC library

Native or one small dependency

In a browser framework, keep the connection in a service or store, not in a component. Create it on a user gesture so audio is allowed to play. Close the peer connection and stop the microphone tracks in the teardown hook: ngOnDestroy in Angular, onUnmounted in Vue, onDestroy in Svelte, an effect cleanup in React.

On native mobile, the WebRTC library adds app size and audio session handling that a socket does not. For a server written in Rust, see streaming a realtime avatar with Rust WebRTC.

gRPC streaming and multiplayer games

gRPC streams run on HTTP/2 over TCP, so they share WebSocket's blocking behavior, and browsers cannot call them without a proxy. Use gRPC between services. For games, send fast-changing state over a WebRTC data channel configured as unordered and unreliable, and keep lobby, chat and turn-based messages on a WebSocket.

One rule to keep. If a human hears it or sees it live, send it over WebRTC. If a program reads it, or both ends are servers, a WebSocket is enough.

Common questions

What are the downsides of using WebRTC?

You have to build or buy signaling, run STUN and TURN servers, and debug failures that only appear on other people's networks. Server-side WebRTC is also heavier than a socket, and calls with more than a few parties need a media server.

Is WebRTC still used?

Yes. It ships in every major browser and is the standard way to carry live audio and video on the web. Voice AI has added to its use: OpenAI's Realtime API connects over WebRTC in the browser.

Does ChatGPT use SSE or WebSocket?

OpenAI's API streams text responses over HTTP as server-sent events. Its Realtime API for speech connects over WebRTC in the browser or WebSocket on the server. OpenAI does not document the ChatGPT app's own transport in those guides, so check your browser's network panel if you need to know.

What is replacing WebSockets?

Nothing has replaced them. The closest candidate is WebTransport, which runs over HTTP/3 and offers multiple streams plus unreliable datagrams. MDN lists it as newly available across browsers since March 2026, so keep a WebSocket fallback.

Is WebRTC built on top of WebSockets?

No. WebRTC media travels as encrypted RTP, normally over UDP. A WebSocket is often used beside it to exchange the offer, answer and ICE candidates, but the standard does not require one.

Can I use WebRTC with a Python backend?

Yes. Agent frameworks such as LiveKit Agents and Pipecat are Python and handle the media transport for you, and Python WebRTC libraries exist if you want to terminate a peer connection yourself.

Add a face to the agent you already run

Keep your transport. Protoface Realtime joins your LiveKit room or Pipecat pipeline and turns the agent's audio into live avatar video.

Start free or see Protoface Realtime.

Michael Trehan

Founder, Protoface

Michael is the founder of Protoface. He was previously a software engineer at Radiant Nuclear and worked in investment banking at JP Morgan.

Keep reading