Flutter WebRTC means the flutter_webrtc plugin: one Dart API for peer connections, media streams and video renderers on iOS, Android, web and desktop. You capture the microphone, open an RTCPeerConnection, exchange an offer and answer over your own signaling channel, then draw the remote track in an RTCVideoView.
How Flutter WebRTC works
The plugin wraps the native WebRTC library on each platform and exposes the same objects a browser has. The flutter_webrtc package page lists audio, video, data channels and simulcast as supported on Android, iOS, web, macOS, Windows and Linux. What the plugin does not give you is signaling: your app and the other side still need a way to trade session descriptions and ICE candidates.
Most tutorials stop at a call between two phones. You are building something different: a Flutter screen where the user talks and a server-side AI avatar answers with live video and audio. The steps are the same until signaling, where a managed room replaces the server you would otherwise write.
The parts of the call:
Local stream. The microphone, and the camera only if the agent needs to see the user.
Peer connection. Carries encrypted media and picks a network path with ICE.
Signaling. Any channel that delivers the offer, the answer and the candidates.
Renderer. A texture the remote video track is painted on.
If the same avatar also has to run in a browser, the companion article covers adding a realtime AI avatar to a React or Angular app.
Setting up flutter_webrtc on iOS and Android
Add the dependency, declare camera and microphone use on both platforms, and set the Android minimum SDK to 23 if your project is lower. The permission entries and Gradle settings come from the plugin's README at version 1.6.2+hotfix.3.
iOS: Info.plist
Add both usage strings to ios/Runner/Info.plist. iOS terminates an app that opens the microphone or camera without them.
The audio background mode is optional and comes from the LiveKit Flutter README. It keeps a voice call running when the user switches apps. For an audio-only agent you can leave the camera key out.
Android: manifest and Gradle
Put the permissions in android/app/src/main/AndroidManifest.xml:
In the app-level build.gradle, set Java 8 compatibility and the minimum SDK:
For Bluetooth headsets the README adds BLUETOOTH and BLUETOOTH_ADMIN with android:maxSdkVersion="30". Release builds need the Proguard rules the README links.
Getting microphone and camera access
Call navigator.mediaDevices.getUserMedia with a constraints map. The first call shows the system permission prompt, and the returned MediaStream holds the tracks you will send.
Use the audio-only version for an avatar. The avatar's face is driven by the agent's audio, so the user's camera adds a permission prompt and upload bandwidth for nothing unless your agent uses vision. Wrap the call in try and catch: a denied permission throws, and your screen should say so and offer a retry.
Creating the peer connection and renderers
Create the renderer first, then the connection, then attach your tracks and an onTrack handler. An RTCVideoRenderer must be initialized before it receives a stream and disposed when the screen closes.
The first function sends the microphone and asks to receive one video stream, the shape of a call with an avatar. The second releases the microphone, the connection and the renderer. The STUN address is a placeholder: use your own STUN and TURN servers.
Signaling: exchanging the offer and answer with the avatar API
Yes, you need signaling. MDN's signaling and video calling guide states that WebRTC does not specify a transport for it, so you either write a signaling server or use a platform whose SDK contains one.
Hand-built signaling over a WebSocket
If you own both ends, a WebSocket and three message types are enough. This version uses the web_socket_channel package:
The app sends an offer, applies the answer, and trades candidates in both directions. Your server must send the answer before its candidates. The article on WebRTC vs WebSocket for realtime AI explains why the socket carries only these messages and never the media.
What changes with a realtime avatar
The Protoface docs describe no endpoint that accepts an SDP offer and list no Flutter SDK. The avatar reaches your app through a LiveKit room. Your backend creates a session, the avatar joins the room as a participant, and your Flutter app joins the same room with the LiveKit Flutter SDK, which runs the offer, answer and ICE exchange for you.

The app never calls the avatar service. It asks your backend for a room token, joins the LiveKit room, and plays the tracks the avatar publishes there.
The app asks your backend to start a conversation. No API key ships in the app.
The backend mints a LiveKit access token for the user and returns it with the room URL. LiveKit's access tokens reference describes the token as a JWT signed with your API secret that encodes the participant's identity, the room name and permissions.
Your agent starts the avatar. With the LiveKit Agents plugin that is
protoface.AvatarSessionandavatar.start(session, room=ctx.room). Without the plugin, the backend callsPOST /v1/sessions.The app connects to the room. The SDK negotiates with the LiveKit server over its own secure WebSocket.
The avatar joins as a participant and publishes audio and video. With the plugin its identity is
protoface-avatar-agent.The app receives a subscribed-track event and renders the video.
The REST request for step 3 builds the session around a transport object:
You mint worker_token yourself, so Protoface never sees your LiveKit secret. The default audio_source is data_stream, which the LiveKit Agents plugin uses. A custom agent that publishes its voice as an ordinary audio track sets it to track. The call returns at once with status: "queued" and a sess_... ID. Poll GET /v1/sessions/{id}: the status moves through starting to running, and first_frame_at is set when the first frame reaches the room. A 503 with code at_capacity carries a Retry-After header. The raw WebSocket transport in the schema is reserved and not available.
The Protoface Realtime integrations page lists the other agent stacks, and the protoface-quickstart repository indexes a starter project for each. None is a Flutter project: the Flutter side is plain LiveKit.
Keep both secrets on the server. The Protoface API key and the LiveKit API secret stay in your backend or agent process. The app receives only a short-lived room token.
Playing the avatar's video and audio stream
With raw flutter_webrtc, put RTCVideoView(remoteRenderer) in your widget tree and call setState after assigning srcObject. With LiveKit, pass the subscribed track to VideoTrackRenderer. Remote audio plays on its own in both cases.
Speaker or earpiece
A call-mode audio session can play through the earpiece, which sounds like silence when the phone is held at arm's length. With raw flutter_webrtc, call Helper.setSpeakerphoneOn(true) once the connection is up. LiveKit takes over the audio session: its audio session guide says speaker output is preferred by default and a wired or Bluetooth headset still wins. To change the route, call AudioManager.instance.setSpeakerOutputPreferred. Do not mix the two APIs, because LiveKit disables the plugin's own audio management when it loads.
Complete Flutter WebRTC example
This main.dart joins a LiveKit room, publishes the microphone and shows the remote video track the avatar publishes. It needs livekit_client in pubspec.yaml and the platform setup from earlier. Version 2.13.0 requires Flutter 3.38 and Dart 3.10 or later. The livekit_client package pins flutter_webrtc as a dependency, so the permissions are the same.
Start the agent from the Protoface realtime quickstart, mint a token for the same room, then run:
In production, replace the two constants with a request to your backend. Add a catch on connect, a RoomDisconnectedEvent handler and a TrackUnsubscribedEvent handler that clears avatar. For full projects, the plugin's maintainers publish the flutter-webrtc-demo repository for peer-to-peer calls, and the LiveKit Flutter SDK ships a conferencing app in its example folder.
flutter_webrtc vs LiveKit and other SDKs
Use raw flutter_webrtc when you own both ends of a one-to-one link and want full control. Use a managed SDK when a server-side agent or avatar has to join the call, because the room, signaling and reconnection are already built.
Question | flutter_webrtc | LiveKit Flutter SDK | WebSocket only |
|---|---|---|---|
Signaling | You build it | Built in | Not needed |
Media handling | Native WebRTC | Native WebRTC, via flutter_webrtc | You write capture, playback and buffering |
Servers you run | Signaling, STUN, TURN | A token endpoint | A socket server |
Avatar or agent in the call | Only if it speaks your signaling | Joins the room as a participant | No video path |
Best for | Custom peer-to-peer links | Voice agents and avatars | Text, events, transcripts |
A WebSocket is the right tool for chat messages, transcripts and control events, and the wrong one for a user's microphone on a mobile network. The comparison of LiveKit vs WebSocket for realtime voice and video apps goes through that trade in detail.
Native Android in Kotlin
The architecture is identical without Flutter. Your backend mints the room token, the agent starts the avatar, and the app joins the room with the LiveKit Android SDK, which renders remote video in a SurfaceViewRenderer or TextureViewRenderer. Hold the room in a ViewModel, not the Activity, so a rotation does not drop the call. Request RECORD_AUDIO before you connect.
Session lifecycle on mobile
Debounce the start button. Two taps should not create two sessions. Avatar sessions count against your plan's concurrency cap.
End what you start. Disconnect when the screen closes. A Protoface session ends on
POST /v1/sessions/{id}/end, at its duration cap, or afteridle_timeout_secondswithout audio from the agent. The default is 30 seconds.Decide what backgrounding means. Keep the call with the audio background mode, or leave the room and rejoin on resume.
Keep app state out of the call. Lesson progress, itinerary or cart data belongs in your backend, so a reconnect does not lose it.
Offer a fallback. If the microphone is denied or the room fails, show text chat.
Common Flutter WebRTC errors and fixes
Most failures come from four places: a renderer that was never initialized, the audio route, a missing permission entry, or a network that blocks UDP.
Symptom | Likely cause | Fix |
|---|---|---|
Black video | Renderer not initialized, or no rebuild after | Await |
No audio on iOS | Sound is routed to the earpiece | Switch the route to the speaker and test on a real device |
iOS app closes on start of call | Usage string missing from | Add the microphone and camera keys |
| Permission missing from the manifest, or denied by the user | Add the permission, or send the user to app settings |
Camera fails in the iOS simulator | The simulator has no camera | Test audio-only there, video on a device |
ICE fails or never connects | No TURN relay on a network that blocks UDP | Add a TURN server that listens on TLS port 443 |
Release build crashes on Android | WebRTC classes stripped by code shrinking | Add the plugin's Proguard rules |
No avatar participant appears | The agent is in a different room | Confirm the agent and the app use the same LiveKit room |
Avatar joins but stays silent | The agent is not producing speech | Check the agent's speech output first |
Measure connection time on your own devices
Log a timestamp when the user taps start, when connect returns and when the first video track is subscribed. Compare the last one with first_frame_at on the Protoface session to see how much of the wait is the avatar starting and how much is the network. Repeat on Wi-Fi and on mobile data.
Common questions
Does flutter_webrtc work on web and desktop as well as iOS and Android?
Yes. The flutter_webrtc package page lists audio, video and data channels as supported on Android, iOS, web, macOS, Windows and Linux. Speaker and earpiece switching only applies on phones.
Do I need a signaling server for Flutter WebRTC?
Yes, something has to carry the offer, the answer and the ICE candidates between the two sides. You can write one over a WebSocket, or use a platform SDK such as LiveKit that includes signaling and leaves you with a token endpoint.
Should I use flutter_webrtc directly or the LiveKit Flutter SDK?
Use flutter_webrtc directly for a one-to-one link where you control both ends. Use the LiveKit Flutter SDK when a server-side voice agent or avatar joins the call, because it is built on flutter_webrtc and adds rooms, signaling and reconnection.
When should a Flutter app use WebSocket instead of WebRTC?
Use a WebSocket for text, transcripts, events and other data that must arrive complete and in order. Use WebRTC for the user's microphone and for live video, where late packets should be skipped and echo cancellation matters.
Is WebRTC the right choice for a Flutter voice or video app?
Yes for live two-way audio and video. It gives you echo cancellation, a jitter buffer and congestion control on mobile networks. For one-way playback of recorded media, ordinary HTTP streaming is simpler.
Where can I find a Flutter WebRTC example on GitHub?
The plugin's maintainers publish the flutter-webrtc-demo repository for peer-to-peer calls. For a room with a server-side agent, the LiveKit Flutter SDK repository includes a conferencing app in its example folder.
Put a face on the agent your Flutter app already calls
Protoface Realtime joins your LiveKit room as a participant and turns your agent's audio into live avatar video. Your Flutter app renders it like any other track.





