OpenTelemetry for GenAI Voice and Avatar Apps

Your agent feels slow and you cannot say which stage is to blame. Trace each turn with standard GenAI spans and add the avatar spans the conventions leave out.

Michael Trehan

Founder, Protoface

Published

July 7, 2026

Updated

October 2, 2026

Round analog gauges on an aircraft cockpit instrument panel
On this page

OpenTelemetry GenAI is the set of semantic conventions and instrumentation libraries that give generative AI calls standard span names, attributes, metrics and events. In a voice or avatar app, a library traces the model call, and you add your own spans for speech to text, text to speech and the avatar session.

What OpenTelemetry GenAI is

OpenTelemetry GenAI is a naming standard plus the libraries that apply it. The conventions define the span, attributes, metrics and events for a model call. Instrumentation libraries patch a model client so those spans appear without changes to your call sites.

  • The conventions have their own repository. They moved out of the main semantic conventions into open-telemetry/semantic-conventions-genai.

  • Everything in it is marked Development. The spans, metrics and events documents all carry that status, so names can still change. Pin your library versions.

The conventions cover one model call. A voice agent with an avatar is a chain of speech to text, model, text to speech and video, so the other stages are yours to name. If you are still choosing how many stages to run, start with the comparison of the OpenAI Realtime API vs STT + LLM + TTS pipelines, because that choice sets how many spans a turn has.

What are the OpenTelemetry conventions for GenAI telemetry?

They define a client span per model call, a small set of gen_ai.* attributes on it, client metrics for duration and tokens, and an opt-in event that carries the full request and response. The span is named {gen_ai.operation.name} {gen_ai.request.model}, for example chat gpt-4o-mini, and its kind is CLIENT.

Name

Signal

Level

What it holds

gen_ai.operation.name

Span attribute

Required

chat, embeddings, execute_tool, invoke_agent and others

gen_ai.provider.name

Span attribute

Required

openai, anthropic, gcp.gemini and others

gen_ai.request.model

Span attribute

Required if available

The model you asked for

gen_ai.conversation.id

Span attribute

Required if available

Your real conversation or thread ID

gen_ai.usage.input_tokens, gen_ai.usage.output_tokens

Span attribute

Recommended

Token counts for the call

gen_ai.response.time_to_first_chunk

Span attribute

Recommended when streaming

Seconds from request to first chunk

error.type

Span attribute

Required on error

Provider error code or exception name

gen_ai.input.messages, gen_ai.output.messages

Span attribute or event

Opt-in

Prompt and response content

gen_ai.client.operation.duration

Histogram, seconds

Recommended

Duration of each call

gen_ai.client.operation.time_to_first_chunk

Histogram, seconds

Recommended when streaming

Wait for the first chunk

gen_ai.client.inference.usage.input_tokens, .output_tokens

Counters

Recommended

Tokens, split by gen_ai.token.modality

gen_ai.client.inference.operation.details

Event

Opt-in

Full request parameters and chat history

The span attributes come from the GenAI spans document, the histograms from the GenAI metrics document, and the counters from the inference token metrics document. error.type is the one stable attribute in the table, because it belongs to the general conventions.

  • Modality. gen_ai.token.modality accepts text, image, audio and unknown, and gen_ai.output.type accepts speech. The spans document also lists gen_ai.usage.audio.input_tokens and gen_ai.usage.audio.output_tokens, so a speech-to-speech model can report audio tokens apart from text.

  • Renamed attributes. gen_ai.system became gen_ai.provider.name, and gen_ai.usage.prompt_tokens and gen_ai.usage.completion_tokens became the input and output token attributes. Libraries lag the documents: release 1.2b0 of the OpenAI instrumentation still records tokens in a gen_ai.client.token.usage histogram split by gen_ai.token.type, which the current token document does not list. Print the names your version emits before you build a dashboard.

  • Conversation ID. The spans document says not to fill gen_ai.conversation.id with a new UUID or a trace ID. Set it only when you hold a real identifier.

Set up OpenTelemetry GenAI instrumentation

Install the SDK, an OTLP exporter and one GenAI instrumentation package, then register the providers before you create the model client. The Python packages live in open-telemetry/opentelemetry-python-genai, which covers OpenAI, Anthropic, Bedrock, Google GenAI, LangChain, LlamaIndex and several agent frameworks. They are beta releases.




from opentelemetry import metrics, trace
from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.genai.openai import OpenAIInstrumentor
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor

resource = Resource.create({"service.name": "voice-agent"})

tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(tracer_provider)

reader = PeriodicExportingMetricReader(OTLPMetricExporter())
metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader]))

OpenAIInstrumentor().instrument()
from opentelemetry import metrics, trace
from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.genai.openai import OpenAIInstrumentor
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor

resource = Resource.create({"service.name": "voice-agent"})

tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(tracer_provider)

reader = PeriodicExportingMetricReader(OTLPMetricExporter())
metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader]))

OpenAIInstrumentor().instrument()
from opentelemetry import metrics, trace
from opentelemetry.exporter.otlp.proto.http.metric_exporter import OTLPMetricExporter
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.genai.openai import OpenAIInstrumentor
from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.export import PeriodicExportingMetricReader
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor

resource = Resource.create({"service.name": "voice-agent"})

tracer_provider = TracerProvider(resource=resource)
tracer_provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(tracer_provider)

reader = PeriodicExportingMetricReader(OTLPMetricExporter())
metrics.set_meter_provider(MeterProvider(resource=resource, metric_readers=[reader]))

OpenAIInstrumentor().instrument()

The Python code sends traces and metrics over OTLP and patches the OpenAI client. Each chat completion, Responses call and embedding produces a span named by the conventions, plus duration and token metrics. For streaming chat completions the OpenAI instrumentation README also lists time to first chunk and time between chunks.

Prompts stay out of telemetry by default. OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT defaults to no_content. Set it to span_only, event_only or span_and_event only where your backend is allowed to store what users say.

That README does not list audio transcription, speech synthesis or realtime sessions, so those stages need spans you write.

Add GenAI attributes to your own spans

When a client has no instrumentation, open a CLIENT span yourself, name it by the convention and set the required attributes before the call. Add token counts after it returns.

from opentelemetry import trace
from opentelemetry.trace import SpanKind, Status, StatusCode

tracer = trace.get_tracer("voice.agent")

def traced_chat(call_model, provider, model, conversation_id):
    with tracer.start_as_current_span(f"chat {model}", kind=SpanKind.CLIENT) as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.provider.name", provider)
        span.set_attribute("gen_ai.request.model", model)
        span.set_attribute("gen_ai.conversation.id", conversation_id)
        try:
            text, input_tokens, output_tokens = call_model()
        except Exception as exc:
            span.set_attribute("error.type", type(exc).__name__)
            span.set_status(Status(StatusCode.ERROR))
            raise
        span.set_attribute("gen_ai.usage.input_tokens", input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", output_tokens)
        return text
from opentelemetry import trace
from opentelemetry.trace import SpanKind, Status, StatusCode

tracer = trace.get_tracer("voice.agent")

def traced_chat(call_model, provider, model, conversation_id):
    with tracer.start_as_current_span(f"chat {model}", kind=SpanKind.CLIENT) as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.provider.name", provider)
        span.set_attribute("gen_ai.request.model", model)
        span.set_attribute("gen_ai.conversation.id", conversation_id)
        try:
            text, input_tokens, output_tokens = call_model()
        except Exception as exc:
            span.set_attribute("error.type", type(exc).__name__)
            span.set_status(Status(StatusCode.ERROR))
            raise
        span.set_attribute("gen_ai.usage.input_tokens", input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", output_tokens)
        return text
from opentelemetry import trace
from opentelemetry.trace import SpanKind, Status, StatusCode

tracer = trace.get_tracer("voice.agent")

def traced_chat(call_model, provider, model, conversation_id):
    with tracer.start_as_current_span(f"chat {model}", kind=SpanKind.CLIENT) as span:
        span.set_attribute("gen_ai.operation.name", "chat")
        span.set_attribute("gen_ai.provider.name", provider)
        span.set_attribute("gen_ai.request.model", model)
        span.set_attribute("gen_ai.conversation.id", conversation_id)
        try:
            text, input_tokens, output_tokens = call_model()
        except Exception as exc:
            span.set_attribute("error.type", type(exc).__name__)
            span.set_status(Status(StatusCode.ERROR))
            raise
        span.set_attribute("gen_ai.usage.input_tokens", input_tokens)
        span.set_attribute("gen_ai.usage.output_tokens", output_tokens)
        return text

call_model is your own function. It calls the provider and returns the reply text with the two token counts from the provider's usage field. Keep error.type to an error code or exception class, never the message.

Trace a voice pipeline: speech to text, model and text to speech

Make each conversational turn one trace, with a root span for the turn and one child per stage.

The span tree of one turn: a voice.turn root with three child spans, and a span link to the avatar session span in its own trace

One turn is one trace. The three stages are children of voice.turn, and the avatar session is reached through a span link because it lives in its own trace.

The span tree for one turn of a three-stage pipeline with an avatar:

  1. voice.turn, the root. Starts when the user stops speaking and ends when the reply audio has been sent. Custom.

  2. voice.stt, child of the turn. Covers transcription. Custom.

  3. chat {model}, child of the turn. Created by the instrumentation library. Standard.

  4. voice.tts, child of the turn. Covers synthesis, with a voice.first_audio event on the first chunk. Custom.

  5. A link from voice.turn to the avatar session span, which lives in its own trace. Custom.

With a speech-to-speech model, items 2 to 4 collapse into one span around the realtime response. Write that one by hand as a CLIENT span with the same GenAI attributes that traced_chat sets.

import time

def run_turn(audio, conversation_id, avatar_session_id, session_link):
    with tracer.start_as_current_span("voice.turn", links=[session_link]) as turn:
        turn.set_attribute("gen_ai.conversation.id", conversation_id)
        turn.set_attribute("avatar.session.id", avatar_session_id)
        started = time.monotonic()

        with tracer.start_as_current_span("voice.stt"):
            text = transcribe(audio)

        reply = generate(text)  # the instrumented client adds the chat span

        with tracer.start_as_current_span("voice.tts") as tts:
            first = True
            for chunk in synthesize(reply):
                if first:
                    tts.add_event("voice.first_audio")
                    waited = time.monotonic() - started
                    turn.set_attribute("voice.time_to_first_audio", waited)
                    first = False
                send_audio(chunk)
import time

def run_turn(audio, conversation_id, avatar_session_id, session_link):
    with tracer.start_as_current_span("voice.turn", links=[session_link]) as turn:
        turn.set_attribute("gen_ai.conversation.id", conversation_id)
        turn.set_attribute("avatar.session.id", avatar_session_id)
        started = time.monotonic()

        with tracer.start_as_current_span("voice.stt"):
            text = transcribe(audio)

        reply = generate(text)  # the instrumented client adds the chat span

        with tracer.start_as_current_span("voice.tts") as tts:
            first = True
            for chunk in synthesize(reply):
                if first:
                    tts.add_event("voice.first_audio")
                    waited = time.monotonic() - started
                    turn.set_attribute("voice.time_to_first_audio", waited)
                    first = False
                send_audio(chunk)
import time

def run_turn(audio, conversation_id, avatar_session_id, session_link):
    with tracer.start_as_current_span("voice.turn", links=[session_link]) as turn:
        turn.set_attribute("gen_ai.conversation.id", conversation_id)
        turn.set_attribute("avatar.session.id", avatar_session_id)
        started = time.monotonic()

        with tracer.start_as_current_span("voice.stt"):
            text = transcribe(audio)

        reply = generate(text)  # the instrumented client adds the chat span

        with tracer.start_as_current_span("voice.tts") as tts:
            first = True
            for chunk in synthesize(reply):
                if first:
                    tts.add_event("voice.first_audio")
                    waited = time.monotonic() - started
                    turn.set_attribute("voice.time_to_first_audio", waited)
                    first = False
                send_audio(chunk)

transcribe, generate, synthesize and send_audio stand for your own stage functions. Because the turn span is current when generate runs, the library's chat span nests under it with no extra code. The link uses the span links API from the OpenTelemetry Python instrumentation guide.

The conventions list no well-known operation name for transcription or synthesis. Leave gen_ai.operation.name off those two spans until the conventions add one. You can still set gen_ai.provider.name and gen_ai.request.model on them so a backend groups all three stages by vendor and model. If a framework such as LiveKit Agents or Pipecat runs the stages for you, check whether it already emits spans for them before you wrap anything by hand.

Add avatar spans for a realtime session

Give the avatar session its own short trace, and connect every turn to it with a span link and a shared session ID attribute. A session can run for many minutes. A parent span held open that long is lost if the process dies before it ends.

Session start

With the Protoface LiveKit plugin, avatar.start(...) creates the session and puts its sess_... ID on avatar.session_id. Wrap that call:

from livekit.plugins import protoface

avatar = protoface.AvatarSession(avatar_id="av_stock_001")
with tracer.start_as_current_span("avatar.session.start") as span:
    span.set_attribute("avatar.id", "av_stock_001")
    await avatar.start(session, room=ctx.room)
    span.set_attribute("avatar.session.id", avatar.session_id)
    session_link = trace.Link(span.get_span_context())
from livekit.plugins import protoface

avatar = protoface.AvatarSession(avatar_id="av_stock_001")
with tracer.start_as_current_span("avatar.session.start") as span:
    span.set_attribute("avatar.id", "av_stock_001")
    await avatar.start(session, room=ctx.room)
    span.set_attribute("avatar.session.id", avatar.session_id)
    session_link = trace.Link(span.get_span_context())
from livekit.plugins import protoface

avatar = protoface.AvatarSession(avatar_id="av_stock_001")
with tracer.start_as_current_span("avatar.session.start") as span:
    span.set_attribute("avatar.id", "av_stock_001")
    await avatar.start(session, room=ctx.room)
    span.set_attribute("avatar.session.id", avatar.session_id)
    session_link = trace.Link(span.get_span_context())

The span times session creation from your agent's side, and each voice.turn links to session_link. If you call POST /v1/sessions yourself, set error.type to the error.code in a failed response, for example concurrent_sessions or session_start_rate_limited. The Protoface docs keep error.code stable and say the message may change. A rising count of concurrency errors is a capacity signal: see load balancing WebRTC sessions across servers.

First video frame

A Protoface session moves through queued, starting and running, and first_frame_at is set when the first frame reaches the room. Read both timestamps from GET /v1/sessions/{id} and write a span with explicit start and end times:

import asyncio, os
from datetime import datetime
import httpx

def to_ns(stamp):
    moment = datetime.fromisoformat(stamp.replace("Z", "+00:00"))
    return int(moment.timestamp() * 1_000_000_000)

async def record_first_frame(session_id):
    url = f"https://api.protoface.com/v1/sessions/{session_id}"
    headers = {"Authorization": f"Bearer {os.environ['PROTOFACE_API_KEY']}"}
    async with httpx.AsyncClient() as client:
        while True:
            response = await client.get(url, headers=headers)
            response.raise_for_status()
            data = response.json()
            done = data["status"] in ("ended", "failed", "canceled")
            if data.get("first_frame_at") or done:
                break
            await asyncio.sleep(1)
    span = tracer.start_span("avatar.first_frame", start_time=to_ns(data["created_at"]))
    span.set_attribute("avatar.session.id", session_id)
    span.set_attribute("avatar.session.status", data["status"])
    if data.get("failure"):
        span.set_attribute("error.type", data["failure"]["code"])
        span.set_status(Status(StatusCode.ERROR))
    frame_at = data.get("first_frame_at")
    span.end(end_time=to_ns(frame_at) if frame_at else None)
import asyncio, os
from datetime import datetime
import httpx

def to_ns(stamp):
    moment = datetime.fromisoformat(stamp.replace("Z", "+00:00"))
    return int(moment.timestamp() * 1_000_000_000)

async def record_first_frame(session_id):
    url = f"https://api.protoface.com/v1/sessions/{session_id}"
    headers = {"Authorization": f"Bearer {os.environ['PROTOFACE_API_KEY']}"}
    async with httpx.AsyncClient() as client:
        while True:
            response = await client.get(url, headers=headers)
            response.raise_for_status()
            data = response.json()
            done = data["status"] in ("ended", "failed", "canceled")
            if data.get("first_frame_at") or done:
                break
            await asyncio.sleep(1)
    span = tracer.start_span("avatar.first_frame", start_time=to_ns(data["created_at"]))
    span.set_attribute("avatar.session.id", session_id)
    span.set_attribute("avatar.session.status", data["status"])
    if data.get("failure"):
        span.set_attribute("error.type", data["failure"]["code"])
        span.set_status(Status(StatusCode.ERROR))
    frame_at = data.get("first_frame_at")
    span.end(end_time=to_ns(frame_at) if frame_at else None)
import asyncio, os
from datetime import datetime
import httpx

def to_ns(stamp):
    moment = datetime.fromisoformat(stamp.replace("Z", "+00:00"))
    return int(moment.timestamp() * 1_000_000_000)

async def record_first_frame(session_id):
    url = f"https://api.protoface.com/v1/sessions/{session_id}"
    headers = {"Authorization": f"Bearer {os.environ['PROTOFACE_API_KEY']}"}
    async with httpx.AsyncClient() as client:
        while True:
            response = await client.get(url, headers=headers)
            response.raise_for_status()
            data = response.json()
            done = data["status"] in ("ended", "failed", "canceled")
            if data.get("first_frame_at") or done:
                break
            await asyncio.sleep(1)
    span = tracer.start_span("avatar.first_frame", start_time=to_ns(data["created_at"]))
    span.set_attribute("avatar.session.id", session_id)
    span.set_attribute("avatar.session.status", data["status"])
    if data.get("failure"):
        span.set_attribute("error.type", data["failure"]["code"])
        span.set_status(Status(StatusCode.ERROR))
    frame_at = data.get("first_frame_at")
    span.end(end_time=to_ns(frame_at) if frame_at else None)

The function polls once a second until the first frame or a terminal status, then records a span from created_at to first_frame_at. Start it with asyncio.create_task(record_first_frame(avatar.session_id)) inside the avatar.session.start block so it joins that trace, and keep a reference to the task so it is not garbage collected. The API reference also names a session.first_frame webhook, which replaces the polling. Both timestamps come from Protoface's clock, so the duration is exact, but the span can sit slightly off from your local spans.

Stream health

The server and the viewer see different things, so record both. On the server, the session object has usage.frames. The API reference calls those counters eventually consistent, so a flat frame count is a hint, not an alarm. Record the final status when the session ends: by default a session closes after 30 seconds without inbound audio, and the Protoface docs say that timer follows the audio your agent sends, not when the visitor stops speaking. Unlogged, that looks like a crash to the user.

In the browser, playback problems such as blocked autoplay or a failed ICE connection never reach your server traces. Send the session ID to the page, log client events against it, and use the guide to troubleshooting ICE, autoplay and audio failures in WebRTC to decide which events to capture. For the avatar setup itself outside Python, see the Protoface Realtime integration for Node.js.

Attribute

Standard or custom

Set on

gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model

Standard, Development

Model spans

gen_ai.conversation.id

Standard, Development

Model and turn spans

error.type

Standard, Stable

Any failed span

voice.time_to_first_audio

Custom

voice.turn

avatar.session.id, avatar.id

Custom

Turn and avatar spans

avatar.session.status

Custom

avatar.first_frame, session end

Keep custom names out of the gen_ai. namespace. A future release of the conventions could define the same name with a different meaning.

Metrics to track for GenAI voice and avatar apps

Take model duration, time to first chunk and token counts from the standard metrics, and add three histograms of your own: turn duration, time to first audio and time to first frame.

Metric

Type

Source

gen_ai.client.operation.duration

Histogram, seconds

Standard

gen_ai.client.operation.time_to_first_chunk

Histogram, seconds

Standard, streaming only

gen_ai.client.inference.usage.input_tokens, .output_tokens

Counters

Standard

voice.turn.duration

Histogram, seconds

Custom

voice.time_to_first_audio

Histogram, seconds

Custom

avatar.time_to_first_frame

Histogram, seconds

Custom

avatar.session.failures

Counter

Custom, by error.type

from opentelemetry import metrics

meter = metrics.get_meter("voice.agent")
first_audio = meter.create_histogram("voice.time_to_first_audio", unit="s")
first_frame = meter.create_histogram("avatar.time_to_first_frame", unit="s")

first_audio.record(waited, {"gen_ai.request.model": model})
first_frame.record(frame_seconds, {"avatar.id": "av_stock_001"})
from opentelemetry import metrics

meter = metrics.get_meter("voice.agent")
first_audio = meter.create_histogram("voice.time_to_first_audio", unit="s")
first_frame = meter.create_histogram("avatar.time_to_first_frame", unit="s")

first_audio.record(waited, {"gen_ai.request.model": model})
first_frame.record(frame_seconds, {"avatar.id": "av_stock_001"})
from opentelemetry import metrics

meter = metrics.get_meter("voice.agent")
first_audio = meter.create_histogram("voice.time_to_first_audio", unit="s")
first_frame = meter.create_histogram("avatar.time_to_first_frame", unit="s")

first_audio.record(waited, {"gen_ai.request.model": model})
first_frame.record(frame_seconds, {"avatar.id": "av_stock_001"})

waited is the value from the turn code, model is the model name you requested, and frame_seconds is first_frame_at minus created_at. Never put a session ID, conversation ID or transcript on a metric: every new value creates a new time series. Those belong on spans.

Your target for first-audio time depends on your models, region and network. Record a week of your own traffic, read the 50th and 95th percentiles, and alert on a change from that baseline.

Export and read the telemetry

Point the OTLP exporter at a Collector or a backend, then debug one slow turn at a time by reading its span tree. The exporters in the setup code read OTEL_EXPORTER_OTLP_ENDPOINT, which the OTLP exporter configuration reference defaults to http://localhost:4318 for HTTP. Use OTEL_EXPORTER_OTLP_HEADERS for a backend's auth header.

export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp.example.com"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer YOUR_TOKEN"
export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp.example.com"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer YOUR_TOKEN"
export OTEL_EXPORTER_OTLP_ENDPOINT="https://otlp.example.com"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer YOUR_TOKEN"

Both values are placeholders. To find the slow stage in a turn:

  1. Query for voice.turn spans and sort by voice.time_to_first_audio, highest first.

  2. Open one trace. The three children sit on one timeline, so the widest bar is the stage to fix.

  3. On the chat span, compare gen_ai.response.time_to_first_chunk with the span's full duration. A long wait for the first chunk points at the model or the prompt size. A short wait and a long tail points at output length.

  4. Check the gaps between children. Time that belongs to no span is your own code: buffering, turn detection or a queue.

  5. If audio was on time but the user saw nothing, follow the link to the session trace and read avatar.first_frame and avatar.session.status.

  6. Filter every span by avatar.session.id to rebuild the whole call in order.

Common questions

Which OpenTelemetry GenAI instrumentation libraries exist for Python?

The opentelemetry-python-genai repository publishes packages for OpenAI, the OpenAI Agents SDK, Anthropic, Amazon Bedrock, Google GenAI, LangChain, LlamaIndex, DSPy and several other frameworks. They are beta releases, so pin the version you deploy.

Are the GenAI semantic conventions stable?

No. The spans, metrics, token metrics and events documents in the GenAI conventions repository are all marked Development, so attribute and metric names can still change between releases.

Do the GenAI conventions capture prompts and responses, and how do I turn that off?

Not by default. Message content is opt-in in the conventions, and the Python OpenAI instrumentation defaults OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT to no_content. Leave it unset, or set it back to no_content, to keep prompts and replies out of your telemetry.

Which observability backends read OpenTelemetry GenAI data?

Any backend that accepts OTLP can store the spans and metrics, because they are ordinary OpenTelemetry signals with agreed names. What differs is whether a backend builds model, token and conversation views on top of the gen_ai.* attributes, so check its documentation for the convention version it reads.

Put a face on the agent you are tracing

Protoface Realtime turns your agent's audio into live avatar video and reports when the first frame lands, so the avatar fits the same trace as the rest of the turn.

Start free or see the Node.js integration.

Michael Trehan

Founder, Protoface

Michael is the founder of Protoface. He was previously a software engineer at Radiant Nuclear and worked in investment banking at JP Morgan.

Keep reading