> ## Documentation Index
> Fetch the complete documentation index at: https://docs.unpod.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# How a call flows

> What runs between the caller audio and your agent text - whether users dial a number or connect from a browser or app.

## What runs the call

Unpod gives your agent voice. It handles everything between the user's audio and your agent's text logic. The Speech Stack is the developer surface of the [Calling Engine](/core-engine/calling-engine) - one channel in the Transport layer of the [Core Engine](/core-engine/overview).

Your agent receives plain text. It returns plain text. The Speech Stack handles everything else: transcription, synthesis, VAD, barge-in detection, endpointing, and transport.

***

## Two Ways Users Connect

<CardGroup cols={2}>
  <Card title="Phone / Telephony" icon="phone">
    Users dial a phone number. Unpod handles SIP, PSTN, number provisioning, and routing. No carrier account or SIP trunk needed.
  </Card>

  <Card title="Browser / App" icon="monitor-smartphone">
    Users connect from a web app or mobile app - same speech pipeline, no phone number required. Today browser sessions run via the hosted Playground; the embeddable `@unpod/web-sdk` is not yet published on npm.
  </Card>
</CardGroup>

Both paths run through the same managed speech pipeline (STT, TTS, VAD, barge-in). Your agent code is identical regardless of how the user connects.

***

## How Audio Flows

<Frame>
  <img src="https://mintcdn.com/unpodai/9OLw2S-v9psMSqik/images/diagrams/unpod-voice-stack.svg?fit=max&auto=format&n=9OLw2S-v9psMSqik&q=85&s=68d4f631702c1116700d8fcf23480f1f" alt="Unpod voice stack diagram showing phone, browser, and mobile entrypoints flowing through the managed speech layer into your AgentRunner and dialog machine." width="1672" height="941" data-path="images/diagrams/unpod-voice-stack.svg" />
</Frame>

<Steps>
  <Step title="User connects">
    Via a phone number (PSTN/SIP), or from a browser/app (via the hosted Playground today; Web SDK not yet published).
  </Step>

  <Step title="Speech pipeline runs">
    Unpod transcribes audio (STT), detects turn end (VAD + endpointing), and handles barge-in interruptions. Your agent gets clean text.
  </Step>

  <Step title="Your agent responds">
    The Unpod orchestrator dispatches the session to your `AgentRunner`. Your entrypoint runs, your dialog machine produces a text reply.
  </Step>

  <Step title="Reply synthesised">
    Unpod converts the reply to speech (TTS) and streams it back to the user.
  </Step>

  <Step title="Session ends">
    Transcript, metrics, and recording are stored and queryable via the management API.
  </Step>
</Steps>

***

## Core Building Blocks

<CardGroup cols={2}>
  <Card title="Voice Profiles" icon="audio-lines" href="/speech-stack/voice-profiles">
    Choose STT and TTS providers. Pre-built profiles or custom combinations with automatic failover.
  </Card>

  <Card title="Agents" icon="robot" href="/speech-stack/agents">
    One `agent_id`, one brain, N voices. Say the brain out loud - a playbook, a prompt, your own runner, or your own endpoint.
  </Card>

  <Card title="Phone Numbers" icon="phone" href="/speech-stack/numbers">
    Provision numbers directly or bring your own. Attach to an agent for inbound calls.
  </Card>

  <Card title="AgentRunner & Sessions" icon="code" href="/speech-stack/agent-runner">
    Install `unpod`, run your `AgentRunner`, and accept sessions from any source.
  </Card>
</CardGroup>

***

## Quickstart Paths

| I want to...              | Start here                                                                                                       |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| Place your first call     | [Quickstart](/get-started/quickstart) - six SDK calls, a Prompt brain, nothing to deploy                         |
| Run my own brain on calls | [Run Your Own Brain](/get-started/first-phone-call) - an AgentRunner serving inbound and outbound                |
| Add voice to a web app    | [Realtime](/get-started/realtime) - browser sessions via the hosted Playground today (Web SDK not yet published) |

***

## Next Steps

<CardGroup cols={2}>
  <Card title="Quickstart" icon="bolt" href="/get-started/quickstart">
    A named agent, a number, a live call - six SDK calls, one file.
  </Card>

  <Card title="Run a SuperDialog agent" icon="layers" href="/speech-stack/level-up-superdialog">
    Drive conversations with structured flows and tools.
  </Card>
</CardGroup>
