Team Chat voice rooms

Persistent voice rooms with shared chat and optional self-hosted audio.

Create a channel in Team Chat and choose Voice room. On iOS and macOS, choose New voice room in the room menu. Public rooms are discoverable across the workspace; private rooms retain the existing member controls. The main agent is initially disabled in a new voice room; room admins can invite it into the text chat through the existing settings. Room history stays in Team Chat and remains available when nobody is talking.

Select Join voice to listen. Everyone joins muted. Enable microphone requests microphone permission and lets other participants hear you. The participant strip shows muted microphones and highlights speakers. Leave voice stops microphone capture and audio playback; text chat stays available. Switching rooms or leaving Team Chat also ends the voice connection in V1. A second device using the same account replaces the first connection.

Voice rooms start with human audio only. Room managers can use + Add agent to add the live voice agent, the text agent, or both. The voice settings icon opens a popover with toggles for live transcription and translation (German, English or Italian). The info icons explain joining and AI listening without taking space from chat. Agents and translation require transcription; disabling transcription also removes both agents and translation. All participants see when AI listening is enabled, including before joining. Camera/video support remains reserved for V2; no feature activates a human microphone.

Optional room assistance

Apply migration 345_team_voice_assistance.sql and deploy both the web app and ClapilotAICore. No additional shared media-server configuration is needed. Configure speech-to-text under ClapilotAICore audio settings. The speaking agent uses the configured realtime voice model and streams audio directly, including interruption when a human speaks. It does not require the separate text-to-speech service. In Settings → ClapilotAICore → Audio → Realtime models, add provider/model pairs and star one preferred model, using the same catalog pattern as image, video and speech models. + Add agent lists those exact realtime models; the room's Voice settings → Live model selector can change the choice or follow the preferred model. Explicit selections stay pinned when the preferred model changes. Disabled providers and removed models are unavailable rather than silently replaced. OpenAI/Gemini entries must be realtime speech models; compatible gateways can use administrator-defined deployment aliases. Migration 346_realtime_model_catalog.sql seeds existing configured defaults and supported OpenAI/Gemini entries once. Provider quota, authentication and unavailable-model errors are shown in the room without exposing provider response bodies or credentials. OpenAI Realtime, GPT-Live (through the shared protocol adapter), Gemini Live and configured OpenAI-compatible/Azure realtime endpoints use their native protocols.

The speaking agent delegates workspace lookups and explicit actions to normal ClapilotAICore runs, with the speaking user's identity and the existing Team Chat permission/approval policy. This tool access also works when the text agent is disabled. Attribution comes from the subscribed human audio track and its verified speech segment, never a user ID suggested by the voice model. Overlapping speakers, changed requests and unavailable transcripts require clarification before tools run. The room lease and speaker access are checked before execution and before publishing results.

The text agent evaluates speaker-labelled transcripts when used alone and stays silent when it cannot contribute something useful. When both agents are enabled, voice delegation owns execution and the text agent option adds the full written backend result; a second agent does not independently repeat the same action. The speaking agent summarizes verified results briefly. Human speech clears local playback; provider-native audio interruption handles barge-in without injecting conversational “stop speaking” commands. Listening silently produces no announcement. Translation uses the configured text model with a validated translation-only response. These shared runtime rules apply to web, iOS and macOS participants.

Each completed utterance appears as a speaker-attributed message in the room chat. Its translation updates that same message, without dispatching ordinary chat agents a second time. The Conversation transcript icon opens a separate scrollable popup on web and macOS, and a full-screen reading view on iPhone. History loads in pages of sixty entries with Load earlier entries; live entries continue to update while reading. Apply migration 348_voice_transcript_chat_messages.sql before deploying. Existing transcript history stays available in the popup; only newly processed utterances are published to chat. Raw microphone audio is mixed on a fixed clock for the live provider and processed in bounded per-speaker utterances for transcription and is not retained as a recording. Enabled transcription sends those utterances to the instance's configured STT provider. The speaking agent additionally sends live microphone audio to the selected realtime provider. Captured transcripts and automatic replies are not mirrored to mapped external messaging channels. Turning assistance off stops new processing; existing transcript text remains room history.

ClapilotAICore runs one visible Clapilot AI media participant per active assisted room. PostgreSQL leases prevent duplicate listeners across runtime replicas. Empty rooms, revoked room permissions, configuration changes and expired leases stop the worker. Room revisions fence in-flight transcription and output. Worker audio subscriptions exclude AI participants to avoid feedback. Startup failures back off for one minute; changing room options retries immediately. The optional LiveKit Node RTC SDK is loaded only when assistance is enabled and participants are present; transport and provider failures leave human voice available and show an assistance error.

Bounded processing supports up to sixteen simultaneous microphone subscriptions and four pending utterances per room. Overload is surfaced as unavailable assistance rather than retaining unbounded audio. There is no automatic replay of captured audio after a worker restart.

Deployment

Apply migration 343_team_chat_voice_rooms.sql before running the updated app. Existing rooms default to text. Configure these server environment variables and restart the web app:

  • TEAM_VOICE_URL: client-reachable wss:// LiveKit endpoint (localhost testing may use ws://).
  • TEAM_VOICE_API_KEY and TEAM_VOICE_API_SECRET: a LiveKit key pair, kept server-side.
  • TEAM_VOICE_INTERNAL_URL: optional server-only HTTP(S) endpoint for room administration; defaults to the public endpoint with HTTP(S) scheme.

Without configuration, text chat works and joining voice returns a localized setup message. No cloud tenant or third-party endpoint is selected automatically. Use a dedicated deployment, or share a server between instances managed by the same operator using separate API keys. Provision trusted TLS, reachable WebRTC UDP/TCP ports, and TURN/TLS for restrictive networks. An HTTP tunnel alone does not relay WebRTC media. See LiveKit deployment and ports.

The docker-compose.voice-local.yml overlay provides an opt-in localhost development server with development-only keys. Never expose those keys publicly or use that overlay on production. Production remains image-based; deploy the web change through the normal main/image/Watchtower path and provision the media endpoint separately.

Persistent self-hosted service

The base Compose file includes a pinned team-voice service behind the voice profile, with restart: unless-stopped. For instances that already use the managed Cloudflare tunnel, run:

node scripts/team-voice-setup.mjs --node-ip <host-ipv4> --turn-domain <turn-tls-hostname>
docker compose -f docker-compose.yml -f docker-compose.local-build.yml up -d --build clapilot
docker compose -f docker-compose.yml -f docker-compose.local-build.yml up -d team-voice

These commands are for a local source-built instance. On production, use the base image-based Compose file after the updated image has been published. The setup script generates a random API key pair in the ignored .env, enables the persistent Compose profile, and adds voice-<instance-hostname> to the existing tunnel and DNS. It preserves other tunnel routes and existing keys on subsequent runs. The standard Cloudflare setup script also preserves the voice route. No credentials are printed or committed.

Signaling uses HTTPS through Cloudflare. Direct media binds only to the supplied host IPv4 address on UDP 7882 and TCP 7881. Internet access requires either reachable media ports or TURN/TLS. The embedded TURN origin is bound to host loopback port 7883 and requires a TLS terminator on public port 443 that sends PROXY protocol v2. Only the configured media host is allowed as a private TURN peer. A configured TURN hostname alone does not start a TLS terminator.

For a development machine already enrolled in Tailscale, an optional persistent Funnel can provide that TLS endpoint without a router port forward:

tailscale funnel --bg --proxy-protocol=2 --tls-terminated-tcp=443 tcp://127.0.0.1:7883

Use the node's *.ts.net hostname as --turn-domain. Tailscale must remain running; the background Funnel resumes after a restart. This relay choice is optional and specific to the instance: other deployments can use their own public media host/TLS terminator. Clients do not need Tailscale. Validate two-way audio with iceTransportPolicy: 'relay', verify the selected candidate's relayProtocol is tls, and verify the public ingress path before describing an Internet deployment as verified. Public Funnel DNS can take up to ten minutes to appear; test with a public resolver so a private tailnet address cannot mask this delay. See Tailscale TCP forwarding and Funnel requirements.

Room tokens expire after 60 seconds and permit subscribing and microphone publishing in exactly one room. Private-room listening requires active membership, including for workspace admins. Room updates rotate the media-room revision and disconnect the old conversation; participants rejoin under the new permissions. Existing clients recheck access every 15 seconds and disconnect when permissions, room revision, or connectivity change. No tokens are returned through agent tools or stored in chat history.

Shared public server

Instances managed by the same operator can use one public voice server. Each instance receives its own LiveKit API key and secret. Media-room names include a namespace derived from the API key, so identical database room IDs on different instances cannot accidentally join or revoke each other's media rooms. Room membership and text chat remain in each instance's database. LiveKit API keys are administrative server credentials: this is a shared operator trust boundary, not hard administrative isolation between untrusted tenants. Use separate media deployments for independently administered tenants.

deploy/team-voice/compose.yml provides LiveKit and Coturn on a Linux server with a public IPv4 address. HTTPS signaling uses the existing Nginx server; raw media uses UDP 7882 and TCP 7881. Coturn uses UDP/TCP 3478 and TURN/TLS 5349, with UDP relay ports 40000–40999. Open those ports in both the host and hosting-provider firewall. TCP 7880 stays bound to loopback. The TURN hostname must use DNS-only addressing, not a Cloudflare HTTP proxy. Participants need no VPN or extra application. Networks allowing only TCP 443 need a separate TURN listener on 443; the supplied configuration preserves an existing web server's use of that port.

Generate private server configuration and one environment snippet per instance:

sudo python3 scripts/shared-team-voice-config.py \
  --directory /opt/clapilot-voice --public-ip <public-ipv4> \
  --hostname <voice-hostname> --turn-hostname <turn-hostname> \
  --instances prod,dev,alx,tme

The generator preserves existing keys and writes secrets only into protected files outside the repository. Copy deploy/team-voice/compose.yml into that directory. Configure the HTTP challenge location from nginx.conf.example and issue a trusted certificate for both hostnames using the certbot/certbot:v5.8.0 image, webroot /var/www/acme mounted from /var/lib/clapilot-voice-acme, and certificate name clapilot-shared-voice. Keep ACME state under /opt/clapilot-voice/letsencrypt, acme-work, and acme-logs. Enable the TLS Nginx server only after the certificate exists and nginx -t passes.

Create /opt/clapilot-voice/turn-tls with owner root:65534, mode 0750. Copy the issued fullchain.pem and privkey.pem there with owner root:65534, mode 0640; the official Coturn image runs as UID/GID 65534. Start the image-based stack with docker compose -f /opt/clapilot-voice/compose.yml up -d. Install the renewal script in /opt/clapilot-voice and its systemd service/timer from deploy/team-voice. Renewal updates Coturn's private certificate copies and reloads Nginx and Coturn without restarting active conversations.

Apply only the four TEAM_VOICE_* lines from each private instance-env/<instance>.env file to the matching instance, then recreate its web container using its existing deployment method. All instances use the shared HTTPS endpoint for both client signaling and server administration. Never copy an instance's key to another instance. Future instances can be added by rerunning the generator with their name and recreating LiveKit after arranging for any active rooms to finish. Verify two-way audio through the public server, force a TURN/TLS connection, and verify instance separation before completing a rollout.

Transport evaluation and ownership

The chosen boundary is an optional LiveKit SFU with pinned JavaScript client/server and Swift SDKs, under Apache-2.0. It supports self-hosting and replaces custom signaling, peer negotiation, NAT traversal, reconnection, and native audio routing code. It owns transient media only; PostgreSQL remains the source of room access and text history. Removing or losing the media service leaves ordinary Team Chat functional. No agent loop, memory, or job dependency is introduced. The adapter is contained in src/lib/team-voice.ts, the web voice component, and the native voice view. A replacement transport can keep the room schema and authenticated join API.

Acceptance checks: two different accounts in one room exchange audio; both join muted; muting/leave stops publishing; private-room outsiders and archived rooms are denied; room changes disconnect old media; text history remains available. Internet/TURN and physical-device audio should also be verified against the deployment used by the team. Codex subscription voice uses the shared realtime catalog with an enabled OpenAI-Codex provider and gpt-live-1-codex. Its server-side PCM adapter uses WebRTC and defaults to cove; startup allows up to 46 seconds. The room keeps the same optional text-agent boundary and permissions. An unavailable subscription does not switch to Platform API billing; account eligibility still requires a real call.