Environment
- Self-hosted stack from
deploy/self-hosting (v1 images pulled 2026-07-03), Ubuntu 24.04, Docker 29.x, default bridge networking, Vultr VPS - Voice region/server configured with the documented public endpoint form:
wss://<domain>/livekit
Problem
LiveKitService derives the server-API URL from the client endpoint (
wss:// →
https://), so the
api and
worker containers make RoomService calls to the instance's
own public hostname. On hosts where Docker's published-port DNAT does not hairpin for containers on the same bridge as
caddy, those calls time out:
"LiveKit listParticipants failed" … fetch failed: Connect Timeout Error (attempted address: <domain>:443, timeout: 10000ms)
"Error disconnecting LiveKit participant" … (same timeout)
curl https://<domain>/api/_health from the host: 200 in ~30 ms- same fetch from inside the api container: timeout (DNS resolves correctly to the public IP; the TCP path is what fails)
User-visible consequences
- Voice reconciliation cannot clean up zombie voice sessions ("Cleaned up zombie voice connections … disconnectedCount: 0" + repeated errors)
- A user whose previous session became a zombie gets "Couldn't connect to voice. Please try again." on every rejoin — while direct client connections (signaling + UDP media) work fine, which makes this confusing to diagnose
Workaround that fixed it cleanly
Give
caddy a network alias equal to the public hostname, so in-network containers resolve it straight to Caddy over the bridge (TLS stays valid because Caddy serves the real certificate):
services:
caddy:
networks:
fluxer:
aliases:
- chat.example.com
After this: server-API calls return 200, reconciliation sweeps report
liveKitDiscoveryErrors: 0, zombie sessions get cleaned, and rejoin works.
Suggestion
Either document this alias in the operator guide (it is zero-cost and safe), or support a separate internal endpoint for server-side LiveKit API calls (e.g. resolving to
http://livekit:7880 in-network) so self-hosted deployments do not depend on hairpin NAT behavior.