Edit history

Self-hosting: API→LiveKit server-API calls fail on hosts where Docker cannot hairpin to the instance's own public hostname has not been edited, so there are no earlier versions.

Current version | Original by Rex
Show

Self-hosting: API→LiveKit server-API calls fail on hosts where Docker cannot hairpin to the instance's own public hostname

Environment

  • Self-hosted stack from deploy/self-hosting (v1 images pulled 2026-07-03), Ubuntu 24.04, Docker 29.x, default bridge networking, Vultr VPS
  • Voice region/server configured with the documented public endpoint form: wss://<domain>/livekit

Problem

LiveKitService derives the server-API URL from the client endpoint (wss:// → https://), so the api and worker containers make RoomService calls to the instance's own public hostname. On hosts where Docker's published-port DNAT does not hairpin for containers on the same bridge as caddy, those calls time out:
"LiveKit listParticipants failed" … fetch failed: Connect Timeout Error (attempted address: <domain>:443, timeout: 10000ms)
"Error disconnecting LiveKit participant" … (same timeout)
  • curl https://<domain>/api/_health from the host: 200 in ~30 ms
  • same fetch from inside the api container: timeout (DNS resolves correctly to the public IP; the TCP path is what fails)

User-visible consequences

  • Voice reconciliation cannot clean up zombie voice sessions ("Cleaned up zombie voice connections … disconnectedCount: 0" + repeated errors)
  • A user whose previous session became a zombie gets "Couldn't connect to voice. Please try again." on every rejoin — while direct client connections (signaling + UDP media) work fine, which makes this confusing to diagnose

Workaround that fixed it cleanly

Give caddy a network alias equal to the public hostname, so in-network containers resolve it straight to Caddy over the bridge (TLS stays valid because Caddy serves the real certificate):
services:
  caddy:
    networks:
      fluxer:
        aliases:
          - chat.example.com
After this: server-API calls return 200, reconciliation sweeps report liveKitDiscoveryErrors: 0, zombie sessions get cleaned, and rejoin works.

Suggestion

Either document this alias in the operator guide (it is zero-cost and safe), or support a separate internal endpoint for server-side LiveKit API calls (e.g. resolving to http://livekit:7880 in-network) so self-hosted deployments do not depend on hairpin NAT behavior.