Polyphon AI

Turn multi-speaker audio into named, structured intelligence — 100% offline

Developer ToolsWeb
#113Day Rank
Visit website

What it does

Polyphon AI is a local, identity-aware transcription and diarization engine. It turns multi-speaker audio into structured, named intelligence — entirely offline, with zero telemetry and no cloud APIs.

Pipelines built on faster-whisper and pyannote leave three problems unsolved: naive timestamp overlap breaks when people interrupt or talk over each other; every file restarts from SPEAKER_00 with no memory across meetings; and forced alignment desyncs when speakers code-switch mid-sentence.

Polyphon fixes all three. Acoustic embedding proximity resolves ambiguous words instead of box overlap. VoiceDB stores persistent voiceprints, so SPEAKER_00 becomes Alice across every recording. Multilingual speech works natively.

Ships with a web workspace, real-time neural diarization, local LLM summarization via Ollama, and a CLI. Runs on commodity hardware.

Product details

Profile

Maker
See Hiong
Platforms
Web
Website
github.com

FAQ

Product questions

Community

Discuss and review Polyphon AI

Read maker updates, ask questions, or leave product feedback for other builders evaluating this launch.

Developer Tools · 0 discussionsSep 7, 2026