
Polyphon AI
Turn multi-speaker audio into named, structured intelligence — 100% offline
Gallery
Polyphon AI screenshots and media
Product screenshots, media, and video shared by the submitter.
What it does
Polyphon AI is a local, identity-aware transcription and diarization engine. It turns multi-speaker audio into structured, named intelligence — entirely offline, with zero telemetry and no cloud APIs.
Pipelines built on faster-whisper and pyannote leave three problems unsolved: naive timestamp overlap breaks when people interrupt or talk over each other; every file restarts from SPEAKER_00 with no memory across meetings; and forced alignment desyncs when speakers code-switch mid-sentence.
Polyphon fixes all three. Acoustic embedding proximity resolves ambiguous words instead of box overlap. VoiceDB stores persistent voiceprints, so SPEAKER_00 becomes Alice across every recording. Multilingual speech works natively.
Ships with a web workspace, real-time neural diarization, local LLM summarization via Ollama, and a CLI. Runs on commodity hardware.
Product details
Profile
- Maker
- See Hiong
- Category
- Developer Tools
- Platforms
- Web
- Website
- github.com
Submitted by
- Launch date
- Launch tags
- Pricing
- Not specified
FAQ
Product questions
Polyphon AI is a local, identity-aware transcription and diarization engine. It turns multi-speaker audio into structured, named intelligence — entirely offline, with zero telemetry and no cloud APIs. Pipelines built on faster-whisper and pyannote leave three problems unsolved: naive timestamp overlap breaks when people interrupt or talk over each other; every file restarts from SPEAKER_00 with no memory across meetings; and forced alignment desyncs when speakers code-switch mid-sentence. Polyphon fixes all three. Acoustic embedding proximity resolves ambiguous words instead of box overlap. VoiceDB stores persistent voiceprints, so SPEAKER_00 becomes Alice across every recording. Multilingual speech works natively. Ships with a web workspace, real-time neural diarization, local LLM summarization via Ollama, and a CLI. Runs on commodity hardware.
Engineers, researchers, and teams handling confidential or multilingual audio who need accurate multi-speaker transcripts that never leave their machine.
Every transcription tool forgets who people are. You run a meeting through Whisper, get back SPEAKER_00 and SPEAKER_01, and manually relabel them — then do it again next week for the same team, because nothing persists. Meanwhile the cloud options that do handle identity require uploading confidential audio to someone else's servers. That's a non-starter for medical notes, legal work, or anything under NDA. And existing local pipelines break in the cases that matter most: people interrupting each other, short backchannels like "yeah" and "right" landing on the wrong speaker, and multilingual conversations where someone switches from English to Mandarin mid-sentence. Polyphon solves identity persistence, reconciliation accuracy, and multilingual robustness — without a single byte leaving your machine.
Community
Discuss and review Polyphon AI
Read maker updates, ask questions, or leave product feedback for other builders evaluating this launch.
Discovery paths



