Formantic Intelligence

Selected work

The kinds of problem we take on.

  1. Adversarial Audio Perturbations for Privacy

    Developed audio perturbation algorithms that raised word error rate from 2% to 30% on commercial ASR systems (Deepgram, Assembly AI, GPT-4o), while the perturbed audio remained intelligible to human listeners. The work was done for a US seed stage startup.

  2. Fine-tuning of whisper models and evaluation infrastructure

    Fine-tuned Whisper models for Indian languages to <2% WER for same-language subtitling and ~40% BLEU for cross-language subtitle translation. Built the accuracy and time-alignment evaluation framework that drove model selection. Built a GPU inference system subtitling 1 hr of audio in under a minute on 4 GPUs, plus cost-optimized video/audio processing (voice activity detection, transcoding, 50+ subtitle styles and animations). Deployed across Azure and AWS.

  3. Speech recognition

    Trained a TDNN speech recognition model (<8% WER) on combined open-source and proprietary data at Motorola Solutions, deployed in production for real-time transcription of 911 calls in CommandCentral software used in ~60% of all 911 calls in the US.

  4. Speaker Diarization

    Trained an x-vector based speaker diarization system on ~10,000 hrs of audio and developed novel synthetic data generation methods for training. Shipped in two Motorola Solutions products; significantly outperformed every commercial model and API evaluated on our data. Also built its production CPU-only inference system.

Have a hard problem?

We'd like to hear about it.

hello@formanticintelligence.com