Dhruv Dholakia
Dhruv founded AsqVox, a voice AI widget that lets a visitor ask a website a question out loud and get an answer grounded in that business own documents. He writes and reviews the glossary published here.
What he writes about
- Voice agents and the pipeline behind them
- Retrieval and grounding for spoken questions
- Latency budgets and conversational timing
- What happens to websites after search stops sending clicks
The glossary exists because most voice AI vocabulary is defined by the companies selling it. Every entry here leads with a plain definition, carries statistics with their source and year, labels vendor numbers as vendor numbers, and says plainly where no reliable figure exists rather than borrowing one.
Recently reviewed
AI disclosure and the EU AI Act
AI disclosure means telling people they are interacting with an AI rather than a human, and the EU AI Act turns that into a legal obligation in defined circumstances when its Article 50 transparency provisions become enforceable on 2 August 2026.
Answer engine optimization
Answer engine optimization is the practice of writing and structuring content so an AI answer engine picks it up and cites it, which means competing for a mention inside someone else's answer rather than for a ranking position on a results page.
Audio codecs
An audio codec is the scheme that squeezes sound into a form a network can carry, and whichever one a voice system runs on settles how the audio sounds, what it costs in bandwidth, and how much of it a recognizer can actually transcribe.
Watermarking of synthetic audio
Watermarking of synthetic audio embeds an imperceptible signal into generated speech so it can later be identified as machine-made, and it is the main technical answer to the question of how anyone will know what was synthesized.
Automatic speech recognition
Automatic speech recognition (ASR), also called speech to text, is the technology that turns spoken audio into written text, and it is the first real step in almost every voice agent.
Barge-in
Barge-in is a voice agent stopping its own speech the instant a person starts talking over it, which is what makes the exchange feel like a conversation rather than a recording you have to sit through.