Sunday, August 2, 2026

How LLMs “hear” audio

For LLMs, audio isn’t waves or sound—it’s transformed into tokens the model can understand. The raw waveform is first converted into spectrograms (a visual-like representation of sound across time and frequency) or encoded using neural codecs. These representations are then compressed into discrete units that act like “audio tokens.” Each token is mapped into an embedding vector, capturing characteristics such as pitch, rhythm, or phoneme structure. The sequence of embeddings allows the model to “hear” patterns—like distinguishing between speech, music, or environmental sounds. With speech, the embeddings align with linguistic units (phonemes, syllables, words), enabling transcription, translation, or even emotion detection. This pipeline—waveforms → tokens → embeddings → meaning—gives models the ability to not just transcribe speech but also understand and reason across audio, bridging sound and language in a unified framework.

No comments:

About Me

My photo
Me? I have always been allured by the Indian Tech Dream. The dream that Indian Tech would provide for a global platform to every aspiring Indian wishing to make a mark & showcase world class software solutions. Yes, I am in pursuit of this dream, with a hope of touching life’s of people in a positive way through software solutions. This blog is primarily to express my work in Tech, the work I have done, the work I have contributed, the work that I have seen, that I term as outstanding software development.