How the song finder listens
A single guess from a single moment of audio is unreliable: talking, an effect or a quiet intro can make any recognition engine answer with something unrelated. So instead of asking once, the song finder cuts your clip into several short windows across its length and prepares each one in more than one way – as recorded, with the level evened out, and with the bass and hiss filtered away. It then asks about them in parallel.
A song is announced only when two independent windows agree on it. If only one window hears something, or two hear different songs, you get the candidates as buttons and pick – rather than being told a confident wrong answer.
Silent stretches are skipped instead of wasting a lookup, and the whole search is time-boxed so you are not left waiting on one slow reply.
More than one engine
Shazam is fast and has an enormous catalogue, but no catalogue has everything – and Shazam's gaps are largest for Persian and other regional music. In Telegram, when the first pass has not settled it, the bot races other engines alongside: ACRCloud, AudD and AcoustID, which fingerprint the recording differently, and a lyrics engine that transcribes the words being sung and searches for them. Their answers are merged by title, and agreement between engines counts just like agreement between windows.
What happens next
Knowing the name is half of it. In Telegram the bot goes on to find that song – the studio recording, checked against the catalogue – and sends it to you as a tagged file with its cover art and a Lyrics button. On the web you get the same answer with buttons to download it, read the lyrics or open it in the bot.
Where the music is in a video
If the song is playing in a video – a reel, a TikTok, a YouTube clip – you don't need to record it. See find a song from a video: send or upload the video itself and the soundtrack is identified directly.
Tips for a good match
- Get close to the speaker and away from conversation.
- Record the chorus if you can – it is the most distinctive part.
- 10–15 seconds is enough; more does not hurt.
- For live performances and covers, recognition finds the original only when the performance is close to it.