
Audio Visual & Text Filtering: Why the Source of a Mention Changes Everything
Text-First, API-Limited Tools Are Built for a World That No Longer Exists
Imagine this: three mentions of your brand pop up at the same time. One is a hashtag tucked into a caption. Another is someone mentioning you out loud, deep into a podcast episode. The third? Your logo flashes for just a couple of seconds in the background of a TikTok. Most monitoring tools will spot one and miss the other two completely. And even if they do catch all three, they usually treat them like they're exactly the same.
But not all mentions are created equal. Where a mention comes from, and how it’s shared, can totally change what it means for your brand.
Why Format Isn’t the Same as Signal
Most social listening tools were designed for a world where everything happened in text: captions, hashtags, headlines, and comments. That made sense at first, but these days, it’s just a small piece of the puzzle. More and more, the conversations that matter most are happening out loud, in visuals, or in on-screen text that never makes it into a caption.
If your tools only read captions, you’re missing out on quantity, and, you’re missing out on the real, unfiltered thoughts people share out loud or show on screen, instead of carefully typing and editing.
What Changes Depending on the Source
Not every mention is created equal. If you treat them all the same, you risk missing the important stuff or letting the real gems get lost in the noise.
Written and captioned mentions are the easiest to spot and double-check, but they’re also the most polished. People edit captions all the time. But what someone says out loud, mid-sentence, in a 40-minute video? That’s usually straight from the heart.
Spoken mentions, picked up by Automatic Speech Recognition (ASR), are usually more honest and off the cuff. Think of a creator sharing their real thoughts in a product review, a podcast host chatting about the latest news, or someone diving deep in a YouTube essay. This is where people say things they might never write down and where you’ll often catch the first hint of a disclosure, a bold claim, or even a brand-damaging comment.
Visual and on-screen text mentions, found using Optical Character Recognition (OCR) and Computer Vision, are a whole different story. Maybe your logo pops up in the background of a video, or someone shares a screenshot of a text conversation, or there’s a handwritten sign at a protest. None of this appears in a caption, and it’s not “written” in the usual way, but it can be the clearest proof of a brand connection you’ll ever see.
Here’s what the numbers show: about 75% of brand mentions in video content live only in the spoken words, not in the title, description, or hashtags. On Instagram, 66% of brand mentions can only be found with OCR or AI-generated image captions. That means two out of three things people say about your brand on Instagram are invisible if your tool only reads text fields.
The Three Technologies Behind Real Multimodal Filtering
To really get the full picture, you need three types of AI working together on the same content, all at once, not in separate corners.
- ASR (Automatic Speech Recognition) transcribes spoken content across podcasts, YouTube videos, TikTok voiceovers, and livestreams, making every spoken word searchable in the same way a caption is.
- OCR (Optical Character Recognition) reads text that appears visually within a frame, screenshots, signage, on-screen graphics, embedded captions, and content that exists only as pixels until OCR turns it into searchable text.
- Computer Vision identifies logos, products, faces, and visual context even when nothing is written or said at all.
Pendulum brings all three together in one Snippet System. That way, a spoken phrase from a podcast, a logo spotted in a video, and text from a screenshot all show up in one place no more jumping between different dashboards.
Why Filtering by Demographic and Geography Matters Once You Can See Everything
Multimodal detection helps you see what’s out there. But the next step is figuring out who’s saying it, and where. Once you can spot every kind of mention, filtering by geography and demographics turns "we found the conversation" into "we know who’s behind it and why it matters."
This difference really matters when you’re trying to separate real insights from background noise, or decide if an alert needs your attention. A spoken mention from a small, local group is a whole different story than the same words shared by a big account with national reach. If your system only looks for keywords, it can’t tell the difference.
Source-Aware Filtering Is What Makes Monitoring Agentic
This is where filtering by source type ties into the bigger move from static rules to agentic monitoring. A rules-based system can flag a keyword, sure. But it can’t tell the difference between a spontaneous comment in a podcast and a carefully crafted caption, or spot that a protest sign caught by OCR is a bigger deal than the same words in a tweet. Real judgment comes from understanding both the source and the context; this is what agentic AI really means for brand intelligence.
Treating every mention like it’s the same isn’t just a small oversight. It’s why so many monitoring systems either flood teams with unhelpful alerts or miss the one mention that really matters just because it was spoken, not typed.
See How Your Social Listening Measures up With Our Free Social Listening Audit
Related Articles
Explore deeper insights and practical perspectives related to this topic.
Times Have Changed. So Should Your Tech Stack
See how brands are upgrading their strategies with Pendulum Social Intelligence.




